# Claude Opus 5

Anthropic · Proprietary · rank 3 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/claude-opus-5  
JSON: https://modelscale.dev/api/model/claude-opus-5

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `claude-opus-5` | — |
| Overall score | 81.87 | Observed 2026-09-22 · source benchlm:models |
| Context window | Unavailable | Unavailable · source benchlm:models |
| Release date | 2026-07-24 | Observed 2026-09-22 · source benchlm:models |
| Access type | Proprietary | — |
| Blended $/1M (75% input / 25% output) | $10.00 | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 85.6 | Observed 2026-09-22 · source benchlm:models |
| Coding | 90.4 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 97.5 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 75.7 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 88.8 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | Unavailable | Unavailable · source benchlm:models |
| Math | Unavailable | Unavailable · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 49.25 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 59 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $5.00 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $25.00 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | $0.50 | Observed 2026-09-22 · source benchlm:pricing |
| Cache write / 1M tokens | $6.25 | Observed 2026-09-22 · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $10.00 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** Anthropic's July 24, 2026 Claude Opus 5 announcement lists the claude-opus-5 API model at $5 input / $25 output per million tokens, the same base price as Opus 4.8. Fast mode runs around 2.5× the default speed and is available at twice the base price. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $28.16 |
| Modelled tokens | 2.82M |

## Benchmark record

97 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| HealthBench (raw) (HealthBench raw score) | 67.1 | 5,000 multi-turn patient conversations | Realistic healthcare conversations | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| HealthBench (length-adjusted) (HealthBench length-adjusted score) | 57.8 | 5,000 multi-turn patient conversations | Realistic healthcare conversations | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| HealthBench Professional (raw) (HealthBench Professional raw score) | 73.4 | 525 physician-authored conversations | Professional clinical tasks | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| BioMysteryBench (human-solvable) (BioMysteryBench Human Solvable) | 90.1 | Human-solvable computational biology investigations | Expert computational biology | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| BioMysteryBench (human-difficult) (BioMysteryBench Human Difficult) | 49.4 | Human-difficult computational biology investigations | Frontier computational biology | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| SpatialBench Verified (LatchBio SpatialBench Verified) | 72.5 | 115 externally validated spatial transcriptomics problems | Professional bioinformatics | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| SingleCellBench (LatchBio SingleCellBench) | 60.6 | 195 single-cell RNA sequencing problems | Professional bioinformatics | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| ProteinGym Hard | 47.7 | Hard protein mutation-effect ranking tasks | Computational protein science | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| Protein Design (Anthropic Protein Design evaluation) | 42.5 | Constrained protein-sequence design tasks | Computational protein design | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| Organic chemistry V2 (Anthropic Organic Chemistry V2 evaluation) | 61.6 | Organic chemistry reasoning tasks | Expert organic chemistry | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| Protocols (troubleshooting) (Molecular Biology Protocols Troubleshooting) | 61.1 | Molecular-biology protocol troubleshooting | Expert laboratory protocols | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| Protocols (understanding) (Benchling Molecular Biology Protocols Understanding) | 78.4 | Molecular-biology protocol extension | Expert laboratory protocols | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| HLE (Humanity's Last Exam) | 64.7 | Expert-level questions | Frontier expert level | [Humanity's Last Exam](https://lastexam.ai/) |
| HLE-Verified | 54.4 | 1,811 verified or revised expert questions | Frontier multidisciplinary expert reasoning | [HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam](https://arxiv.org/abs/2602.13964) |
| LABBench2 (LABBench2: An Improved Benchmark for AI Systems Performing Biology Research) | 84.2 | Nearly 1,900 biology-research tasks | Real-world biology research | [LABBench2: An Improved Benchmark for AI Systems Performing Biology Research](https://arxiv.org/abs/2604.09554) |
| Artificial Analysis Intelligence Index | 50.8 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 93.2 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 54.9 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | 37.1 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 60.9 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 60.8 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| HealthBench Professional | 59.8 | Clinician chat tasks | Professional clinical workflows | [HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats](https://arxiv.org/abs/2604.27470) |
| HLE w/o tools (Humanity's Last Exam without tools) | 56.3 | Expert-level questions | Frontier expert level | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| GPQA Diamond (Vals) (GPQA Diamond, Vals AI run) | 93.4 | Graduate-level science questions | Expert reasoning | [Vals AI GPQA Diamond, Vals AI run leaderboard](https://www.vals.ai/benchmarks/gpqa) |
| MMLU-Pro (Vals) (MMLU-Pro, Vals AI run) | 91.6 | Academic multiple-choice questions | Broad academic knowledge | [Vals AI MMLU-Pro, Vals AI run leaderboard](https://www.vals.ai/benchmarks/mmlu_pro) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| ProgramBench (episode 1) (ProgramBench hidden-test pass rate after episode 1) | 83.0 | 166 golden program-reconstruction tasks | Long-context clean-room software engineering | [ProgramBench: Can language models rebuild programs from scratch?](https://arxiv.org/abs/2605.03546) |
| SWE-bench Verified (Software Engineering Benchmark Verified) | 96 | 500 verified issues | Professional software engineering | [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770) |
| SWE-bench Pro | 79.2 | 1,865 repository problems | Long-horizon professional engineering | [SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?](https://arxiv.org/abs/2509.16941) |
| VulcanBench v3 | 87.0 | 23 post-cutoff repository tasks in the v3 report | Professional multi-file software engineering | [VulcanBench](https://github.com/morganlinton/VulcanBench/tree/main) |
| VulcanBench CII v1 (VulcanBench Coding Intelligence Index v1) | 96.4 | 38 validated post-cutoff repository tasks | Mid-band frontier software engineering | [CII v1 frontier results](https://github.com/morganlinton/VulcanBench/blob/main/docs/results/cii-v1-2026-08/README.md) |
| FrontierCode 1.1 Main | 53.4 | 100 private Main tasks (150 in Extended) | Frontier coding-agent quality | [FrontierCode leaderboard](https://cognition.com/frontiercode) |
| FrontierCode 1.1 Extended | 63.6 | 150 private software-engineering tasks | Frontier coding-agent quality | [GPT-5.6 models are now available in Devin](https://devin.ai/blog/gpt-5-6) |
| SWE Multilingual | 89.5 | Multilingual software-engineering tasks | Professional software engineering | [MiniMax M2.7: Early Echoes of Self-Evolution](https://www.minimax.io/news/minimax-m27-en) |
| SWE Multimodal (SWE-bench Multimodal) | 59.4 | Multimodal software engineering tasks | Frontier multimodal coding | [SWE-bench Multimodal](https://www.swebench.com/multimodal) |
| ProgramBench (ProgramBench: Can Language Models Rebuild Programs From Scratch?) | 93.0 | 200 program reconstruction tasks | Full-repository software architecture | [ProgramBench: Can Language Models Rebuild Programs From Scratch?](https://programbench.com/static/paper.pdf) |
| FrontierSWE v2 | 52.0 | 34 ultra-long-horizon engineering and research tasks | Ultra-long-horizon frontier software engineering | [FrontierSWE v2](https://www.frontierswe.com/blog/v2) |
| Bug Hunt Bench | 27 | 105 planted bugs across two production TypeScript repositories | Blind production-repository bug finding and repair | [Bug Hunt Bench method, caveats, and definitions](https://bughunt.productcompass.pm/method) |
| AA Coding Index (Artificial Analysis Coding Index) | 78.0 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 56.4 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| LiveCodeBench (Vals) (LiveCodeBench, Vals AI run) | 89.0 | Competitive programming problems (easy, medium, hard) | Frontier coding | [Vals AI LiveCodeBench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/lcb) |
| SWE-bench (Vals) (SWE-bench, Vals AI run) | 97.0 | Real repository issues by human time bucket | Frontier coding agents | [Vals AI SWE-bench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/swebench) |
| DeepSWE | 68.8 | 113 software engineering tasks across 91 repositories and 5 languages | Long-horizon software engineering | [DeepSWE benchmark blog](https://deepswe.datacurve.ai/blog) |

### Mathematics

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| IMO 2026 (International Mathematical Olympiad 2026) | 42 | 6 proof-based problems | International olympiad mathematics | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| RiemannBench (no tools) (RiemannBench without tools) | 60.0 | 25 private research-level mathematics problems | Research mathematics | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| RiemannBench (tools) (RiemannBench with tools) | 79.0 | 25 private research-level mathematics problems | Research mathematics | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| ArXivMath Jun. 2026 (no tools) (ArXivMath June 2026 without tools) | 90.8 | 49 recent research-mathematics problems | Research mathematics | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| ArXivMath Jun. 2026 (tools) (ArXivMath June 2026 with tools) | 91.3 | 49 recent research-mathematics problems | Research mathematics | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| ARC-AGI-1 (ARC-AGI-1 Semi-Private Evaluation) | 97.50 | Semi-private ARC-AGI-1 evaluation set | Abstract visual reasoning | [ARC Prize leaderboard](https://arcprize.org/) |
| ARC-AGI-2 (Abstraction and Reasoning Corpus for AGI v2) | 90.4 | Visual pattern completion and abstract reasoning | Expert-level — hardest public reasoning benchmark | [ARC-AGI-2: A Harder General Intelligence Benchmark](https://arcprize.org/arc-agi/2/) |
| ARC-AGI-3 (Abstraction and Reasoning Corpus for AGI v3) | 30.16 | Interactive game-like tasks with hidden rules | Frontier agentic reasoning | [ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence](https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf) |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 79.3 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| MLCR-AA (Medical Long Context Reasoning (MLCR-AA)) | 55.6 | Long, fragmented medical-record reasoning | Long-context medical reasoning | [Medical Long Context Reasoning (MLCR-AA)](https://artificialanalysis.ai/evaluations/mlcr-aa) |
| CritPt (Critical Physics Tasks) | 29.1 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Multilingual

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GMMLU (Global MMLU) | 92.5 | Knowledge questions across 42 languages | Multilingual knowledge | [Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation](https://arxiv.org/abs/2412.03304) |
| MILU (Multi-task Indic Language Understanding Benchmark) | 92.1 | Knowledge tasks across 11 languages | Multilingual Indic knowledge | [MILU: A Multi-task Indic language understanding benchmark](https://arxiv.org/abs/2411.02538) |
| INCLUDE | 89.8 | Cross-lingual understanding | Broad multilingual capability | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| DRACO (Data Research and Analysis with Complex Operations) | 88.6 | Agentic data research and analysis tasks | Professional data analysis | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| BrowseComp (10-agent, prerelease) (Multi-Agent BrowseComp — 10-agent team prerelease configuration) | 93.6 | BrowseComp web-research tasks | Long-horizon web research | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| MCP-Atlas claim coverage (MCP-Atlas mean claim coverage) | 89.1 | Production-like multi-server MCP workflows | Real-world tool use | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| LAB all-pass (Anthropic harness) (Legal Agent Benchmark all-pass rate — Anthropic harness) | 23.58 | 1,235 legal-agent tasks | Professional legal work | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| LAB criterion-pass (Anthropic harness) (Legal Agent Benchmark mean criterion-pass rate — Anthropic harness) | 93.74 | 1,235 legal-agent tasks | Professional legal work | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| LAB all-pass (Harvey held-out) (Legal Agent Benchmark all-pass rate — Harvey held-out set) | 11.7 | Harvey-held-out legal-agent tasks | Professional legal work | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| LAB criterion-pass (Harvey held-out) (Legal Agent Benchmark mean criterion-pass rate — Harvey held-out set) | 94.1 | Harvey-held-out legal-agent tasks | Professional legal work | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| Toolathlon Verified Pass@3 | 87.0 | 108 verified real-world tool-use tasks | Long-horizon application tool use | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| Toolathlon Verified Pass³ (Toolathlon Verified Pass cubed) | 73.1 | 108 verified real-world tool-use tasks | Long-horizon application tool use | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| Toolathlon Verified avg. turns (Toolathlon Verified average assistant turns) | 23.5 | 108 verified real-world tool-use tasks | Long-horizon application tool use | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| Terminal-Bench 3.0 | 42.7 | 74 professional computer-work tasks across 7 domains | Frontier autonomous knowledge work | [Terminal-Bench 3.0](https://www.frontierbench.ai/) |
| AA Briefcase (Artificial Analysis Briefcase) | 1720 | Professional knowledge-work tasks | Professional work | [Artificial Analysis Briefcase Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-briefcase) |
| AA AutomationBench (Artificial Analysis AutomationBench) | 56.6 | Business-process automation tasks | Agentic automation | [Artificial Analysis AutomationBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/automationbench-aa) |
| AA EnterpriseOps-Gym (Artificial Analysis EnterpriseOps-Gym) | 47.5 | Enterprise operations workflows | Enterprise agent operations | [Artificial Analysis EnterpriseOps-Gym Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa) |
| AA Harvey LAB (Artificial Analysis Harvey LAB-AA) | 93.5 | Legal agent tasks | Professional legal work | [Artificial Analysis Harvey LAB-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/harvey-lab-aa) |
| AA Tau3 Banking (Artificial Analysis Tau3-Banking) | 42.1 | Banking tool-use workflows | Agentic banking workflows | [Artificial Analysis Tau3-Banking Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/tau3-banking) |
| BrowseComp | 90.8 | Research questions requiring browsing | Hard web research | [BrowseComp](https://openai.com/index/browsecomp/) |
| HLE w/ tools (Humanity's Last Exam with tools) | 64.7 | Expert questions with tool use | Frontier tool-augmented reasoning | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA | 1862 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 60.4 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA Agentic Index (Artificial Analysis Agentic Index) | 56.2 | Cross-benchmark agentic index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA Terminal-Bench 4.0 (Artificial Analysis Terminal-Bench v4.0) | 49.0 | Terminal-based agent tasks | Agentic software engineering | [Artificial Analysis Terminal-Bench v4.0 Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/terminalbench-v4-0) |
| GDP.pdf (Artificial Analysis GDP.pdf) | 21.6 | Professional document-production tasks | Professional knowledge work | [GDP.pdf Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gdp-pdf) |
| AA-AnalystAgent (Artificial Analysis AnalystAgent) | 53.8 | Spreadsheet and document analysis questions | Business and data analysis | [AA-AnalystAgent Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-analyst-agent) |
| OSWorld 2.0 | 70.6 | 108 long-horizon computer-use workflows | Long-horizon professional workflows | [OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks](https://arxiv.org/abs/2606.29537) |
| MCP Atlas | 85.8 | Tool-integrated agent tasks | Advanced tool use | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| Toolathlon-Verified | 80.6 | Verified multi-tool workflows | Advanced tool use | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| AutomationBench | 26.0 | 600 public automation tasks | Long-horizon automation | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| DeepSearchQA | 95.0 | Agentic browsing and list-answer questions | Agentic web research | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 84.6 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |
| ApprenticeBench (ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job) | 36 | 100 vendor bills processed in sequence inside a simulated construction company | Long-horizon computer use with offline and online continual learning | [ApprenticeBench: a step change in AI's job readiness](https://neocognition.io/blog/apprentice-bench/) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Chartography (no tools) (Chartography without tools) | 29.6 | 100 specialized chart types | Professional chart reasoning | [Chartography](https://surgehq.ai/blog/chartography) |
| Chartography (tools) (Chartography with image and code tools) | 83.0 | 100 specialized chart types | Professional chart reasoning | [Chartography](https://surgehq.ai/blog/chartography) |
| BenchCAD Vision2Code (no tools) (BenchCAD Vision2Code voxel IoU without tools) | 0.366 | 1,000-file Vision2Code subset | Programmatic CAD generation | [BenchCAD: A comprehensive, industry-standard benchmark for programmatic CAD](https://arxiv.org/abs/2605.10865) |
| BenchCAD Vision2Code (tools) (BenchCAD Vision2Code voxel IoU with tools) | 0.821 | 1,000-file Vision2Code subset | Programmatic CAD generation | [BenchCAD: A comprehensive, industry-standard benchmark for programmatic CAD](https://arxiv.org/abs/2605.10865) |
| GDP.pdf (no tools) (GDP.pdf mean criteria pass rate without tools) | 83.4 | 100 professional document prompts | Professional document reasoning | [GDP.pdf](https://surgehq.ai/blog/gdp-pdf-can-100b-ai-models-master-the-documents-that-run-the-world) |
| GDP.pdf (tools) (GDP.pdf mean criteria pass rate with tools) | 85.5 | 100 professional document prompts | Professional document reasoning | [GDP.pdf](https://surgehq.ai/blog/gdp-pdf-can-100b-ai-models-master-the-documents-that-run-the-world) |
| OfficeQA | 78.1 | Historical Treasury Bulletin questions | Professional document reasoning | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 84.7 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| Design Arena Website (Design Arena Website Elo) | 1320 | Website generation comparisons | Design and website generation | [OpenRouter Grok 4.3 benchmarks](https://openrouter.ai/x-ai/grok-4.3/benchmarks) |
| OfficeQA Pro | 66.9 | Document and spreadsheet tasks | Enterprise grounded reasoning | [OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning](https://arxiv.org/abs/2603.08655) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

| Event | Type | Confirmation | Announced | Effective | Replacement | Source |
| --- | --- | --- | --- | --- | --- | --- |
| Claude Opus 5 | Release | confirmed | Unavailable | 2026-07-24 | — | [Anthropic](https://www.anthropic.com/news/claude-opus-5) |

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
