Claude Mythos 5
Anthropic · Proprietary · bench-align-v5
Capability shape
Seven axes from the ranking source. A missing axis is drawn as a gap.
Runtime service evidence
Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim.
Time to first token: latency to the first answer chunk (Artificial Analysis, via BenchLM). Reasoning models include thinking time, so values can run to tens or hundreds of seconds.
Evidence key: ObservedUnavailable
| Measurement | Value | Observed | Last good | Evidence |
|---|---|---|---|---|
| Time to first token | Unavailable | Unavailable | Unavailable | |
| Throughput | Unavailable | Unavailable | Unavailable |
Regional or per-endpoint measurements appear only when the API supplies them; none are modelled here.
Endpoint and price matrix
Every published price component, including cache reads and writes.
| Component | USD | Evidence |
|---|---|---|
| Input / 1M tokens | $10.00 | |
| Output / 1M tokens | $50.00 | |
| Cache read / 1M tokens | $1.00 | |
| Cache write / 1M tokens | Unavailable | |
| Blended / 1M (75% input / 25% output) | $20.00 |
Workload-aware monthly cost example
10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. This uses the same calculator as the cost simulator, so an unavailable applicable rate makes the example unavailable too.
Benchmark record
20 matched benchmark rows with their published value, unit, and provenance.
Knowledge
| Benchmark | Value | Tasks | Difficulty | Provenance |
|---|---|---|---|---|
| GPQA Graduate-Level Google-Proof Q&A | 94.1 | 448 questions | Graduate level | GPQA: A Graduate-Level Google-Proof Q&A Benchmark |
| HLE Humanity's Last Exam | 64.5 | Expert-level questions | Frontier expert level | Humanity's Last Exam |
| HLE w/o tools Humanity's Last Exam without tools | 59 | Expert-level questions | Frontier expert level | Introducing GPT-5.4 mini and nano |
Coding
| Benchmark | Value | Tasks | Difficulty | Provenance |
|---|---|---|---|---|
| Terminal-Bench 2.0 Terminal-Bench 2.0 | 88.0 | Terminal-based software tasks | Professional software engineering | Terminal-Bench 2.0 |
| SWE-bench Verified Software Engineering Benchmark Verified | 95.5 | 500 verified issues | Professional software engineering | SWE-bench: Can Language Models Resolve Real-World GitHub Issues? |
| SWE-bench Pro SWE-bench Pro | 80.3 | 1,865 repository problems | Long-horizon professional engineering | SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? |
Mathematics
| Benchmark | Value | Tasks | Difficulty | Provenance |
|---|---|---|---|---|
| USAMO 2026 United States of America Mathematical Olympiad 2026 | 97.6 | 6 proof-based problems | International olympiad level | United States of America Mathematical Olympiad |
Multilingual
| Benchmark | Value | Tasks | Difficulty | Provenance |
|---|---|---|---|---|
| SWE Multilingual SWE-bench Multilingual | 92.2 | 300 problems across 9 languages | Professional multilingual software engineering | SWE-bench Multilingual |
Agentic
| Benchmark | Value | Tasks | Difficulty | Provenance |
|---|---|---|---|---|
| Terminal-Bench 2.0 Terminal-Bench 2.0 | 88 | Terminal-based software tasks | Professional software engineering | Terminal-Bench 2.0 |
| BrowseComp BrowseComp | 88 | Research questions requiring browsing | Hard web research | BrowseComp |
| OSWorld-Verified OSWorld-Verified | 85 | 369 real-world computer tasks (361 when eight Google Drive tasks are excluded) | Multi-step desktop and cross-application workflows | OSWorld |
| CyberGym CyberGym | 83.8 | 1,507 vulnerability analysis instances | Real-world cybersecurity | CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale |
Multimodal & Grounded
| Benchmark | Value | Tasks | Difficulty | Provenance |
|---|---|---|---|---|
| CharXiv CharXiv Reasoning | 93.5 | Scientific chart reasoning | Scientific visualization reasoning | CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs |
| CharXiv w/o tools CharXiv Reasoning without tools | 88.9 | Scientific chart reasoning (tool-free) | Scientific visualization reasoning | CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs |
| SWE-bench Multimodal SWE-bench Multimodal | 54.9 | Multimodal software engineering tasks | Frontier multimodal coding | SWE-bench Multimodal |
external
| Benchmark | Value | Tasks | Difficulty | Provenance |
|---|---|---|---|---|
| ExploitBench ExploitBench v8-bench | 78 | V8 exploit synthesis runs | Browser exploitation and cybersecurity | ExploitBench |
| The Last Ones completion The Last Ones Cyber Range Completion Rate | 60 | 10 long-horizon runs | Long-horizon cyber operations | UK AISI and CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities |
| Firefox 147 exploits Firefox 147 Working Exploit Rate | 88.4 | 250 Firefox 147 vulnerability trials | Browser exploit development | Claude Fable 5 and Claude Mythos 5 |
| Anthropic OSS-Fuzz crash Anthropic OSS-Fuzz Any-Crash Rate | 80.0 | Approximately 830 OSS-Fuzz entry points | Vulnerability discovery | Claude Fable 5 and Claude Mythos 5 |
| Anthropic OSS-Fuzz write primitive Anthropic OSS-Fuzz Write-Primitive-or-Higher Rate | 32.4 | Approximately 830 OSS-Fuzz entry points | Vulnerability exploitation | Claude Fable 5 and Claude Mythos 5 |
Lifecycle and limitations log
Lifecycle events the source associates with this model, plus what this profile does not claim.