# GPT-5.6 Sol

OpenAI · Proprietary · rank 5 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/gpt-5-6-sol  
JSON: https://modelscale.dev/api/model/gpt-5-6-sol

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `gpt-5-6-sol` | — |
| Overall score | 80.65 | Observed 2026-09-22 · source benchlm:models |
| Context window | 1.05M tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-07-09 | Observed 2026-09-22 · source benchlm:models |
| Access type | Proprietary | — |
| Blended $/1M (75% input / 25% output) | $8.00 | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 93.6 | Observed 2026-09-22 · source benchlm:models |
| Coding | 62.6 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 83.5 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 70.4 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 87.6 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | 87.7 | Observed 2026-09-22 · source benchlm:models |
| Math | 96.8 | Observed 2026-09-22 · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 122.52 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 77 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $4.00 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $20.00 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | $0.40 | Observed 2026-09-22 · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $8.00 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** OpenAI's API pricing page lists GPT-5.6 Sol at $4.00 input / $0.40 cached input / $20.00 output per million tokens for short-context requests, and states this promotional pricing is available at least through November 21, 2026. Prompts above 272K input tokens use the published $8.00 input / $0.80 cached input / $30.00 output long-context tier; cache writes cost 1.25x uncached input. Observed September 11, 2026; the earlier $5.00 / $30.00 list rate is retained in the price history. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $22.53 |
| Modelled tokens | 2.82M |

## Benchmark record

72 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GPQA (Graduate-Level Google-Proof Q&A) | 94.6 | 448 questions | Graduate level | [GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022) |
| GPQA-D (GPQA Diamond) | 94.6 | Graduate-level science questions | Graduate level | [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking) |
| HLE-Verified | 54.5 | 1,811 verified or revised expert questions | Frontier multidisciplinary expert reasoning | [HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam](https://arxiv.org/abs/2602.13964) |
| LABBench2 (LABBench2: An Improved Benchmark for AI Systems Performing Biology Research) | 82.1 | Nearly 1,900 biology-research tasks | Real-world biology research | [LABBench2: An Improved Benchmark for AI Systems Performing Biology Research](https://arxiv.org/abs/2604.09554) |
| Artificial Analysis Intelligence Index | 58.9 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 94.1 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 49.5 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | 22.0 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 59.4 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 92.2 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| HealthBench Hard | 33.1 | 1,000 health prompts | Advanced health reasoning | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| HealthBench Professional | 60.5 | Clinician chat tasks | Professional clinical workflows | [HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats](https://arxiv.org/abs/2604.27470) |
| GPQA Diamond (Vals) (GPQA Diamond, Vals AI run) | 95.2 | Graduate-level science questions | Expert reasoning | [Vals AI GPQA Diamond, Vals AI run leaderboard](https://www.vals.ai/benchmarks/gpqa) |
| MMLU-Pro (Vals) (MMLU-Pro, Vals AI run) | 89.1 | Academic multiple-choice questions | Broad academic knowledge | [Vals AI MMLU-Pro, Vals AI run leaderboard](https://www.vals.ai/benchmarks/mmlu_pro) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 2.0 | 91.9 | Terminal-based software tasks | Professional software engineering | [Terminal-Bench 2.0](https://www.tbench.ai/) |
| SWE-bench Pro | 64.6 | 1,865 repository problems | Long-horizon professional engineering | [SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?](https://arxiv.org/abs/2509.16941) |
| VulcanBench v3 | 87.0 | 23 post-cutoff repository tasks in the v3 report | Professional multi-file software engineering | [VulcanBench](https://github.com/morganlinton/VulcanBench/tree/main) |
| VulcanBench CII v1 (VulcanBench Coding Intelligence Index v1) | 86.5 | 38 validated post-cutoff repository tasks | Mid-band frontier software engineering | [CII v1 frontier results](https://github.com/morganlinton/VulcanBench/blob/main/docs/results/cii-v1-2026-08/README.md) |
| FrontierCode 1.1 Extended | 60.6 | 150 private software-engineering tasks | Frontier coding-agent quality | [GPT-5.6 models are now available in Devin](https://devin.ai/blog/gpt-5-6) |
| FrontierSWE v2 | 32.2 | 34 ultra-long-horizon engineering and research tasks | Ultra-long-horizon frontier software engineering | [FrontierSWE v2](https://www.frontierswe.com/blog/v2) |
| Bug Hunt Bench | 42 | 105 planted bugs across two production TypeScript repositories | Blind production-repository bug finding and repair | [Bug Hunt Bench method, caveats, and definitions](https://bughunt.productcompass.pm/method) |
| AA Coding Index (Artificial Analysis Coding Index) | 77.4 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 57.1 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| LiveCodeBench (Vals) (LiveCodeBench, Vals AI run) | 82.6 | Competitive programming problems (easy, medium, hard) | Frontier coding | [Vals AI LiveCodeBench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/lcb) |
| SWE-bench (Vals) (SWE-bench, Vals AI run) | 96.2 | Real repository issues by human time bucket | Frontier coding agents | [Vals AI SWE-bench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/swebench) |
| DeepSWE | 72.7 | 113 software engineering tasks across 91 repositories and 5 languages | Long-horizon software engineering | [DeepSWE benchmark blog](https://deepswe.datacurve.ai/blog) |

### Mathematics

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| FrontierMath (legacy) (FrontierMath legacy aggregate) | 89 | Historical aggregate | Research-level mathematics | [FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI](https://epoch.ai/frontiermath) |
| FrontierMath v2 (Tiers 1-3) (FrontierMath v2 Tiers 1-3) | 89.000 | 295 private advanced mathematics problems | From olympiad-plus to early research mathematics | [FrontierMath v2 benchmark hub](https://epoch.ai/benchmarks/frontiermath-tier-4-v2) |
| FrontierMath v2 (Tier 4) (FrontierMath v2 Tier 4) | 83.000 | 43 private extreme-difficulty mathematics problems | Research-level mathematics requiring hours or days of expert work | [FrontierMath Tier 4 v2 leaderboard](https://epoch.ai/benchmarks/frontiermath-tier-4-v2?view=graph&tab=leaderboard) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| ARC-AGI-2 (Abstraction and Reasoning Corpus for AGI v2) | 92.5 | Visual pattern completion and abstract reasoning | Expert-level — hardest public reasoning benchmark | [ARC-AGI-2: A Harder General Intelligence Benchmark](https://arcprize.org/arc-agi/2/) |
| ARC-AGI-3 (Abstraction and Reasoning Corpus for AGI v3) | 7.78 | Interactive game-like tasks with hidden rules | Frontier agentic reasoning | [ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence](https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf) |
| GeneBench-Pro | 28.7 | 129 genomics statistical-analysis workflows | Long-horizon scientific reasoning | [GeneBench-Pro: Evaluating Multistage Statistical Reasoning](https://cdn.openai.com/pdf/21938268-21af-442f-af93-3b2249afb241/genebench-pro.pdf) |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 84.0 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| MLCR-AA (Medical Long Context Reasoning (MLCR-AA)) | 26.1 | Long, fragmented medical-record reasoning | Long-context medical reasoning | [Medical Long Context Reasoning (MLCR-AA)](https://artificialanalysis.ai/evaluations/mlcr-aa) |
| CritPt (Critical Physics Tasks) | 32.3 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Instruction Following

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA-IFBench (Artificial Analysis IFBench) | 72.7 | Verifiable instruction constraints | Instruction precision | [Artificial Analysis IFBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/ifbench) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 3.0 | 34.6 | 74 professional computer-work tasks across 7 domains | Frontier autonomous knowledge work | [Terminal-Bench 3.0](https://www.frontierbench.ai/) |
| AA Briefcase (Artificial Analysis Briefcase) | 1487 | Professional knowledge-work tasks | Professional work | [Artificial Analysis Briefcase Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-briefcase) |
| AA AutomationBench (Artificial Analysis AutomationBench) | 60.1 | Business-process automation tasks | Agentic automation | [Artificial Analysis AutomationBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/automationbench-aa) |
| AA EnterpriseOps-Gym (Artificial Analysis EnterpriseOps-Gym) | 42.9 | Enterprise operations workflows | Enterprise agent operations | [Artificial Analysis EnterpriseOps-Gym Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa) |
| AA Harvey LAB (Artificial Analysis Harvey LAB-AA) | 87.2 | Legal agent tasks | Professional legal work | [Artificial Analysis Harvey LAB-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/harvey-lab-aa) |
| AA ITBench (Artificial Analysis ITBench-AA) | 56.2 | IT incident-response tasks | Enterprise IT operations | [Artificial Analysis ITBench-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/itbench-aa) |
| AA Tau3 Banking (Artificial Analysis Tau3-Banking) | 44.3 | Banking tool-use workflows | Agentic banking workflows | [Artificial Analysis Tau3-Banking Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/tau3-banking) |
| Terminal-Bench 2.0 | 91.9 | Terminal-based software tasks | Professional software engineering | [Terminal-Bench 2.0](https://www.tbench.ai/) |
| BrowseComp | 92.2 | Research questions requiring browsing | Hard web research | [BrowseComp](https://openai.com/index/browsecomp/) |
| GDPval-AA | 1735 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 54.4 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA Agentic Index (Artificial Analysis Agentic Index) | 50.5 | Cross-benchmark agentic index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA Terminal-Bench 4.0 (Artificial Analysis Terminal-Bench v4.0) | 39.9 | Terminal-based agent tasks | Agentic software engineering | [Artificial Analysis Terminal-Bench v4.0 Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/terminalbench-v4-0) |
| GDP.pdf (Artificial Analysis GDP.pdf) | 27.2 | Professional document-production tasks | Professional knowledge work | [GDP.pdf Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gdp-pdf) |
| AA-AnalystAgent (Artificial Analysis AnalystAgent) | 47.5 | Spreadsheet and document analysis questions | Business and data analysis | [AA-AnalystAgent Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-analyst-agent) |
| OSWorld 2.0 | 62.6 | 108 long-horizon computer-use workflows | Long-horizon professional workflows | [OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks](https://arxiv.org/abs/2606.29537) |
| CyberGym | 84.5 | 1,507 vulnerability analysis instances | Real-world cybersecurity | [CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale](https://www.cybergym.io/) |
| ExploitGym | 33.7 | 898 exploitation tasks | Advanced cybersecurity exploitation | [ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?](https://arxiv.org/abs/2605.11086) |
| Toolathlon | 58 | Multi-tool workflows | Advanced tool use | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| τ²-bench results (τ²-Bench Tool-Agent-User Evaluation) | 85.1 | Airline, retail, and telecom customer-service task sets | Dual-control customer-service workflows | [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment](https://arxiv.org/abs/2506.07982) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 85.8 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |
| ApprenticeBench (ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job) | 26 | 100 vendor bills processed in sequence inside a simulated construction company | Long-horizon computer use with offline and online continual learning | [ApprenticeBench: a step change in AI's job readiness](https://neocognition.io/blog/apprentice-bench/) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MMMU-Pro (Massive Multi-discipline Multimodal Understanding Pro) | 83 | Multimodal academic reasoning | Frontier multimodal | [MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark](https://arxiv.org/abs/2409.02813) |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 83.4 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| MMMU-Pro w/ Python (MMMU-Pro with Python) | 84.6 | Multimodal academic reasoning | Frontier multimodal | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |

### external

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| ExploitBench (ExploitBench v8-bench) | 74 | V8 exploit synthesis runs | Browser exploitation and cybersecurity | [ExploitBench](https://exploitbench.ai/) |
| The Last Ones completion (The Last Ones Cyber Range Completion Rate) | 70 | 10 long-horizon runs | Long-horizon cyber operations | [UK AISI and CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities) |
| ACCR standard (Advanced Cyber Completion Rate — Standard Access) | 1.5 | Internal advanced-cyber request set | Cyber access and safeguards | [Expanding Daybreak as the cyber defense window narrows](https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/) |
| ACCR Daybreak Blue (Advanced Cyber Completion Rate — Daybreak Blue) | 2.0 | Internal advanced-cyber request set | Cyber access and safeguards | [Expanding Daybreak as the cyber defense window narrows](https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/) |
| SEC-Bench Pro | 71.2 | Security engineering tasks | Advanced cybersecurity | [Introducing GPT-5.6](https://openai.com/index/gpt-5-6/) |
| FrontierCyber | 9.6 | 197 dynamic cyber tasks | Easy through elite cyber operations | [FrontierCyber](https://www.irregular.com/research/frontiercyber) |
| CyScenarioBench success (CyScenarioBench Average Success Rate) | 28 | 11 cyber scenarios | Long-horizon cybersecurity | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |
| CyScenarioBench solved (CyScenarioBench Scenarios Ever Solved) | 7 | 11 cyber scenarios | Long-horizon cybersecurity | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |
| Atomic network attacks (Atomic Network Attack Simulation) | 98 | Atomic cyber tasks | Network attack simulation | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |
| Atomic vulnerability research (Atomic Vulnerability Research and Exploitation) | 91 | Atomic cyber tasks | Vulnerability research and exploitation | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |
| Atomic evasion (Atomic Evasion) | 56 | Atomic cyber tasks | Cybersecurity evasion | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
