# GPT-6 Astra

OpenAI · Proprietary · rank 2 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/gpt-6-astra  
JSON: https://modelscale.dev/api/model/gpt-6-astra

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `gpt-6-astra` | — |
| Overall score | 83.79 | Observed 2026-09-22 · source benchlm:models |
| Context window | 1.05M tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-09-03 | Observed 2026-09-22 · source benchlm:models |
| Access type | Proprietary | — |
| Blended $/1M (75% input / 25% output) | $20.00 | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 87.1 | Observed 2026-09-22 · source benchlm:models |
| Coding | 75.2 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 86.4 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 89.5 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 82.9 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | Unavailable | Unavailable · source benchlm:models |
| Math | 85 | Observed 2026-09-22 · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 253.04 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 69 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $10.00 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $50.00 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | $1.00 | Observed 2026-09-22 · source benchlm:pricing |
| Cache write / 1M tokens | $12.50 | Observed 2026-09-22 · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $20.00 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** OpenAI's GPT-6 Astra model page lists $10.00 input / $1.00 cached input / $12.50 cache writes / $50.00 output per million tokens, a 1,050,000-token context window, 922,000 maximum input tokens, and 128,000 maximum output tokens. Prompts above 272K input tokens are charged at 2x input and cache rates and 1.5x output for the full request; Batch and Flex are 50% of Standard and Fast mode is 2x. The September 3, 2026 price sheet lists Astra at 2.5x GPT-5.6 Sol's current $4.00 / $20.00 rates. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $56.32 |
| Modelled tokens | 2.82M |

## Benchmark record

59 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GPQA (Graduate-Level Google-Proof Q&A) | 96 | 448 questions | Graduate level | [GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022) |
| GPQA-D (GPQA Diamond) | 96.0 | Graduate-level science questions | Graduate level | [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking) |
| Artificial Analysis Intelligence Index | 52.7 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 96.1 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 54.7 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | 43.4 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 62.6 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 51.3 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| HealthBench Hard | 36.3 | 1,000 health prompts | Advanced health reasoning | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| HealthBench Professional | 63.4 | Clinician chat tasks | Professional clinical workflows | [HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats](https://arxiv.org/abs/2604.27470) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| FrontierCode 1.1 Main | 53.3 | 100 private Main tasks (150 in Extended) | Frontier coding-agent quality | [FrontierCode leaderboard](https://cognition.com/frontiercode) |
| FrontierCode 1.1 Extended | 64.5 | 150 private software-engineering tasks | Frontier coding-agent quality | [GPT-5.6 models are now available in Devin](https://devin.ai/blog/gpt-5-6) |
| FrontierSWE v2 | 65.5 | 34 ultra-long-horizon engineering and research tasks | Ultra-long-horizon frontier software engineering | [FrontierSWE v2](https://www.frontierswe.com/blog/v2) |
| AA Coding Index (Artificial Analysis Coding Index) | 76.9 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 56.5 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| DeepSWE | 74.1 | 113 software engineering tasks across 91 repositories and 5 languages | Long-horizon software engineering | [DeepSWE benchmark blog](https://deepswe.datacurve.ai/blog) |

### Mathematics

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| FrontierMath v2 (Tier 4) (FrontierMath v2 Tier 4) | 97.600 | 43 private extreme-difficulty mathematics problems | Research-level mathematics requiring hours or days of expert work | [FrontierMath Tier 4 v2 leaderboard](https://epoch.ai/benchmarks/frontiermath-tier-4-v2?view=graph&tab=leaderboard) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| ARC-AGI-1 (ARC-AGI-1 Semi-Private Evaluation) | 98.50 | Semi-private ARC-AGI-1 evaluation set | Abstract visual reasoning | [ARC Prize leaderboard](https://arcprize.org/) |
| MRCR v2 256K-512K (OpenAI MRCR v2 8-needle 256K-512K) | 100.0 | 100 eight-needle retrieval examples | Very long-context retrieval | [Muse Spark 1.3 Evaluation Methodology](https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology) |
| MRCR v2 512K-1M (OpenAI MRCR v2 8-needle 512K-1M) | 96.3 | 100 eight-needle retrieval examples | Million-token retrieval | [Muse Spark 1.3 Evaluation Methodology](https://research.meta.ai/static/muse-spark-1-3-multimodal-evaluation-methodology) |
| ARC-AGI-2 (Abstraction and Reasoning Corpus for AGI v2) | 95 | Visual pattern completion and abstract reasoning | Expert-level — hardest public reasoning benchmark | [ARC-AGI-2: A Harder General Intelligence Benchmark](https://arcprize.org/arc-agi/2/) |
| ARC-AGI-3 (Abstraction and Reasoning Corpus for AGI v3) | 62.71 | Interactive game-like tasks with hidden rules | Frontier agentic reasoning | [ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence](https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf) |
| GeneBench-Pro | 37.8 | 129 genomics statistical-analysis workflows | Long-horizon scientific reasoning | [GeneBench-Pro: Evaluating Multistage Statistical Reasoning](https://cdn.openai.com/pdf/21938268-21af-442f-af93-3b2249afb241/genebench-pro.pdf) |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 80.7 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| MLCR-AA (Medical Long Context Reasoning (MLCR-AA)) | 35.0 | Long, fragmented medical-record reasoning | Long-context medical reasoning | [Medical Long Context Reasoning (MLCR-AA)](https://artificialanalysis.ai/evaluations/mlcr-aa) |
| CritPt (Critical Physics Tasks) | 31.7 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 4.0 | 57.90 | 66 professional computer-work tasks | Frontier autonomous knowledge work | [Terminal-Bench 4.0](https://www.tbench.ai/news/terminal-bench-4-0) |
| Terminal-Bench-Science 0.1 | 64.6 | 70 expert-curated scientific research workflows | Frontier scientific research workflows | [Terminal-Bench-Science 0.1](https://www.terminal-bench-science.ai/) |
| AA Briefcase (Artificial Analysis Briefcase) | 1569 | Professional knowledge-work tasks | Professional work | [Artificial Analysis Briefcase Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-briefcase) |
| AA AutomationBench (Artificial Analysis AutomationBench) | 68.5 | Business-process automation tasks | Agentic automation | [Artificial Analysis AutomationBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/automationbench-aa) |
| AA Tau3 Banking (Artificial Analysis Tau3-Banking) | 41.4 | Banking tool-use workflows | Agentic banking workflows | [Artificial Analysis Tau3-Banking Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/tau3-banking) |
| BrowseComp | 91.5 | Research questions requiring browsing | Hard web research | [BrowseComp](https://openai.com/index/browsecomp/) |
| HLE w/ tools (Humanity's Last Exam with tools) | 57.2 | Expert questions with tool use | Frontier tool-augmented reasoning | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA | 1542 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 52.1 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA Agentic Index (Artificial Analysis Agentic Index) | 51.5 | Cross-benchmark agentic index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA Terminal-Bench 4.0 (Artificial Analysis Terminal-Bench v4.0) | 59.1 | Terminal-based agent tasks | Agentic software engineering | [Artificial Analysis Terminal-Bench v4.0 Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/terminalbench-v4-0) |
| GDP.pdf (Artificial Analysis GDP.pdf) | 31.0 | Professional document-production tasks | Professional knowledge work | [GDP.pdf Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gdp-pdf) |
| AA-AnalystAgent (Artificial Analysis AnalystAgent) | 51.2 | Spreadsheet and document analysis questions | Business and data analysis | [AA-AnalystAgent Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-analyst-agent) |
| OSWorld 2.0 | 72.6 | 108 long-horizon computer-use workflows | Long-horizon professional workflows | [OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks](https://arxiv.org/abs/2606.29537) |
| ExploitGym | 42.4 | 898 exploitation tasks | Advanced cybersecurity exploitation | [ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?](https://arxiv.org/abs/2605.11086) |
| AutomationBench | 41.4 | 600 public automation tasks | Long-horizon automation | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| Agents' Last Exam | 59.3 | Agent tasks | Advanced agentic work | [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 87.3 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |
| ApprenticeBench (ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job) | 68 | 100 vendor bills processed in sequence inside a simulated construction company | Long-horizon computer use with offline and online continual learning | [ApprenticeBench: a step change in AI's job readiness](https://neocognition.io/blog/apprentice-bench/) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| BenchCAD Vision2Code (tools) (BenchCAD Vision2Code voxel IoU with tools) | 0.959 | 1,000-file Vision2Code subset | Programmatic CAD generation | [BenchCAD: A comprehensive, industry-standard benchmark for programmatic CAD](https://arxiv.org/abs/2605.10865) |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 86.9 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| ScreenSpot Pro | 92.7 | 1,581 grounding instructions | Professional GUI grounding | [ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use](https://arxiv.org/abs/2504.07981) |

### external

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| SRE-Bench (SRE-Bench binary reverse engineering) | 88.0 | 262 binary reverse-engineering instances | Binary reverse engineering | [GPT-6 Astra System Card](https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra.pdf) |
| ExploitBench (ExploitBench v8-bench) | 100 | V8 exploit synthesis runs | Browser exploitation and cybersecurity | [ExploitBench](https://exploitbench.ai/) |
| ACCR standard (Advanced Cyber Completion Rate — Standard Access) | 3.5 | Internal advanced-cyber request set | Cyber access and safeguards | [Expanding Daybreak as the cyber defense window narrows](https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/) |
| ACCR Daybreak Blue (Advanced Cyber Completion Rate — Daybreak Blue) | 3.5 | Internal advanced-cyber request set | Cyber access and safeguards | [Expanding Daybreak as the cyber defense window narrows](https://openai.com/index/expanding-daybreak-as-the-cyber-defense-window-narrows/) |
| SEC-Bench Pro | 85.4 | Security engineering tasks | Advanced cybersecurity | [Introducing GPT-5.6](https://openai.com/index/gpt-5-6/) |
| FrontierCyber | 38.1 | 197 dynamic cyber tasks | Easy through elite cyber operations | [FrontierCyber](https://www.irregular.com/research/frontiercyber) |
| CyScenarioBench success (CyScenarioBench Average Success Rate) | 59 | 11 cyber scenarios | Long-horizon cybersecurity | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |
| CyScenarioBench solved (CyScenarioBench Scenarios Ever Solved) | 9 | 11 cyber scenarios | Long-horizon cybersecurity | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |
| Atomic network attacks (Atomic Network Attack Simulation) | 100 | Atomic cyber tasks | Network attack simulation | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |
| Atomic vulnerability research (Atomic Vulnerability Research and Exploitation) | 100 | Atomic cyber tasks | Vulnerability research and exploitation | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |
| Atomic evasion (Atomic Evasion) | 52 | Atomic cyber tasks | Cybersecurity evasion | [Assessing GPT-5.6 Sol](https://www.irregular.com/research/assessing-gpt-5.6-sol) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

| Event | Type | Confirmation | Announced | Effective | Replacement | Source |
| --- | --- | --- | --- | --- | --- | --- |
| GPT-6 Astra | Release | confirmed | Unavailable | 2026-09-03 | — | [OpenAI GPT-6 Astra model documentation](https://developers.openai.com/api/docs/models/gpt-6-astra) |

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
