# o3-mini

OpenAI · Proprietary · rank 141 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/o3-mini  
JSON: https://modelscale.dev/api/model/o3-mini

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `o3-mini` | — |
| Overall score | 47.03 | Observed 2026-09-22 · source benchlm:models |
| Context window | 200K tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2025-01-31 | Observed 2026-09-22 · source benchlm:models |
| Access type | Proprietary | — |
| Blended $/1M (75% input / 25% output) | $1.93 | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | Unavailable | Unavailable · source benchlm:models |
| Coding | 23 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 66.3 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | Unavailable | Unavailable · source benchlm:models |
| Multimodal & Grounded | Unavailable | Unavailable · source benchlm:models |
| Instruction Following | Unavailable | Unavailable · source benchlm:models |
| Math | Unavailable | Unavailable · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 7.12 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 160 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $1.10 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $4.40 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | $0.55 | Observed 2026-09-22 · source openrouter:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $1.93 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** OpenAI's o3-mini model page lists $1.10 input / $4.40 output per million tokens. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $5.42 |
| Modelled tokens | 2.82M |

## Benchmark record

9 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MMLU (Massive Multitask Language Understanding) | 86.9 | 57 subjects | Elementary to professional level | [Measuring Massive Multitask Language Understanding](https://arxiv.org/abs/2009.03300) |
| GPQA (Graduate-Level Google-Proof Q&A) | 77.2 | 448 questions | Graduate level | [GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022) |
| Artificial Analysis Intelligence Index | 12.5 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 74.8 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 7.9 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| SWE-bench Verified (Software Engineering Benchmark Verified) | 49.3 | 500 verified issues | Professional software engineering | [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770) |

### Mathematics

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AIME 2024 (American Invitational Mathematics Examination 2024) | 87.3 | 15 problems | High school olympiad level | [American Invitational Mathematics Examination](https://www.maa.org/math-competitions/aime) |

### Instruction Following

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| IFEval (Instruction-Following Eval) | 93.9 | 541 prompts across 25 instruction types | Instruction precision | [Instruction-Following Evaluation for Large Language Models](https://arxiv.org/abs/2311.07911) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| τ²-bench results (τ²-Bench Tool-Agent-User Evaluation) | 28.7 | Airline, retail, and telecom customer-service task sets | Dual-control customer-service workflows | [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment](https://arxiv.org/abs/2506.07982) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
