# GPT-5.4 Pro

OpenAI · Proprietary · rank 57 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/gpt-5-4-pro  
JSON: https://modelscale.dev/api/model/gpt-5-4-pro

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `gpt-5-4-pro` | — |
| Overall score | 60.8 | Observed 2026-09-22 · source benchlm:models |
| Context window | 1.05M tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-03-05 | Observed 2026-09-22 · source benchlm:models |
| Access type | Proprietary | — |
| Blended $/1M (75% input / 25% output) | $67.50 | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 79.3 | Observed 2026-09-22 · source benchlm:models |
| Coding | Unavailable | Unavailable · source benchlm:models |
| Knowledge | 84.3 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 70.2 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | Unavailable | Unavailable · source benchlm:models |
| Instruction Following | Unavailable | Unavailable · source benchlm:models |
| Math | 68.8 | Observed 2026-09-22 · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 151.79 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 74 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $30.00 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $180.00 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $67.50 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** OpenAI's current API pricing page lists GPT-5.4 Pro at $30.00 input / $180.00 output per million tokens for short-context requests, with higher pricing for long-context requests and no cached-input rate. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $190.08 |
| Modelled tokens | 2.82M |

## Benchmark record

11 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| HLE (Humanity's Last Exam) | 58.7 | Expert-level questions | Frontier expert level | [Humanity's Last Exam](https://lastexam.ai/) |
| FrontierScience | 36.7 | Research-level science tasks | Research frontier | [FrontierScience](https://openai.com/index/frontierscience/) |
| FrontierScience Research | 36.7 | Scientific research problems | Frontier scientific research | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| HLE w/o tools (Humanity's Last Exam without tools) | 42.7 | Expert-level questions | Frontier expert level | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |

### Mathematics

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| IPhO 2025 (Theory) (International Physics Olympiad 2025 (Theory)) | 93.5 | 3 olympiad theory problems | International olympiad physics | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| FrontierMath (legacy) (FrontierMath legacy aggregate) | 50 | Historical aggregate | Research-level mathematics | [FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI](https://epoch.ai/frontiermath) |
| FrontierMath v2 (Tiers 1-3) (FrontierMath v2 Tiers 1-3) | 50.000 | 295 private advanced mathematics problems | From olympiad-plus to early research mathematics | [FrontierMath v2 benchmark hub](https://epoch.ai/benchmarks/frontiermath-tier-4-v2) |
| FrontierMath v2 (Tier 4) (FrontierMath v2 Tier 4) | 37.500 | 43 private extreme-difficulty mathematics problems | Research-level mathematics requiring hours or days of expert work | [FrontierMath Tier 4 v2 leaderboard](https://epoch.ai/benchmarks/frontiermath-tier-4-v2?view=graph&tab=leaderboard) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| ARC-AGI-2 (Abstraction and Reasoning Corpus for AGI v2) | 83.3 | Visual pattern completion and abstract reasoning | Expert-level — hardest public reasoning benchmark | [ARC-AGI-2: A Harder General Intelligence Benchmark](https://arcprize.org/arc-agi/2/) |
| CritPt (Critical Physics Tasks) | 30.0 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| BrowseComp | 89.3 | Research questions requiring browsing | Hard web research | [BrowseComp](https://openai.com/index/browsecomp/) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
