# Gemini 3.1 Pro

Google · Proprietary · rank 14 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/gemini-3-1-pro  
JSON: https://modelscale.dev/api/model/gemini-3-1-pro

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `gemini-3-1-pro` | — |
| Overall score | 70.54 | Observed 2026-09-22 · source benchlm:models |
| Context window | 1M tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-02-19 | Observed 2026-09-22 · source benchlm:models |
| Access type | Proprietary | — |
| Blended $/1M (75% input / 25% output) | $4.50 | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 72 | Observed 2026-09-22 · source benchlm:models |
| Coding | 76.4 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 78.9 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 50.7 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 79.1 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | Unavailable | Unavailable · source benchlm:models |
| Math | 54.2 | Observed 2026-09-22 · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 63.62 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 124 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $2.00 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $12.00 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | $0.20 | Observed 2026-09-22 · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $4.50 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** Google's Gemini Developer API pricing page lists gemini-3.1-pro-preview at $2.00 input / $0.20 cached input / $12.00 output per million tokens for prompts up to 200K tokens, rising to $4.00 / $0.40 / $18.00 above 200K. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $12.67 |
| Modelled tokens | 2.82M |

## Benchmark record

46 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GPQA-D (GPQA Diamond) | 94.3 | Graduate-level science questions | Graduate level | [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking) |
| Artificial Analysis Intelligence Index | 29.7 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 94.1 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 47.0 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | 31.9 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 54.9 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 50.9 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| HealthBench Hard | 20.6 | 1,000 health prompts | Advanced health reasoning | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| MedXpertQA (Text) (MedXpertQA Text) | 71.5 | 2,450 medical multiple-choice questions | Professional medical knowledge | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| HLE w/o tools (Humanity's Last Exam without tools) | 45.4 | Expert-level questions | Frontier expert level | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| GPQA Diamond (Vals) (GPQA Diamond, Vals AI run) | 95.5 | Graduate-level science questions | Expert reasoning | [Vals AI GPQA Diamond, Vals AI run leaderboard](https://www.vals.ai/benchmarks/gpqa) |
| MMLU-Pro (Vals) (MMLU-Pro, Vals AI run) | 91.0 | Academic multiple-choice questions | Broad academic knowledge | [Vals AI MMLU-Pro, Vals AI run leaderboard](https://www.vals.ai/benchmarks/mmlu_pro) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| LiveCodeBench Pro | 82.9 | Quarter-specific contest programming sets | High-end contest programming | [LiveCodeBench Pro: How Do Olympiad Medalists Judge LLMs in Competitive Programming?](https://arxiv.org/abs/2506.11928) |
| Vibe Code Bench (Vibe Code Bench v1.1) | 32.03 | End-to-end web application builds | End-to-end software delivery | [Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development](https://www.vals.ai/benchmarks/vibe-code) |
| React Native Evals | 78.9 | React Native app implementation tasks | Production mobile app engineering | [React Native Evals](https://rn-evals.vercel.app/) |
| AA Coding Index (Artificial Analysis Coding Index) | 68.8 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 58.7 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| LiveCodeBench (Vals) (LiveCodeBench, Vals AI run) | 88.5 | Competitive programming problems (easy, medium, hard) | Frontier coding | [Vals AI LiveCodeBench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/lcb) |
| SWE-bench (Vals) (SWE-bench, Vals AI run) | 78.8 | Real repository issues by human time bucket | Frontier coding agents | [Vals AI SWE-bench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/swebench) |

### Mathematics

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| FrontierMath v2 (Tiers 1-3) (FrontierMath v2 Tiers 1-3) | 36.900 | 295 private advanced mathematics problems | From olympiad-plus to early research mathematics | [FrontierMath v2 benchmark hub](https://epoch.ai/benchmarks/frontiermath-tier-4-v2) |
| FrontierMath v2 (Tier 4) (FrontierMath v2 Tier 4) | 16.700 | 43 private extreme-difficulty mathematics problems | Research-level mathematics requiring hours or days of expert work | [FrontierMath Tier 4 v2 leaderboard](https://epoch.ai/benchmarks/frontiermath-tier-4-v2?view=graph&tab=leaderboard) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| ARC-AGI-2 (Abstraction and Reasoning Corpus for AGI v2) | 77.08 | Visual pattern completion and abstract reasoning | Expert-level — hardest public reasoning benchmark | [ARC-AGI-2: A Harder General Intelligence Benchmark](https://arcprize.org/arc-agi/2/) |
| ARC-AGI-3 (Abstraction and Reasoning Corpus for AGI v3) | 0.42 | Interactive game-like tasks with hidden rules | Frontier agentic reasoning | [ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence](https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf) |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 82.0 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| CritPt (Critical Physics Tasks) | 17.7 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Instruction Following

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA-IFBench (Artificial Analysis IFBench) | 77.1 | Verifiable instruction constraints | Instruction precision | [Artificial Analysis IFBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/ifbench) |

### Multilingual

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA Global-MMLU-Lite (Artificial Analysis Global-MMLU-Lite) | 93.2 | Multilingual knowledge questions | Multilingual professional knowledge | [Artificial Analysis Global-MMLU-Lite Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/global-mmlu-lite) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GDPval-AA | 904 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 13.8 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA Agentic Index (Artificial Analysis Agentic Index) | 10.3 | Cross-benchmark agentic index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| APEX-Agents-AA | 32.0 | 452 professional-services agent tasks | Long-horizon workplace agent tasks | [APEX-Agents-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/apex-agents-aa) |
| Gert Labs (Gert Labs Composite Game Benchmark) | 56.87 | Novel game environments | Agentic coding and decision-making | [Gert Labs rankings](https://gertlabs.com/rankings) |
| τ²-bench results (τ²-Bench Tool-Agent-User Evaluation) | 95.6 | Airline, retail, and telecom customer-service task sets | Dual-control customer-service workflows | [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment](https://arxiv.org/abs/2506.07982) |
| DeepSearchQA | 69.7 | Agentic browsing and list-answer questions | Agentic web research | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| Claw-Eval | 57.8 | 300 tasks, 2,159 rubrics | Real-world general, multi-turn, and native multimodal agent execution | [Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents](https://arxiv.org/abs/2604.06132) |
| ResearchClawBench | 13.3 | 40 tasks across 10 scientific domains | Scientific research re-discovery | [ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research](https://arxiv.org/abs/2606.07591) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 70.8 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MMMU-Pro (Massive Multi-discipline Multimodal Understanding Pro) | 83.9 | Multimodal academic reasoning | Frontier multimodal | [MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark](https://arxiv.org/abs/2409.02813) |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 82.4 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| Design Arena Website (Design Arena Website Elo) | 1264 | Website generation comparisons | Design and website generation | [OpenRouter Grok 4.3 benchmarks](https://openrouter.ai/x-ai/grok-4.3/benchmarks) |
| ERQA | 69.4 | Evidence-based visual QA | Grounded multimodal reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| ScreenSpot Pro | 84.4 | 1,581 grounding instructions | Professional GUI grounding | [ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use](https://arxiv.org/abs/2504.07981) |
| MedXpertQA (MM) (MedXpertQA Multimodal) | 81.3 | 2,000 multimodal medical questions | Clinical multimodal reasoning | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| ZeroBench | 29.0 | 100 visual reasoning questions | Tool-augmented visual reasoning | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| SimpleVQA | 72.4 | Visual QA tasks | General visual understanding | [GLM-5V-Turbo](https://docs.z.ai/guides/vlm/glm-5v-turbo) |
| CharXiv (CharXiv Reasoning) | 80.2 | Scientific chart reasoning | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
