# Gemini 3.7 Flash

Google · Proprietary · rank 22 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/gemini-3-7-flash  
JSON: https://modelscale.dev/api/model/gemini-3-7-flash

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `gemini-3-7-flash` | — |
| Overall score | 68.72 | Observed 2026-09-22 · source benchlm:models |
| Context window | 1M tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-08-13 | Observed 2026-09-22 · source benchlm:models |
| Access type | Proprietary | — |
| Blended $/1M (75% input / 25% output) | $1.50 | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 65.7 | Observed 2026-09-22 · source benchlm:models |
| Coding | 76.9 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 82.3 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 77.2 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 82.6 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | Unavailable | Unavailable · source benchlm:models |
| Math | Unavailable | Unavailable · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 11.96 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 335 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $0.75 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $3.75 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | $0.075 | Observed 2026-09-22 · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $1.50 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** Google lists introductory Gemini 3.7 Flash rates through December 31, 2026: $0.75 input, $0.075 cached input, and $3.75 output per million tokens. Flex and Batch input and output rates are half those amounts. On January 1, 2027, standard rates rise to $1.50 input, $0.15 cached input, and $7.50 output; Flex and Batch rates rise to $0.75 input, $0.075 cached input, and $3.75 output. Priority rates are $1.35 input, $0.135 cached input, and $6.75 output during the promotion, doubling in 2027. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $4.22 |
| Modelled tokens | 2.82M |

## Benchmark record

40 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| BioMysteryBench (human-solvable) (BioMysteryBench Human Solvable) | 87.1 | Human-solvable computational biology investigations | Expert computational biology | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| BioMysteryBench (human-difficult) (BioMysteryBench Human Difficult) | 43.5 | Human-difficult computational biology investigations | Frontier computational biology | [Claude Opus 5 System Card](https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb9f5bdeaf48/Claude%20Opus%205%20System%20Card.pdf) |
| HLE-Verified | 53.6 | 1,811 verified or revised expert questions | Frontier multidisciplinary expert reasoning | [HLE-Verified: A Systematic Verification and Structured Revision of Humanity's Last Exam](https://arxiv.org/abs/2602.13964) |
| LABBench2 (LABBench2: An Improved Benchmark for AI Systems Performing Biology Research) | 82.1 | Nearly 1,900 biology-research tasks | Real-world biology research | [LABBench2: An Improved Benchmark for AI Systems Performing Biology Research](https://arxiv.org/abs/2604.09554) |
| Artificial Analysis Intelligence Index | 39.1 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 94.5 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 47.9 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | 26.5 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 55.3 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 64.5 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| GPQA Diamond (Vals) (GPQA Diamond, Vals AI run) | 93.9 | Graduate-level science questions | Expert reasoning | [Vals AI GPQA Diamond, Vals AI run leaderboard](https://www.vals.ai/benchmarks/gpqa) |
| MMLU-Pro (Vals) (MMLU-Pro, Vals AI run) | 90.1 | Academic multiple-choice questions | Broad academic knowledge | [Vals AI MMLU-Pro, Vals AI run leaderboard](https://www.vals.ai/benchmarks/mmlu_pro) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 2.1 (Terminal-Bench 2.1 (provider run)) | 85.8 | Terminal-based software-agent tasks | Professional software engineering | [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/) |
| FrontierCode 1.1 Main | 43.6 | 100 private Main tasks (150 in Extended) | Frontier coding-agent quality | [FrontierCode leaderboard](https://cognition.com/frontiercode) |
| FrontierSWE v2 | 20.3 | 34 ultra-long-horizon engineering and research tasks | Ultra-long-horizon frontier software engineering | [FrontierSWE v2](https://www.frontierswe.com/blog/v2) |
| AA Coding Index (Artificial Analysis Coding Index) | 76.1 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 57.2 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| LiveCodeBench (Vals) (LiveCodeBench, Vals AI run) | 88.7 | Competitive programming problems (easy, medium, hard) | Frontier coding | [Vals AI LiveCodeBench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/lcb) |
| SWE-bench (Vals) (SWE-bench, Vals AI run) | 80.8 | Real repository issues by human time bucket | Frontier coding agents | [Vals AI SWE-bench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/swebench) |
| DeepSWE | 65.3 | 113 software engineering tasks across 91 repositories and 5 languages | Long-horizon software engineering | [DeepSWE benchmark blog](https://deepswe.datacurve.ai/blog) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MRCR v2 64K-128K (OpenAI MRCR v2 8-needle 64K-128K) | 97 | 8-needle retrieval tasks | Long-context reasoning | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 81.7 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| CritPt (Critical Physics Tasks) | 14.3 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 3.0 | 14.9 | 74 professional computer-work tasks across 7 domains | Frontier autonomous knowledge work | [Terminal-Bench 3.0](https://www.frontierbench.ai/) |
| AA Harvey LAB (Artificial Analysis Harvey LAB-AA) | 90.7 | Legal agent tasks | Professional legal work | [Artificial Analysis Harvey LAB-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/harvey-lab-aa) |
| Terminal-Bench 2.1 (Terminal-Bench 2.1 (provider run)) | 85.8 | Terminal-based software-agent tasks | Professional software engineering | [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/) |
| GDPval-AA | 1525 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 43.6 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA Agentic Index (Artificial Analysis Agentic Index) | 36.4 | Cross-benchmark agentic index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-AnalystAgent (Artificial Analysis AnalystAgent) | 60.0 | Spreadsheet and document analysis questions | Business and data analysis | [AA-AnalystAgent Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-analyst-agent) |
| OSWorld 2.0 | 47.9 | 108 long-horizon computer-use workflows | Long-horizon professional workflows | [OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks](https://arxiv.org/abs/2606.29537) |
| AutomationBench | 30.4 | 600 public automation tasks | Long-horizon automation | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| Agents' Last Exam | 26.3 | Agent tasks | Advanced agentic work | [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 77.5 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |
| ApprenticeBench (ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job) | 16 | 100 vendor bills processed in sequence inside a simulated construction company | Long-horizon computer use with offline and online continual learning | [ApprenticeBench: a step change in AI's job readiness](https://neocognition.io/blog/apprentice-bench/) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 85.5 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| Design Arena Website (Design Arena Website Elo) | 1314 | Website generation comparisons | Design and website generation | [OpenRouter Grok 4.3 benchmarks](https://openrouter.ai/x-ai/grok-4.3/benchmarks) |
| LVBench | 85.4 | Long-form video question answering | Extended temporal reasoning | [Qwen3.8-Max: A New Bar for Coding and Cowork](https://qwen.ai/blog?id=qwen3.8) |
| CharXiv (CharXiv Reasoning) | 88.7 | Scientific chart reasoning | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |
| CharXiv w/o tools (CharXiv Reasoning without tools) | 84.5 | Scientific chart reasoning (tool-free) | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
