# Qwen3.7 Plus

Alibaba · Proprietary · rank 52 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/qwen3-7-plus  
JSON: https://modelscale.dev/api/model/qwen3-7-plus

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `qwen3-7-plus` | — |
| Overall score | 61.8 | Observed 2026-09-22 · source benchlm:models |
| Context window | 1M tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-06-03 | Observed 2026-09-22 · source benchlm:models |
| Access type | Proprietary | — |
| Blended $/1M (75% input / 25% output) | Unavailable | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 58.1 | Observed 2026-09-22 · source benchlm:models |
| Coding | 57.9 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 59.6 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 74 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 72.4 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | 89.2 | Observed 2026-09-22 · source benchlm:models |
| Math | 78.2 | Observed 2026-09-22 · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 31.11 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 69 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Output / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache read / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | Unavailable | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | Unavailable |
| Modelled tokens | Unavailable |
| Reason | The applicable input rate is unavailable. |

## Benchmark record

70 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GPQA (Graduate-Level Google-Proof Q&A) | 90.3 | 448 questions | Graduate level | [GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022) |
| GPQA-D (GPQA Diamond) | 90.3 | Graduate-level science questions | Graduate level | [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking) |
| SuperGPQA (SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines) | 71.4 | 285 disciplines | Graduate level | [SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines](https://arxiv.org/abs/2502.14739) |
| MMLU-Pro (Massive Multitask Language Understanding Professional) | 88.5 | Multiple subjects | Professional level | [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://arxiv.org/abs/2406.01574) |
| HLE (Humanity's Last Exam) | 34.7 | Expert-level questions | Frontier expert level | [Humanity's Last Exam](https://lastexam.ai/) |
| Artificial Analysis Intelligence Index | 25.2 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 90.0 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 35.6 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | 1.1 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 22.5 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 27.7 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| MMLU-Redux | 94.5 | Broad academic QA | Advanced general knowledge | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| MMMLU | 89.0 | Multilingual academic QA | Broad multilingual knowledge | [MMMLU](https://huggingface.co/datasets/openai/MMMLU) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 2.0 | 70.3 | Terminal-based software tasks | Professional software engineering | [Terminal-Bench 2.0](https://www.tbench.ai/) |
| SWE-bench Verified (Software Engineering Benchmark Verified) | 77.7 | 500 verified issues | Professional software engineering | [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770) |
| LiveCodeBench (LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code) | 89.6 | Continuously updated contest problems | Competitive programming level | [LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code](https://arxiv.org/abs/2403.07974) |
| SWE-bench Pro | 57.6 | 1,865 repository problems | Long-horizon professional engineering | [SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?](https://arxiv.org/abs/2509.16941) |
| SWE Multilingual | 75.8 | Multilingual software-engineering tasks | Professional software engineering | [MiniMax M2.7: Early Echoes of Self-Evolution](https://www.minimax.io/news/minimax-m27-en) |
| NL2Repo | 41.1 | Natural language to repository tasks | System-level software comprehension | [MiniMax M2.7: Early Echoes of Self-Evolution](https://www.minimax.io/news/minimax-m27-en) |
| SciCode (Scientific Code Benchmark) | 51.3 | — | — | [BenchLM](https://benchlm.ai/benchmarks/scicode) |
| AA Coding Index (Artificial Analysis Coding Index) | 55.9 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 46.1 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |

### Mathematics

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| HMMT Feb 2026 (Harvard-MIT Mathematics Tournament February 2026) | 92.9 | Competition math problems | Olympiad-style mathematics | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| IMOAnswerBench | 86.0 | Advanced mathematical answer generation | Olympiad-level mathematics | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| Apex | 22.7 | Advanced mathematical reasoning | Frontier math reasoning | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MRCRv2 | 91.7 | Long-context retrieval | Hard long-context | [Introducing GPT-5.2 and GPT-5.2 Pro](https://openai.com/index/introducing-gpt-5-2/) |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 73.0 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| CritPt (Critical Physics Tasks) | 9.1 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Instruction Following

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| IFEval (Instruction-Following Eval) | 94.6 | 541 prompts across 25 instruction types | Instruction precision | [Instruction-Following Evaluation for Large Language Models](https://arxiv.org/abs/2311.07911) |
| IFBench (Instruction Following Benchmark) | 79.1 | — | — | [BenchLM](https://benchlm.ai/benchmarks/ifbench) |
| AA-IFBench (Artificial Analysis IFBench) | 78.0 | Verifiable instruction constraints | Instruction precision | [Artificial Analysis IFBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/ifbench) |

### Multilingual

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MMLU-ProX | 85.4 | Multilingual professional QA | Professional multilingual | [MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation](https://arxiv.org/abs/2503.10497) |
| NOVA-63 | 58.8 | Broad multilingual evaluation | Broad multilingual capability | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| INCLUDE | 83.0 | Cross-lingual understanding | Broad multilingual capability | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| PolyMath | 84.0 | Multilingual math problems | Advanced multilingual reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| MAXIFE | 88.8 | Multilingual instruction following | Advanced multilingual instruction following | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 2.0 | 70.3 | Terminal-based software tasks | Professional software engineering | [Terminal-Bench 2.0](https://www.tbench.ai/) |
| GDPval-AA | 886 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 12.8 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA Agentic Index (Artificial Analysis Agentic Index) | 19.7 | Cross-benchmark agentic index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| APEX-Agents-AA | 22.4 | 452 professional-services agent tasks | Long-horizon workplace agent tasks | [APEX-Agents-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/apex-agents-aa) |
| OSWorld-Verified | 73.3 | 369 real-world computer tasks (361 when eight Google Drive tasks are excluded) | Multi-step desktop and cross-application workflows | [OSWorld](https://os-world.github.io/) |
| OSWorld 2.0 | 2.8 | 108 long-horizon computer-use workflows | Long-horizon professional workflows | [OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks](https://arxiv.org/abs/2606.29537) |
| AndroidWorld | 81.0 | Android app workflows | Complex mobile task completion | [GLM-5V-Turbo](https://docs.z.ai/guides/vlm/glm-5v-turbo) |
| MCP Atlas | 73.2 | Tool-integrated agent tasks | Advanced tool use | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| τ²-bench results (τ²-Bench Tool-Agent-User Evaluation) | 93 | Airline, retail, and telecom customer-service task sets | Dual-control customer-service workflows | [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment](https://arxiv.org/abs/2506.07982) |
| BFCL v4 (Berkeley Function Calling Leaderboard v4) | 72.9 | Function-calling tasks | Advanced tool use | [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking) |
| Claw-Eval | 62.7 | 300 tasks, 2,159 rubrics | Real-world general, multi-turn, and native multimodal agent execution | [Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents](https://arxiv.org/abs/2604.06132) |
| QwenClawBench | 61.8 | Real-world agent workflows | Broad real-world agentic execution | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| QwenWebBench | 1536 | Web artifacts and interactive deliverables | Artifact generation | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| VITA-Bench | 45.6 | Interactive consumer-service agent tasks | Long-horizon real-world workflows | [VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications](https://vitabench.github.io/) |
| DeepPlanning | 62.3 | Travel planning and constrained shopping | Constrained agent planning | [DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints](https://arxiv.org/abs/2601.18137) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 52.8 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MMMU-Pro (Massive Multi-discipline Multimodal Understanding Pro) | 79 | Multimodal academic reasoning | Frontier multimodal | [MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark](https://arxiv.org/abs/2409.02813) |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 80.5 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| OCRBench V2 | 70.7 | Image OCR tasks | Native visual text understanding | [OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning](https://arxiv.org/abs/2501.00321) |
| Design Arena Website (Design Arena Website Elo) | 1283 | Website generation comparisons | Design and website generation | [OpenRouter Grok 4.3 benchmarks](https://openrouter.ai/x-ai/grok-4.3/benchmarks) |
| OmniDocBench 1.5 | 91.4 | Document understanding tasks | Grounded document reasoning | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| RealWorldQA | 86.9 | Real-world visual question answering | General visual reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| Video-MME (with subtitle) (Video-MME with subtitle) | 88.0 | Video understanding | Multimodal video reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| MathVision | 90.3 | Visually grounded math problems | Advanced multimodal mathematics | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| ODINW13 | 51.1 | Out-of-distribution object understanding | Robust visual grounding | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| ERQA | 69.8 | Evidence-based visual QA | Grounded multimodal reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| VideoMMMU | 85.4 | Video-grounded expert reasoning | Frontier multimodal video reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| MLVU (M-Avg) (MLVU mean average) | 87.4 | General video understanding | Broad multimodal video reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| ScreenSpot Pro | 79.0 | 1,581 grounding instructions | Professional GUI grounding | [ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use](https://arxiv.org/abs/2504.07981) |
| MedXpertQA (MM) (MedXpertQA Multimodal) | 71.0 | 2,000 multimodal medical questions | Clinical multimodal reasoning | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| MMSearch-Plus | 41.4 | Hard multimodal search tasks | Advanced multimodal search | [GLM-5V-Turbo](https://docs.z.ai/guides/vlm/glm-5v-turbo) |
| SimpleVQA | 81.7 | Visual QA tasks | General visual understanding | [GLM-5V-Turbo](https://docs.z.ai/guides/vlm/glm-5v-turbo) |
| CharXiv (CharXiv Reasoning) | 85.9 | Scientific chart reasoning | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
