# Qwen3.8-Flash-Next

Alibaba · Open Weight · rank 79 · bench-align-v5 · Self-hosted

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/qwen3-8-flash-next  
JSON: https://modelscale.dev/api/model/qwen3-8-flash-next

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `qwen3-8-flash-next` | — |
| Overall score | 57.03 | Observed 2026-09-22 · source benchlm:models |
| Context window | 262K tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-08-26 | Observed 2026-09-22 · source benchlm:models |
| Access type | Open Weight | — |
| Blended $/1M (75% input / 25% output) | Unavailable | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 41 | Observed 2026-09-22 · source benchlm:models |
| Coding | 54.9 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 52.8 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 75.8 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 83.1 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | 87.2 | Observed 2026-09-22 · source benchlm:models |
| Math | Unavailable | Unavailable · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 34.08 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 64 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Output / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache read / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | Unavailable | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** Qwen publishes Qwen3.8-Flash-Next under the Qwen Community 1.0 license for self-hosting and does not publish a distinct first-party hosted token rate for this exact experimental checkpoint. BenchLM represents the open-weight row as self-host/free-per-token before infrastructure costs. Qwen Cloud's production Qwen3.8-Flash is a separate model based on this architecture, so this row does not inherit its price or default 1M context. No hosted token rate was published for this model, so its per-token price is unavailable rather than zero.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | Unavailable |
| Modelled tokens | Unavailable |
| Reason | The applicable input rate is unavailable. |

## Benchmark record

38 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GPQA (Graduate-Level Google-Proof Q&A) | 91.7 | 448 questions | Graduate level | [GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022) |
| GPQA-D (GPQA Diamond) | 91.7 | Graduate-level science questions | Graduate level | [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking) |
| HLE (Humanity's Last Exam) | 35.9 | Expert-level questions | Frontier expert level | [Humanity's Last Exam](https://lastexam.ai/) |
| Artificial Analysis Intelligence Index | 39.8 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 92.3 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 38.0 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | -9.7 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 24.5 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 45.3 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| HLE w/o tools (Humanity's Last Exam without tools) | 35.9 | Expert-level questions | Frontier expert level | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| LiveCodeBench v6 | 91.9 | Fresh programming problems | Competitive programming level | [LiveCodeBench official repository and release documentation](https://github.com/LiveCodeBench/LiveCodeBench) |
| SWE-bench Pro | 62.5 | 1,865 repository problems | Long-horizon professional engineering | [SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?](https://arxiv.org/abs/2509.16941) |
| SWE Multilingual | 81 | Multilingual software-engineering tasks | Professional software engineering | [MiniMax M2.7: Early Echoes of Self-Evolution](https://www.minimax.io/news/minimax-m27-en) |
| NL2Repo | 48.1 | Natural language to repository tasks | System-level software comprehension | [MiniMax M2.7: Early Echoes of Self-Evolution](https://www.minimax.io/news/minimax-m27-en) |
| AA Coding Index (Artificial Analysis Coding Index) | 73.0 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 50.6 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| DeepSWE | 58.7 | 113 software engineering tasks across 91 repositories and 5 languages | Long-horizon software engineering | [DeepSWE benchmark blog](https://deepswe.datacurve.ai/blog) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 79.7 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| CritPt (Critical Physics Tasks) | 11.1 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Instruction Following

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| IFBench (Instruction Following Benchmark) | 81.3 | — | — | [BenchLM](https://benchlm.ai/benchmarks/ifbench) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA Briefcase (Artificial Analysis Briefcase) | 1597 | Professional knowledge-work tasks | Professional work | [Artificial Analysis Briefcase Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-briefcase) |
| GDPval-AA | 1648 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 55.6 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| OSWorld 2.0 | 19.4 | 108 long-horizon computer-use workflows | Long-horizon professional workflows | [OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks](https://arxiv.org/abs/2606.29537) |
| JobBench | 55.7 | 130 tasks across 35 occupations | Professional multi-source workflows | [JobBench: Aligning Agent Work With Human Will](https://arxiv.org/abs/2605.26329) |
| AndroidWorld | 84.5 | Android app workflows | Complex mobile task completion | [GLM-5V-Turbo](https://docs.z.ai/guides/vlm/glm-5v-turbo) |
| Toolathlon-Verified | 73.5 | Verified multi-tool workflows | Advanced tool use | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| Agents' Last Exam | 51.2 | Agent tasks | Advanced agentic work | [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/) |
| CoWorkBench | 73.9 | Long-horizon professional workflows | Cross-domain professional work | [Qwen3.8-Max: A New Bar for Coding and Cowork](https://qwen.ai/blog?id=qwen3.8) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 79.8 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| MathVision w/ Python (MathVision with Python) | 95.7 | Visual mathematics problems with Python | Advanced multimodal mathematics | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| RealWorldQA | 88.5 | Real-world visual question answering | General visual reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| MathVision | 90.6 | Visually grounded math problems | Advanced multimodal mathematics | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| ERQA | 72.3 | Evidence-based visual QA | Grounded multimodal reasoning | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| LVBench | 76.6 | Long-form video question answering | Extended temporal reasoning | [Qwen3.8-Max: A New Bar for Coding and Cowork](https://qwen.ai/blog?id=qwen3.8) |
| Vision2Web | 64.0 | Screenshot-to-web tasks | Multimodal web generation | [GLM-5V-Turbo](https://docs.z.ai/guides/vlm/glm-5v-turbo) |
| CharXiv (CharXiv Reasoning) | 90.6 | Scientific chart reasoning | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |
| CharXiv w/o tools (CharXiv Reasoning without tools) | 84.6 | Scientific chart reasoning (tool-free) | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
