# Qwen3.5-35B-A3B

Alibaba · Open Weight · rank 99 · bench-align-v5 · Self-hosted

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/qwen3-5-35b-a3b  
JSON: https://modelscale.dev/api/model/qwen3-5-35b-a3b

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `qwen3-5-35b-a3b` | — |
| Overall score | 53.74 | Observed 2026-09-22 · source benchlm:models |
| Context window | 262K tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-03-04 | Observed 2026-09-22 · source benchlm:models |
| Access type | Open Weight | — |
| Blended $/1M (75% input / 25% output) | Unavailable | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 13.7 | Observed 2026-09-22 · source benchlm:models |
| Coding | 47.6 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 70 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 40.2 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 67.2 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | 87.4 | Observed 2026-09-22 · source benchlm:models |
| Math | Unavailable | Unavailable · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 15.67 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 147 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Output / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache read / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | Unavailable | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** Open-weight model. Self-hosted or third-party hosted costs vary by provider and infrastructure. No hosted token rate was published for this model, so its per-token price is unavailable rather than zero.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | Unavailable |
| Modelled tokens | Unavailable |
| Reason | The applicable input rate is unavailable. |

## Benchmark record

27 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GPQA (Graduate-Level Google-Proof Q&A) | 84.2 | 448 questions | Graduate level | [GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022) |
| SuperGPQA (SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines) | 63.4 | 285 disciplines | Graduate level | [SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines](https://arxiv.org/abs/2502.14739) |
| MMLU-Pro (Massive Multitask Language Understanding Professional) | 85.3 | Multiple subjects | Professional level | [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://arxiv.org/abs/2406.01574) |
| Artificial Analysis Intelligence Index | 19.3 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 84.5 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 21.0 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | -48.1 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 20.1 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 85.4 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| SWE-bench Verified (Software Engineering Benchmark Verified) | 69.2 | 500 verified issues | Professional software engineering | [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770) |
| SWE-Rebench | 53.7 | Fresh GitHub issues (rolling window) | Professional software engineering | [SWE-Rebench: Contamination-Free Evaluation of Software Engineering Agents](https://swe-rebench.com/) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| LongBench v2 | 59 | Long-context tasks | Hard long-context | [LongBench v2](https://arxiv.org/abs/2412.15204) |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 72.0 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| CritPt (Critical Physics Tasks) | 0.9 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Instruction Following

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| IFEval (Instruction-Following Eval) | 91.9 | 541 prompts across 25 instruction types | Instruction precision | [Instruction-Following Evaluation for Large Language Models](https://arxiv.org/abs/2311.07911) |
| AA-IFBench (Artificial Analysis IFBench) | 72.5 | Verifiable instruction constraints | Instruction precision | [Artificial Analysis IFBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/ifbench) |

### Multilingual

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MMLU-ProX | 81 | Multilingual professional QA | Professional multilingual | [MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation](https://arxiv.org/abs/2503.10497) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 2.0 | 40.5 | Terminal-based software tasks | Professional software engineering | [Terminal-Bench 2.0](https://www.tbench.ai/) |
| BrowseComp | 61 | Research questions requiring browsing | Hard web research | [BrowseComp](https://openai.com/index/browsecomp/) |
| Gert Labs (Gert Labs Composite Game Benchmark) | 28.96 | Novel game environments | Agentic coding and decision-making | [Gert Labs rankings](https://gertlabs.com/rankings) |
| OSWorld-Verified | 54.5 | 369 real-world computer tasks (361 when eight Google Drive tasks are excluded) | Multi-step desktop and cross-application workflows | [OSWorld](https://os-world.github.io/) |
| τ²-bench results (τ²-Bench Tool-Agent-User Evaluation) | 89.2 | Airline, retail, and telecom customer-service task sets | Dual-control customer-service workflows | [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment](https://arxiv.org/abs/2506.07982) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MMMU (Massive Multi-discipline Multimodal Understanding) | 81.4 | Multimodal academic reasoning | Frontier multimodal | [MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI](https://arxiv.org/abs/2401.05508) |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 72.7 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| MathVision | 83.9 | Visually grounded math problems | Advanced multimodal mathematics | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| MMVU (Multimodal Multi-disciplinary Video Understanding) | 72.3 | Video understanding | Multi-disciplinary multimodal video reasoning | [Kimi K2.5 benchmark release surface](https://www.kimi.com/blog/kimi-k2-5.html) |
| V* | 92.7 | Frontier multimodal reasoning tasks | Frontier multimodal | [GLM-5V-Turbo](https://docs.z.ai/guides/vlm/glm-5v-turbo) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
