Skip to main content
ModelScale

Qwen3.8 Max

Alibaba · Open Weight · rank 10 · bench-align-v5

Canonical idqwen3-8-max
Overall score71.76
Context window1M tokens
Release date2026-08-03
Access typeOpen Weight
Blended $/1MUnavailable75% input / 25% output

Capability shape

Seven axes from the ranking source. A missing axis is drawn as a gap.

Capability evidence

Agentic83.6
Coding68
Knowledge67.4
Reasoning86.5
Multimodal & Grounded87.4
Instruction Following90.7
MathUnavailable

Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim.

Time to first token: latency to the first answer chunk (Artificial Analysis, via BenchLM). Reasoning models include thinking time, so values can run to tens or hundreds of seconds.

Evidence key: Observed

Runtime measurements
MeasurementValueObservedLast goodEvidence
Time to first token51.78 s2026-09-172026-09-17
Throughput41 tok/s2026-09-172026-09-17

Regional or per-endpoint measurements appear only when the API supplies them; none are modelled here.

Endpoint and price matrix

Every published price component, including cache reads and writes.

Price components
ComponentUSDEvidence
Input / 1M tokensUnavailable
Output / 1M tokensUnavailable
Cache read / 1M tokensUnavailable
Cache write / 1M tokensUnavailable
Blended / 1M (75% input / 25% output)Unavailable
Self-hosted listingQwen publishes the Qwen3.8-2.4T-A95B checkpoint for self-hosting under the custom Qwen3.8-Max License. Alibaba Cloud Model Studio's official pricing page separately lists the hosted qwen3.8-max SKU in both non-thinking and thinking modes for the 0 < Token <= 1M tier. The China and global tables list CNY 12 input / CNY 36 output per million tokens; the US international table lists CNY 14.988 input / CNY 44.965 output. The USD numeric fields stay null because we do not convert a non-USD first-party price.No hosted token rate was published for this model, so its per-token price is unavailable rather than zero.

Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. This uses the same calculator as the cost simulator, so an unavailable applicable rate makes the example unavailable too.

Modelled monthly costUnavailableThe applicable input rate is unavailable.
Modelled tokensUnavailable
Open the simulatorChange this workload

Benchmark record

60 matched benchmark rows with their published value, unit, and provenance.

Knowledge

Knowledge benchmarks
BenchmarkValueTasksDifficultyProvenance
GPQA
Graduate-Level Google-Proof Q&A
92.6448 questionsGraduate levelGPQA: A Graduate-Level Google-Proof Q&A Benchmark
GPQA-D
GPQA Diamond
92.6Graduate-level science questionsGraduate levelTrinity-Large-Thinking: Scaling an Open Source Frontier Agent
HLE
Humanity's Last Exam
43.6Expert-level questionsFrontier expert levelHumanity's Last Exam
HLE w/o tools
Humanity's Last Exam without tools
43.6Expert-level questionsFrontier expert levelIntroducing GPT-5.4 mini and nano
GPQA Diamond (Vals)
GPQA Diamond, Vals AI run
93.7Graduate-level science questionsExpert reasoningVals AI GPQA Diamond, Vals AI run leaderboard
MMLU-Pro (Vals)
MMLU-Pro, Vals AI run
88.6Academic multiple-choice questionsBroad academic knowledgeVals AI MMLU-Pro, Vals AI run leaderboard

Coding

Coding benchmarks
BenchmarkValueTasksDifficultyProvenance
Terminal-Bench 2.1
Terminal-Bench 2.1 (provider run)
86.6Terminal-based software-agent tasksProfessional software engineeringDeepSeek V4 Flash 0731 update
SWE-bench Pro
SWE-bench Pro
67.71,865 repository problemsLong-horizon professional engineeringSWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
VulcanBench v3
VulcanBench v3
81.223 post-cutoff repository tasks in the v3 reportProfessional multi-file software engineeringVulcanBench
OpenHarmony Bench
OpenHarmony Bench v1.0
60.8153 app-development and bug-fix tasksEnd-to-end OpenHarmony application developmentOpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development
FrontierSWE
FrontierSWE
73.517 ultra-long-horizon engineering and research tasksUltra-long-horizon frontier software engineeringFrontierSWE: Benchmarking coding agents at the limits of human abilities
FrontierSWE v2
FrontierSWE v2
15.834 ultra-long-horizon engineering and research tasksUltra-long-horizon frontier software engineeringFrontierSWE v2
MLS-Bench Lite
MLS-Bench Lite
41.030 machine-learning research tasksML research and systems engineeringMLS-Bench
PaperBench
PaperBench
93.0AI research-paper reproductionFrontier autonomous research and engineeringQwen3.8-Max: A New Bar for Coding and Cowork
QwenReactBench
QwenReactBench
1724Bilingual React project constructionProduction frontend developmentQwen3.8-Max: A New Bar for Coding and Cowork
NL2Repo
NL2Repo
55.9Natural language to repository tasksSystem-level software comprehensionMiniMax M2.7: Early Echoes of Self-Evolution
LiveCodeBench (Vals)
LiveCodeBench, Vals AI run
87.9Competitive programming problems (easy, medium, hard)Frontier codingVals AI LiveCodeBench, Vals AI run leaderboard
SWE-bench (Vals)
SWE-bench, Vals AI run
85.6Real repository issues by human time bucketFrontier coding agentsVals AI SWE-bench, Vals AI run leaderboard
DeepSWE
DeepSWE
56.6113 software engineering tasks across 91 repositories and 5 languagesLong-horizon software engineeringDeepSWE benchmark blog

Reasoning

Reasoning benchmarks
BenchmarkValueTasksDifficultyProvenance
LongBench v2
LongBench v2
66.3Long-context tasksHard long-contextLongBench v2
MRCRv2
MRCRv2
92.9Long-context retrievalHard long-contextIntroducing GPT-5.2 and GPT-5.2 Pro

Instruction Following

Instruction Following benchmarks
BenchmarkValueTasksDifficultyProvenance
IFBench
Instruction Following Benchmark
82.8BenchLM

Agentic

Agentic benchmarks
BenchmarkValueTasksDifficultyProvenance
Terminal-Bench 2.1
Terminal-Bench 2.1 (provider run)
86.6Terminal-based software-agent tasksProfessional software engineeringDeepSeek V4 Flash 0731 update
HLE w/ tools
Humanity's Last Exam with tools
56.2Expert questions with tool useFrontier tool-augmented reasoningDeepSeek-V4 Technical Report
OSWorld-Verified
OSWorld-Verified
86.1369 real-world computer tasks (361 when eight Google Drive tasks are excluded)Multi-step desktop and cross-application workflowsOSWorld
OSWorld 2.0
OSWorld 2.0
19.4108 long-horizon computer-use workflowsLong-horizon professional workflowsOSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
JobBench
JobBench
53.4130 tasks across 35 occupationsProfessional multi-source workflowsJobBench: Aligning Agent Work With Human Will
AndroidWorld
AndroidWorld
85.3Android app workflowsComplex mobile task completionGLM-5V-Turbo
Toolathlon-Verified
Toolathlon-Verified
72.5Verified multi-tool workflowsAdvanced tool useKimi K3: Open Frontier Intelligence
AutomationBench
AutomationBench
27.3600 public automation tasksLong-horizon automationKimi K3: Open Frontier Intelligence
Agents' Last Exam
Agents' Last Exam
52.4Agent tasksAdvanced agentic workDeepSeek V4 Flash 0731 update
WideResearch
WideResearch
81.9Open-ended research tasksBroad research-agent workflowsQwen3.6 launch benchmarks
CoWorkBench
CoWorkBench
74.8Long-horizon professional workflowsCross-domain professional workQwen3.8-Max: A New Bar for Coding and Cowork
MobileWorld
MobileWorld
77.8Interactive mobile-device workflowsLong-horizon mobile computer useQwen3.8-Max: A New Bar for Coding and Cowork
WebArena-Verified
WebArena-Verified Browser Agent Benchmark
66.8812 verified tasks; separate 258-task Hard subsetAudited stateful browser workWebArena-Verified: A Fully Audited Benchmark for Web Agents
Terminal-Bench 2.1 (Vals)
Terminal-Bench 2.1, Vals AI run
67.4Difficult terminal tasksFrontier agenticVals AI Terminal-Bench 2.1, Vals AI run leaderboard

Multimodal & Grounded

Multimodal & Grounded benchmarks
BenchmarkValueTasksDifficultyProvenance
MMMU-Pro
Massive Multi-discipline Multimodal Understanding Pro
82.3Multimodal academic reasoningFrontier multimodalMMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
OCRBench V2
OCRBench V2
74.2Image OCR tasksNative visual text understandingOCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
MathVision w/ Python
MathVision with Python
97.7Visual mathematics problems with PythonAdvanced multimodal mathematicsKimi K3: Open Frontier Intelligence
BabyVision w/ Python
BabyVision with Python
91.3Visual perception tasks with PythonFine-grained visual perceptionKimi K3: Open Frontier Intelligence
ZeroBench w/ Python
ZeroBench_main with Python
49.0Visual reasoning questions with PythonTool-augmented visual reasoningKimi K3: Open Frontier Intelligence
PerceptionBench
PerceptionBench (Internal)
63.5Internal atomic visual-perception tasksFine-grained visual perceptionKimi K3: Open Frontier Intelligence
OmniDocBench 1.5
OmniDocBench 1.5
92.1Document understanding tasksGrounded document reasoningIntroducing GPT-5.4 mini and nano
RealWorldQA
RealWorldQA
88.0Real-world visual question answeringGeneral visual reasoningQwen3.6 launch benchmarks
Video-MME (with subtitle)
Video-MME with subtitle
90.4Video understandingMultimodal video reasoningQwen3.6 launch benchmarks
MathVision
MathVision
95.2Visually grounded math problemsAdvanced multimodal mathematicsQwen3.6 launch benchmarks
CC-OCR
CC-OCR
79.6Optical character recognitionDocument readingQwen3.6 launch benchmarks
ERQA
ERQA
77.8Evidence-based visual QAGrounded multimodal reasoningQwen3.6 launch benchmarks
VideoMMMU
VideoMMMU
88.7Video-grounded expert reasoningFrontier multimodal video reasoningQwen3.6 launch benchmarks
MLVU (M-Avg)
MLVU mean average
90.8General video understandingBroad multimodal video reasoningQwen3.6 launch benchmarks
LVBench
LVBench
81.8Long-form video question answeringExtended temporal reasoningQwen3.8-Max: A New Bar for Coding and Cowork
MMVU
Multimodal Multi-disciplinary Video Understanding
82.4Video understandingMulti-disciplinary multimodal video reasoningKimi K2.5 benchmark release surface
ScreenSpot Pro
ScreenSpot Pro
84.51,581 grounding instructionsProfessional GUI groundingScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
MedXpertQA (MM)
MedXpertQA Multimodal
80.42,000 multimodal medical questionsClinical multimodal reasoningMuse Spark Eval Methodology
ZeroBench
ZeroBench
24.0100 visual reasoning questionsTool-augmented visual reasoningMuse Spark Eval Methodology
Vision2Web
Vision2Web
69.0Screenshot-to-web tasksMultimodal web generationGLM-5V-Turbo
SimpleVQA
SimpleVQA
75.0Visual QA tasksGeneral visual understandingGLM-5V-Turbo
CharXiv
CharXiv Reasoning
93.5Scientific chart reasoningScientific visualization reasoningCharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
CharXiv w/o tools
CharXiv Reasoning without tools
88.4Scientific chart reasoning (tool-free)Scientific visualization reasoningCharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
BabyVision
BabyVision
82.0Visual perception tasksFine-grained visual perceptionMuse Spark 1.1 Evaluation Report

Lifecycle and limitations log

Lifecycle events the source associates with this model, plus what this profile does not claim.

No lifecycle event references this model.

That is not evidence the model has no lifecycle plan — only that this source published none.

What this profile does not claimValues are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned — the attempt timestamp is in each badge. Last attempted fetch for this model’s score: 2026-09-17 17:17 UTC.