# Kimi K3

Moonshot AI · Pending · rank 7 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/kimi-3  
JSON: https://modelscale.dev/api/model/kimi-3

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `kimi-3` | — |
| Overall score | 74.36 | Observed 2026-09-22 · source benchlm:models |
| Context window | 1.05M tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-07-16 | Observed 2026-09-22 · source benchlm:models |
| Access type | Pending | — |
| Blended $/1M (75% input / 25% output) | $6.00 | Derived from the input and output rates below |
| Cost per successful task (LiveBench) | $0.35 | Observed 2026-09-22 · source livebench:cost |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 91 | Observed 2026-09-22 · source benchlm:models |
| Coding | 73.3 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 83.7 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 78.5 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | 89.4 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | Unavailable | Unavailable · source benchlm:models |
| Math | Unavailable | Unavailable · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 49.73 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 44 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $3.00 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $15.00 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | $0.30 | Observed 2026-09-22 · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $6.00 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | $0.35 | Observed 2026-09-22 · source livebench:cost |

**Self-hosted listing.** Moonshot AI's official Kimi K3 pricing page lists $0.30 cache-hit input, $3.00 cache-miss input, and $15.00 output per million tokens for model ID kimi-k3, with an exact 1,048,576-token context window. OpenRouter independently lists route moonshotai/kimi-k3 with one Moonshot AI INT4 endpoint at the same rates and context length. Prices exclude applicable taxes. Moonshot schedules the full weight release for July 27, 2026, so source type remains pending until the files are public. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $16.90 |
| Modelled tokens | 2.82M |

## Benchmark record

72 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GPQA (Graduate-Level Google-Proof Q&A) | 93.5 | 448 questions | Graduate level | [GPQA: A Graduate-Level Google-Proof Q&A Benchmark](https://arxiv.org/abs/2311.12022) |
| GPQA-D (GPQA Diamond) | 93.5 | Graduate-level science questions | Graduate level | [Trinity-Large-Thinking: Scaling an Open Source Frontier Agent](https://www.arcee.ai/blog/trinity-large-thinking) |
| HLE (Humanity's Last Exam) | 56 | Expert-level questions | Frontier expert level | [Humanity's Last Exam](https://lastexam.ai/) |
| Artificial Analysis Intelligence Index | 43.6 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 93.5 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 46.9 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | 19.7 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 47.6 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 53.2 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| HLE w/o tools (Humanity's Last Exam without tools) | 43.5 | Expert-level questions | Frontier expert level | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| GPQA Diamond (Vals) (GPQA Diamond, Vals AI run) | 92.9 | Graduate-level science questions | Expert reasoning | [Vals AI GPQA Diamond, Vals AI run leaderboard](https://www.vals.ai/benchmarks/gpqa) |
| MMLU-Pro (Vals) (MMLU-Pro, Vals AI run) | 88.0 | Academic multiple-choice questions | Broad academic knowledge | [Vals AI MMLU-Pro, Vals AI run leaderboard](https://www.vals.ai/benchmarks/mmlu_pro) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| VulcanBench v3 | 73.7 | 23 post-cutoff repository tasks in the v3 report | Professional multi-file software engineering | [VulcanBench](https://github.com/morganlinton/VulcanBench/tree/main) |
| OpenHarmony Bench (OpenHarmony Bench v1.0) | 57.3 | 153 app-development and bug-fix tasks | End-to-end OpenHarmony application development | [OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development](https://arxiv.org/abs/2608.16022) |
| ProgramBench (ProgramBench: Can Language Models Rebuild Programs From Scratch?) | 77.8 | 200 program reconstruction tasks | Full-repository software architecture | [ProgramBench: Can Language Models Rebuild Programs From Scratch?](https://programbench.com/static/paper.pdf) |
| PostTrain Bench | 36.6 | Post-training software-engineering tasks | Frontier software engineering | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| FrontierSWE | 81.2 | 17 ultra-long-horizon engineering and research tasks | Ultra-long-horizon frontier software engineering | [FrontierSWE: Benchmarking coding agents at the limits of human abilities](https://www.frontierswe.com/blog) |
| FrontierSWE v2 | 25.9 | 34 ultra-long-horizon engineering and research tasks | Ultra-long-horizon frontier software engineering | [FrontierSWE v2](https://www.frontierswe.com/blog/v2) |
| Kimi Code Bench v2 | 72.9 | Realistic coding-agent tasks | Production software engineering | [Kimi K2.7 Code](https://huggingface.co/moonshotai/Kimi-K2.7-Code) |
| MLS-Bench Lite | 48.3 | 30 machine-learning research tasks | ML research and systems engineering | [MLS-Bench](https://mls-bench.com/) |
| AA Coding Index (Artificial Analysis Coding Index) | 76.2 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 59.5 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| LiveCodeBench (Vals) (LiveCodeBench, Vals AI run) | 87.2 | Competitive programming problems (easy, medium, hard) | Frontier coding | [Vals AI LiveCodeBench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/lcb) |
| SWE-bench (Vals) (SWE-bench, Vals AI run) | 93.4 | Real repository issues by human time bucket | Frontier coding agents | [Vals AI SWE-bench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/swebench) |
| DeepSWE | 67.5 | 113 software engineering tasks across 91 repositories and 5 languages | Long-horizon software engineering | [DeepSWE benchmark blog](https://deepswe.datacurve.ai/blog) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 88.7 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| MLCR-AA (Medical Long Context Reasoning (MLCR-AA)) | 38.3 | Long, fragmented medical-record reasoning | Long-context medical reasoning | [Medical Long Context Reasoning (MLCR-AA)](https://artificialanalysis.ai/evaluations/mlcr-aa) |
| CritPt (Critical Physics Tasks) | 23.4 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA Briefcase (Artificial Analysis Briefcase) | 1510 | Professional knowledge-work tasks | Professional work | [Artificial Analysis Briefcase Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-briefcase) |
| AA AutomationBench (Artificial Analysis AutomationBench) | 58.3 | Business-process automation tasks | Agentic automation | [Artificial Analysis AutomationBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/automationbench-aa) |
| AA EnterpriseOps-Gym (Artificial Analysis EnterpriseOps-Gym) | 45.3 | Enterprise operations workflows | Enterprise agent operations | [Artificial Analysis EnterpriseOps-Gym Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa) |
| AA Harvey LAB (Artificial Analysis Harvey LAB-AA) | 94.6 | Legal agent tasks | Professional legal work | [Artificial Analysis Harvey LAB-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/harvey-lab-aa) |
| AA ITBench (Artificial Analysis ITBench-AA) | 47.7 | IT incident-response tasks | Enterprise IT operations | [Artificial Analysis ITBench-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/itbench-aa) |
| AA Tau3 Banking (Artificial Analysis Tau3-Banking) | 46.0 | Banking tool-use workflows | Agentic banking workflows | [Artificial Analysis Tau3-Banking Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/tau3-banking) |
| Terminal-Bench 2.0 | 88.3 | Terminal-based software tasks | Professional software engineering | [Terminal-Bench 2.0](https://www.tbench.ai/) |
| BrowseComp | 91.2 | Research questions requiring browsing | Hard web research | [BrowseComp](https://openai.com/index/browsecomp/) |
| GDPval-AA | 1524 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 51.2 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA Agentic Index (Artificial Analysis Agentic Index) | 50.6 | Cross-benchmark agentic index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA Terminal-Bench 4.0 (Artificial Analysis Terminal-Bench v4.0) | 12.6 | Terminal-based agent tasks | Agentic software engineering | [Artificial Analysis Terminal-Bench v4.0 Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/terminalbench-v4-0) |
| GDP.pdf (Artificial Analysis GDP.pdf) | 22.0 | Professional document-production tasks | Professional knowledge work | [GDP.pdf Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gdp-pdf) |
| AA-AnalystAgent (Artificial Analysis AnalystAgent) | 38.8 | Spreadsheet and document analysis questions | Business and data analysis | [AA-AnalystAgent Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-analyst-agent) |
| APEX-Agents-AA | 41.3 | 452 professional-services agent tasks | Long-horizon workplace agent tasks | [APEX-Agents-AA Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/apex-agents-aa) |
| JobBench | 52.9 | 130 tasks across 35 occupations | Professional multi-source workflows | [JobBench: Aligning Agent Work With Human Will](https://arxiv.org/abs/2605.26329) |
| MCP Atlas | 84.2 | Tool-integrated agent tasks | Advanced tool use | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| Toolathlon-Verified | 73.2 | Verified multi-tool workflows | Advanced tool use | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| AutomationBench | 30.8 | 600 public automation tasks | Long-horizon automation | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| APEX-Agents | 37.6 | Professional-services agent tasks | Long-horizon professional work | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| SpreadsheetBench 2 | 34.8 | Spreadsheet analysis and editing tasks | Professional spreadsheet work | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| DECK-Bench (DECK-Bench (Internal)) | 73.5 | Internal presentation workflows | Professional presentation creation | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| DeepSearchQA | 95.0 | Agentic browsing and list-answer questions | Agentic web research | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 80.9 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |
| ApprenticeBench (ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job) | 18 | 100 vendor bills processed in sequence inside a simulated construction company | Long-horizon computer use with offline and online continual learning | [ApprenticeBench: a step change in AI's job readiness](https://neocognition.io/blog/apprentice-bench/) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MMMU-Pro (Massive Multi-discipline Multimodal Understanding Pro) | 81.6 | Multimodal academic reasoning | Frontier multimodal | [MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark](https://arxiv.org/abs/2409.02813) |
| AA-MMMU-Pro (Artificial Analysis MMMU-Pro) | 80.5 | Multimodal academic reasoning | Frontier multimodal | [Artificial Analysis MMMU-Pro Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/mmmu-pro) |
| Design Arena Website (Design Arena Website Elo) | 1351 | Website generation comparisons | Design and website generation | [OpenRouter Grok 4.3 benchmarks](https://openrouter.ai/x-ai/grok-4.3/benchmarks) |
| OfficeQA Pro | 63.3 | Document and spreadsheet tasks | Enterprise grounded reasoning | [OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning](https://arxiv.org/abs/2603.08655) |
| MathVision w/ Python (MathVision with Python) | 97.8 | Visual mathematics problems with Python | Advanced multimodal mathematics | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| BabyVision w/ Python (BabyVision with Python) | 85.7 | Visual perception tasks with Python | Fine-grained visual perception | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| ZeroBench w/ Python (ZeroBench_main with Python) | 41.0 | Visual reasoning questions with Python | Tool-augmented visual reasoning | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| WorldVQA ForceAnswer | 51.0 | Atomic visual world-knowledge questions | Fine-grained visual knowledge | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| OmniDocBench | 91.1 | Complex document-understanding tasks | Grounded document reasoning | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| PerceptionBench (PerceptionBench (Internal)) | 58.5 | Internal atomic visual-perception tasks | Fine-grained visual perception | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| MMMU-Pro w/ Python (MMMU-Pro with Python) | 83.4 | Multimodal academic reasoning | Frontier multimodal | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| MathVision | 94.3 | Visually grounded math problems | Advanced multimodal mathematics | [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6) |
| ZeroBench | 23.0 | 100 visual reasoning questions | Tool-augmented visual reasoning | [Muse Spark Eval Methodology](https://ai.meta.com/static-resource/muse-spark-eval-methodology) |
| CharXiv (CharXiv Reasoning) | 91.3 | Scientific chart reasoning | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |
| CharXiv w/o tools (CharXiv Reasoning without tools) | 84.8 | Scientific chart reasoning (tool-free) | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |

### external

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| ExploitBench (ExploitBench v8-bench) | 32 | V8 exploit synthesis runs | Browser exploitation and cybersecurity | [ExploitBench](https://exploitbench.ai/) |
| ACE solved (ACE Cyber Range Challenges Solved) | 0 | 41 advanced cyber-range challenges | Advanced cyber operations | [UK AISI and CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities) |
| The Last Ones steps (The Last Ones Average Progress) | 17 | 32-step long-horizon cyber range | Long-horizon cyber operations | [UK AISI and CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities) |
| The Last Ones completion (The Last Ones Cyber Range Completion Rate) | 10 | 10 long-horizon runs | Long-horizon cyber operations | [UK AISI and CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities](https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
