# Kimi K2.7 Code

Moonshot AI · Open Weight · rank 33 · bench-align-v5

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/kimi-k2-7-code  
JSON: https://modelscale.dev/api/model/kimi-k2-7-code

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `kimi-k2-7-code` | — |
| Overall score | 65.94 | Observed 2026-09-22 · source benchlm:models |
| Context window | 256K tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-06-12 | Observed 2026-09-22 · source benchlm:models |
| Access type | Open Weight | — |
| Blended $/1M (75% input / 25% output) | $1.71 | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 71 | Observed 2026-09-22 · source benchlm:models |
| Coding | 46 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 65.2 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | 75.5 | Observed 2026-09-22 · source benchlm:models |
| Multimodal & Grounded | Unavailable | Unavailable · source benchlm:models |
| Instruction Following | 75.1 | Observed 2026-09-22 · source benchlm:models |
| Math | Unavailable | Unavailable · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | 38.26 s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |
| Throughput | 63 tok/s | 2026-09-22 | 2026-09-22 | Observed 2026-09-22 · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | $0.95 | Observed 2026-09-22 · source benchlm:pricing |
| Output / 1M tokens | $4.00 | Observed 2026-09-22 · source benchlm:pricing |
| Cache read / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | $1.71 | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** Moonshot's Kimi API platform lists kimi-k2.7-code at $0.95 cache-miss input / $4.00 output per million tokens, with cache-hit input priced at $0.19 per million tokens. The rates above are a hosted price matched from another provider, not a first-party list price.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | $4.82 |
| Modelled tokens | 2.82M |

## Benchmark record

26 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Artificial Analysis Intelligence Index | 25.8 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 89.6 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 35.0 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | -10.2 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| AA-Omniscience Accuracy (Artificial Analysis Omniscience Accuracy) | 39.6 | Knowledge questions | Broad knowledge | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA-Omniscience Hallucination Rate (Artificial Analysis Omniscience Hallucination Rate) | 82.4 | Knowledge questions | Factuality | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| OpenHarmony Bench (OpenHarmony Bench v1.0) | 52.1 | 153 app-development and bug-fix tasks | End-to-end OpenHarmony application development | [OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development](https://arxiv.org/abs/2608.16022) |
| ProgramBench (ProgramBench: Can Language Models Rebuild Programs From Scratch?) | 53.6 | 200 program reconstruction tasks | Full-repository software architecture | [ProgramBench: Can Language Models Rebuild Programs From Scratch?](https://programbench.com/static/paper.pdf) |
| Kimi Code Bench v2 | 62.0 | Realistic coding-agent tasks | Production software engineering | [Kimi K2.7 Code](https://huggingface.co/moonshotai/Kimi-K2.7-Code) |
| MLS-Bench Lite | 35.1 | 30 machine-learning research tasks | ML research and systems engineering | [MLS-Bench](https://mls-bench.com/) |
| AA Coding Index (Artificial Analysis Coding Index) | 60.8 | Cross-benchmark coding index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| AA-SciCode (Artificial Analysis SciCode) | 47.8 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| LiveCodeBench (Vals) (LiveCodeBench, Vals AI run) | 82.1 | Competitive programming problems (easy, medium, hard) | Frontier coding | [Vals AI LiveCodeBench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/lcb) |
| SWE-bench (Vals) (SWE-bench, Vals AI run) | 78.2 | Real repository issues by human time bucket | Frontier coding agents | [Vals AI SWE-bench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/swebench) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA-LCR (Artificial Analysis Long Context Reasoning) | 79.3 | Long-context reasoning tasks | Long-context reasoning | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| CritPt (Critical Physics Tasks) | 10.0 | Research-level physics questions | Research-level physics reasoning | [CritPt Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/critpt) |

### Instruction Following

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA-IFBench (Artificial Analysis IFBench) | 63.1 | Verifiable instruction constraints | Instruction precision | [Artificial Analysis IFBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/ifbench) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| GDPval-AA | 1114 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA (GDPval-AA normalized) | 26.3 | Economically valuable tasks | Professional agentic workflows | [Artificial Analysis model benchmarks](https://artificialanalysis.ai/models/grok-4-3) |
| AA Agentic Index (Artificial Analysis Agentic Index) | 22.5 | Cross-benchmark agentic index | Display-only external reference | [Artificial Analysis model leaderboards](https://artificialanalysis.ai/leaderboards/models) |
| MCP Atlas | 76 | Tool-integrated agent tasks | Advanced tool use | [Introducing GPT-5.4 mini and nano](https://openai.com/index/introducing-gpt-5-4-mini-and-nano/) |
| Kimi Claw 24/7 (Kimi Claw 24/7 Bench) | 46.9 | 17 professional scenarios, 610 evaluation points | Long-horizon agentic work | [Kimi K2.7 Code](https://huggingface.co/moonshotai/Kimi-K2.7-Code) |
| MCP Mark Verified (MCPMark-Verified) | 81.1 | MCP tool-use tasks across five server environments | Advanced tool use | [MCPMark](https://mcpmark.ai/) |
| τ²-bench results (τ²-Bench Tool-Agent-User Evaluation) | 90.1 | Airline, retail, and telecom customer-service task sets | Dual-control customer-service workflows | [τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment](https://arxiv.org/abs/2506.07982) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 67.0 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Design Arena Website (Design Arena Website Elo) | 1278 | Website generation comparisons | Design and website generation | [OpenRouter Grok 4.3 benchmarks](https://openrouter.ai/x-ai/grok-4.3/benchmarks) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
