# GLM-5.3-Flash

Z.AI · Open Weight · rank 32 · bench-align-v5 · Self-hosted

> Every figure below is reproduced as its upstream source published it: nothing is modelled, estimated, interpolated or converted. `Unavailable` means no source published the value — it is never a zero. Each value carries its evidence state and the date it was observed.

Page: https://modelscale.dev/models/glm-5-3-flash  
JSON: https://modelscale.dev/api/model/glm-5-3-flash

## Facts

| Field | Value | Evidence |
| --- | --- | --- |
| Canonical id | `glm-5-3-flash` | — |
| Overall score | 66.19 | Observed 2026-09-22 · source benchlm:models |
| Context window | 1M tokens | Observed 2026-09-22 · source benchlm:models |
| Release date | 2026-08-26 | Observed 2026-09-22 · source benchlm:models |
| Access type | Open Weight | — |
| Blended $/1M (75% input / 25% output) | Unavailable | Derived from the input and output rates below |

## Capability evidence

Seven axes from the ranking source. An axis the source did not score is unavailable, not zero.

| Axis | Score | Evidence |
| --- | --- | --- |
| Agentic | 80.6 | Observed 2026-09-22 · source benchlm:models |
| Coding | 56.4 | Observed 2026-09-22 · source benchlm:models |
| Knowledge | 67.5 | Observed 2026-09-22 · source benchlm:models |
| Reasoning | Unavailable | Unavailable · source benchlm:models |
| Multimodal & Grounded | 80.4 | Observed 2026-09-22 · source benchlm:models |
| Instruction Following | Unavailable | Unavailable · source benchlm:models |
| Math | Unavailable | Unavailable · source benchlm:models |

## Runtime service evidence

Measured values with the date they were observed. Nothing is inferred from a sibling model or a provider claim. Regional or per-endpoint measurements appear only when the API supplies them; none are modelled.

| Measurement | Value | Observed | Last good | Evidence |
| --- | --- | --- | --- | --- |
| Time to first token | Unavailable | Unavailable | Unavailable | Unavailable · source benchlm:speed |
| Throughput | Unavailable | Unavailable | Unavailable | Unavailable · source benchlm:speed |

## Endpoint and price matrix

Every published price component, including cache reads and writes.

| Component | USD | Evidence |
| --- | --- | --- |
| Input / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Output / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache read / 1M tokens | Unavailable | Unavailable · source benchlm:pricing |
| Cache write / 1M tokens | Unavailable | Unavailable · source openrouter:pricing |
| Blended / 1M (75% input / 25% output) | Unavailable | Derived — from the input and output rates above; it has no source record of its own |
| Cost per successful task (LiveBench) | Unavailable | Unavailable · source livebench:table |

**Self-hosted listing.** Z.AI publishes GLM-5.3-Flash under MIT for self-hosting, so BenchLM represents the open checkpoint as self-host/free-per-token before infrastructure costs. The exact glm-5.3-flash model is also available through Z.AI services and the GLM Coding Plan, but Z.AI's public per-million-token pricing table does not yet list this SKU. The hosted API must not be read as free from this self-host placeholder. No hosted token rate was published for this model, so its per-token price is unavailable rather than zero.

## Workload-aware monthly cost example

10 conversations per day × 8 messages × 22 active days, 1200 input and 400 output tokens per message, no cache. Derived here from the published rates above by this site's own calculator — not a figure any source published.

| Field | Value |
| --- | --- |
| Modelled monthly cost | Unavailable |
| Modelled tokens | Unavailable |
| Reason | The applicable input rate is unavailable. |

## Benchmark record

33 matched benchmark rows with their published value, unit, and provenance.

### Knowledge

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Artificial Analysis Intelligence Index | 41.8 | Cross-benchmark intelligence index | Display-only external reference | [Artificial Analysis](https://artificialanalysis.ai/) |
| AA-GPQA Diamond (Artificial Analysis GPQA Diamond) | 91.2 | Graduate-level science questions | Graduate-level science reasoning | [Artificial Analysis GPQA Diamond Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gpqa-diamond) |
| AA-HLE (Artificial Analysis Humanity's Last Exam) | 39.9 | Expert-level questions | Frontier expert reasoning | [Artificial Analysis Humanity's Last Exam Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/hle) |
| AA-Omniscience Index (Artificial Analysis Omniscience Index) | 7.5 | Knowledge questions | Broad factual knowledge | [AA-Omniscience: Knowledge and Hallucination Benchmark](https://artificialanalysis.ai/evaluations/omniscience) |
| GPQA Diamond (Vals) (GPQA Diamond, Vals AI run) | 86.4 | Graduate-level science questions | Expert reasoning | [Vals AI GPQA Diamond, Vals AI run leaderboard](https://www.vals.ai/benchmarks/gpqa) |
| MMLU-Pro (Vals) (MMLU-Pro, Vals AI run) | 86.1 | Academic multiple-choice questions | Broad academic knowledge | [Vals AI MMLU-Pro, Vals AI run leaderboard](https://www.vals.ai/benchmarks/mmlu_pro) |

### Coding

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Terminal-Bench 2.1 (Terminal-Bench 2.1 (provider run)) | 84.3 | Terminal-based software-agent tasks | Professional software engineering | [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/) |
| OpenHarmony Bench (OpenHarmony Bench v1.0) | 57.3 | 153 app-development and bug-fix tasks | End-to-end OpenHarmony application development | [OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development](https://arxiv.org/abs/2608.16022) |
| NL2Repo | 56.3 | Natural language to repository tasks | System-level software comprehension | [MiniMax M2.7: Early Echoes of Self-Evolution](https://www.minimax.io/news/minimax-m27-en) |
| AA-SciCode (Artificial Analysis SciCode) | 51.6 | Scientific coding subproblems | Scientific programming | [Artificial Analysis SciCode Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/scicode) |
| LiveCodeBench (Vals) (LiveCodeBench, Vals AI run) | 80.5 | Competitive programming problems (easy, medium, hard) | Frontier coding | [Vals AI LiveCodeBench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/lcb) |
| SWE-bench (Vals) (SWE-bench, Vals AI run) | 92.0 | Real repository issues by human time bucket | Frontier coding agents | [Vals AI SWE-bench, Vals AI run leaderboard](https://www.vals.ai/benchmarks/swebench) |
| DeepSWE | 63.4 | 113 software engineering tasks across 91 repositories and 5 languages | Long-horizon software engineering | [DeepSWE benchmark blog](https://deepswe.datacurve.ai/blog) |

### Reasoning

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| MLCR-AA (Medical Long Context Reasoning (MLCR-AA)) | 51.1 | Long, fragmented medical-record reasoning | Long-context medical reasoning | [Medical Long Context Reasoning (MLCR-AA)](https://artificialanalysis.ai/evaluations/mlcr-aa) |

### Agentic

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| AA Briefcase (Artificial Analysis Briefcase) | 1459 | Professional knowledge-work tasks | Professional work | [Artificial Analysis Briefcase Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/aa-briefcase) |
| AA AutomationBench (Artificial Analysis AutomationBench) | 60.4 | Business-process automation tasks | Agentic automation | [Artificial Analysis AutomationBench Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/automationbench-aa) |
| AA EnterpriseOps-Gym (Artificial Analysis EnterpriseOps-Gym) | 33.2 | Enterprise operations workflows | Enterprise agent operations | [Artificial Analysis EnterpriseOps-Gym Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/enterprise-ops-gym-aa) |
| AA Tau3 Banking (Artificial Analysis Tau3-Banking) | 47.2 | Banking tool-use workflows | Agentic banking workflows | [Artificial Analysis Tau3-Banking Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/tau3-banking) |
| Terminal-Bench 2.1 (Terminal-Bench 2.1 (provider run)) | 84.3 | Terminal-based software-agent tasks | Professional software engineering | [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/) |
| HLE w/ tools (Humanity's Last Exam with tools) | 55.3 | Expert questions with tool use | Frontier tool-augmented reasoning | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| GDPval-AA | 1773 | Agentic real-world work tasks | Professional agentic workflows | [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf) |
| AA Terminal-Bench 4.0 (Artificial Analysis Terminal-Bench v4.0) | 32.8 | Terminal-based agent tasks | Agentic software engineering | [Artificial Analysis Terminal-Bench v4.0 Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/terminalbench-v4-0) |
| GDP.pdf (Artificial Analysis GDP.pdf) | 15.4 | Professional document-production tasks | Professional knowledge work | [GDP.pdf Benchmark Leaderboard](https://artificialanalysis.ai/evaluations/gdp-pdf) |
| Toolathlon-Verified | 78.4 | Verified multi-tool workflows | Advanced tool use | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| AutomationBench | 48.8 | 600 public automation tasks | Long-horizon automation | [Kimi K3: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) |
| Agents' Last Exam | 26.3 | Agent tasks | Advanced agentic work | [DeepSeek V4 Flash 0731 update](https://api-docs.deepseek.com/zh-cn/updates/) |
| Terminal-Bench 2.1 (Vals) (Terminal-Bench 2.1, Vals AI run) | 62.9 | Difficult terminal tasks | Frontier agentic | [Vals AI Terminal-Bench 2.1, Vals AI run leaderboard](https://www.vals.ai/benchmarks/terminal-bench-2-1) |

### Multimodal & Grounded

| Benchmark | Value | Tasks | Difficulty | Provenance |
| --- | --- | --- | --- | --- |
| Chartography (tools) (Chartography with image and code tools) | 78.0 | 100 specialized chart types | Professional chart reasoning | [Chartography](https://surgehq.ai/blog/chartography) |
| Design Arena Website (Design Arena Website Elo) | 1283 | Website generation comparisons | Design and website generation | [OpenRouter Grok 4.3 benchmarks](https://openrouter.ai/x-ai/grok-4.3/benchmarks) |
| OfficeQA Pro | 62.4 | Document and spreadsheet tasks | Enterprise grounded reasoning | [OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning](https://arxiv.org/abs/2603.08655) |
| MMVU (Multimodal Multi-disciplinary Video Understanding) | 80.5 | Video understanding | Multi-disciplinary multimodal video reasoning | [Kimi K2.5 benchmark release surface](https://www.kimi.com/blog/kimi-k2-5.html) |
| CharXiv (CharXiv Reasoning) | 89.4 | Scientific chart reasoning | Scientific visualization reasoning | [CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs](https://charxiv.github.io/) |
| BabyVision | 53.4 | Visual perception tasks | Fine-grained visual perception | [Muse Spark 1.1 Evaluation Report](https://ai.meta.com/static-resource/muse-spark-1-1-evaluation-report) |

## Lifecycle and limitations log

Lifecycle events the source associates with this model.

No lifecycle event references this model. That is not evidence the model has no lifecycle plan — only that this source published none.

## What this profile does not claim

Values are reproduced exactly as their sources published them, in the units those sources declared; none are converted, interpolated, or averaged across providers. Any field marked unavailable was attempted and not returned. Last attempted fetch for this model's score: 2026-09-22 10:17 UTC.
