Models workbench

Price, capability, and service constraints in one working surface. Build a two-to-four model comparison from published records.

Benchmark evidence · Stale · checked Sep 14, 2026

Benchmark revision benchmark_2f59d4658a789000c14f0d0c7a177e62 · cache benchmark_2f59d4658a789000c14f0d0c7a177e62+cache-20260914055705000-c21dcf92-e210-43a5-bb34-f9c3e2271701 · pinned catalog catalog_e7f5de78d2108d49f3be10ea_62bbb2e3 · current catalog catalog_e7f5de78d2108d49f3be10ea_62bbb2e3.

Visible models
100
Frontier set
Loading…
Selection
0/4

Price–performance frontier

Composite quality against blended cost per one million tokens.

Loading the complete price–performance publication…

Quick comparison

Select two to four models from the catalog. The selection stays in the ordered models URL parameter across catalog views and filters.

Selected models
    0 / 4

    0 of 4 models selected.

    Catalog

    100 of 5442 matching source-scoped directory records · sorted by rank · profile links use canonical routes. Weekly rank is a benchmark directory order, not usage popularity.

    This is a bounded directory page. More published records are available. Use the page controls; search and access filters query the complete published directory. Your selected models remain in the URL.

    Anthropic · Proprietary

    Claude Mythos 5

    Score
    83.40
    Runtime
    Not reported
    Input / 1M
    $10
    Output / 1M
    $50
    Context
    1M
    Evidence
    supported
    Anthropic · Proprietary

    Claude Fable 5

    Score
    83.15
    Runtime
    Not reported
    Input / 1M
    $10
    Output / 1M
    $50
    Context
    1M
    Evidence
    supported
    Anthropic · Proprietary

    Claude Opus 5

    Score
    83.06
    Runtime
    Not reported
    Input / 1M
    $5
    Output / 1M
    $25
    Context
    1M
    Evidence
    supported
    OpenAI · Proprietary

    GPT-5.6 Sol

    Score
    82.20
    Runtime
    Not reported
    Input / 1M
    $5
    Output / 1M
    $30
    Context
    1.1M
    Evidence
    supported
    Moonshot AI · Unknown

    Kimi K3

    Score
    80.61
    Runtime
    Not reported
    Input / 1M
    $3
    Output / 1M
    $15
    Context
    1.1M
    Evidence
    supported
    Alibaba · Open weights

    Qwen3.8 Max

    Score
    79.22
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Tencent · Open weights

    Hy4 preview

    Score
    79.16
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Meta · Proprietary

    Muse Spark 1.1

    Score
    77.14
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Anthropic · Proprietary

    Claude Opus 4.8

    Score
    76.60
    Runtime
    Not reported
    Input / 1M
    $5
    Output / 1M
    $25
    Context
    1M
    Evidence
    supported
    Google · Proprietary

    Gemini 3.6 Flash

    Score
    75.66
    Runtime
    Not reported
    Input / 1M
    $1.5
    Output / 1M
    $7.5
    Context
    1M
    Evidence
    supported
    xAI · Proprietary

    Grok 4.5

    Score
    75.65
    Runtime
    Not reported
    Input / 1M
    $2
    Output / 1M
    $6
    Context
    500K
    Evidence
    supported
    OpenAI · Proprietary

    GPT-5.4

    Score
    73.56
    Runtime
    Not reported
    Input / 1M
    $2.5
    Output / 1M
    $15
    Context
    1.1M
    Evidence
    supported
    OpenAI · Proprietary

    GPT-5.6 Terra

    Score
    72.95
    Runtime
    Not reported
    Input / 1M
    $2.5
    Output / 1M
    $15
    Context
    1.1M
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.5

    Score
    72.92
    Runtime
    Not reported
    Input / 1M
    $5
    Output / 1M
    $30
    Context
    1M
    Evidence
    estimated
    Anthropic · Proprietary

    Claude Opus 4.7 (Adaptive)

    Score
    72.58
    Runtime
    Not reported
    Input / 1M
    $5
    Output / 1M
    $25
    Context
    1M
    Evidence
    estimated
    Alibaba · Open weights

    Qwen3.8-27B

    Score
    72.51
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Anthropic · Proprietary

    Claude Opus 4.7

    Score
    72.33
    Runtime
    Not reported
    Input / 1M
    $5
    Output / 1M
    $25
    Context
    1M
    Evidence
    supported
    Alibaba · Proprietary

    Qwen3.7 Max

    Score
    71.79
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Meta · Proprietary

    Muse Spark

    Score
    71.02
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Xiaomi · Proprietary

    MiMo-V2.5-Pro

    Score
    69.41
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    MiniMax · Open weights

    MiniMax M3

    Score
    68.73
    Runtime
    Not reported
    Input / 1M
    $0.3
    Output / 1M
    $1.2
    Context
    1M
    Evidence
    supported
    Dots Studio · Open weights

    dots3-note Preview

    Score
    68.66
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Ornith AI · Open weights

    Ornith-1.5-397B

    Score
    68.60
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Anthropic · Proprietary

    Claude Opus 4.6

    Score
    68.26
    Runtime
    Not reported
    Input / 1M
    $5
    Output / 1M
    $25
    Context
    1M
    Evidence
    supported
    Tencent · Open weights

    Hy3

    Score
    68.16
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Google · Proprietary

    Gemini 3 Pro

    Score
    67.64
    Runtime
    Not reported
    Input / 1M
    $2
    Output / 1M
    $12
    Context
    2M
    Evidence
    supported
    OpenAI · Proprietary

    GPT-5.6 Luna

    Score
    67.35
    Runtime
    Not reported
    Input / 1M
    $1
    Output / 1M
    $6
    Context
    1.1M
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.2 Pro

    Score
    67.30
    Runtime
    Not reported
    Input / 1M
    $21
    Output / 1M
    $168
    Context
    400K
    Evidence
    supported
    Xiaomi · Proprietary

    MiMo-V2-Pro

    Score
    67.15
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Z.AI · Open weights

    GLM-5.1

    Score
    67.04
    Runtime
    Not reported
    Input / 1M
    $1.4
    Output / 1M
    $4.4
    Context
    203K
    Evidence
    supported
    Thinking Machines Lab · Open weights

    Inkling

    Score
    67.02
    Runtime
    Not reported
    Input / 1M
    $1.87
    Output / 1M
    $4.68
    Context
    1M
    Evidence
    supported
    OpenAI · Proprietary

    GPT-5.4 nano

    Score
    66.78
    Runtime
    Not reported
    Input / 1M
    $0.2
    Output / 1M
    $1.25
    Context
    400K
    Evidence
    supported
    Z.AI · Proprietary

    GLM-5-Turbo

    Score
    66.30
    Runtime
    Not reported
    Input / 1M
    $1.2
    Output / 1M
    $4
    Context
    200K
    Evidence
    supported
    OpenAI · Proprietary

    GPT-5.3 Codex

    Score
    66.16
    Runtime
    Not reported
    Input / 1M
    $1.75
    Output / 1M
    $14
    Context
    400K
    Evidence
    supported
    Alibaba · Proprietary

    Qwen3.7 Plus

    Score
    65.93
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Z.AI · Open weights

    GLM-5

    Score
    65.89
    Runtime
    Not reported
    Input / 1M
    $1
    Output / 1M
    $3.2
    Context
    200K
    Evidence
    supported
    Google · Proprietary

    Gemini 3.5 Flash-Lite

    Score
    65.50
    Runtime
    Not reported
    Input / 1M
    $0.3
    Output / 1M
    $2.5
    Context
    1M
    Evidence
    supported
    Anthropic · Proprietary

    Claude Opus 4.6 (Adaptive)

    Score
    65.06
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Anthropic · Proprietary

    Claude Sonnet 5

    Score
    65.02
    Runtime
    Not reported
    Input / 1M
    $2
    Output / 1M
    $10
    Context
    1M
    Evidence
    estimated
    Alibaba · Proprietary

    Qwen3.6 Plus

    Score
    64.99
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Anthropic · Proprietary

    Claude Sonnet 4.6

    Score
    64.77
    Runtime
    Not reported
    Input / 1M
    $3
    Output / 1M
    $15
    Context
    200K
    Evidence
    supported
    Google · Proprietary

    Gemini 3.5 Flash

    Score
    64.73
    Runtime
    Not reported
    Input / 1M
    $1.5
    Output / 1M
    $9
    Context
    1M
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.5 Pro

    Score
    64.43
    Runtime
    Not reported
    Input / 1M
    $30
    Output / 1M
    $180
    Context
    1M
    Evidence
    estimated
    xAI · Proprietary

    Grok 4.3

    Score
    64.25
    Runtime
    Not reported
    Input / 1M
    $1.25
    Output / 1M
    $2.5
    Context
    1M
    Evidence
    supported
    Thinking Machines Lab · Open weights

    Inkling-Small

    Score
    64.02
    Runtime
    Not reported
    Input / 1M
    $0.58
    Output / 1M
    $1.44
    Context
    1M
    Evidence
    supported
    Anthropic · Proprietary

    Claude Opus 4.5

    Score
    63.91
    Runtime
    Not reported
    Input / 1M
    $5
    Output / 1M
    $25
    Context
    200K
    Evidence
    supported
    xAI · Proprietary

    Grok 4.6

    Score
    63.42
    Runtime
    Not reported
    Input / 1M
    $2
    Output / 1M
    $6
    Context
    500K
    Evidence
    estimated
    Z.AI · Open weights

    GLM-5.2

    Score
    63.40
    Runtime
    Not reported
    Input / 1M
    $1.4
    Output / 1M
    $4.4
    Context
    1M
    Evidence
    estimated
    MiniMax · Open weights

    MiniMax M2.7

    Score
    63.29
    Runtime
    Not reported
    Input / 1M
    $0.3
    Output / 1M
    $1.2
    Context
    200K
    Evidence
    supported
    Z.AI · Open weights

    GLM-5.3

    Score
    62.84
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Xiaomi · Proprietary

    MiMo-V2-Omni

    Score
    62.61
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Z.AI · Proprietary

    GLM-5V-Turbo

    Score
    62.53
    Runtime
    Not reported
    Input / 1M
    $1.2
    Output / 1M
    $4
    Context
    200K
    Evidence
    supported
    Google · Proprietary

    Gemini 3 Pro Deep Think

    Score
    62.13
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Meta · Proprietary

    Muse Spark 1.2

    Score
    61.71
    Runtime
    Not reported
    Input / 1M
    $1.25
    Output / 1M
    $4.25
    Context
    1M
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.4 Pro

    Score
    61.53
    Runtime
    Not reported
    Input / 1M
    $30
    Output / 1M
    $180
    Context
    1.1M
    Evidence
    estimated
    Google · Unknown

    Gemini 3.7 Flash

    Score
    61.40
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Alibaba · Open weights

    Qwen3.8-Flash-Next

    Score
    61.32
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Z.AI · Open weights

    GLM-5.3-Flash

    Score
    61.27
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    DeepSeek · Proprietary

    DeepSeek V4 Pro 0813

    Score
    61.02
    Runtime
    Not reported
    Input / 1M
    $0.435
    Output / 1M
    $0.87
    Context
    1M
    Evidence
    estimated
    InternScience · Open weights

    Agents-A1

    Score
    61.02
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    InternScience · Open weights

    Agents-A1-F16-GGUF

    Score
    61.02
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    InternScience · Open weights

    Agents-A1-FP8

    Score
    61.02
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    InternScience · Open weights

    Agents-A1-Q4_K_M-GGUF

    Score
    61.02
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    InternScience · Open weights

    Agents-A1-Q8_0-GGUF

    Score
    61.02
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Z.AI · Open weights

    GLM-4.7

    Score
    60.91
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    xAI · Proprietary

    Grok 4.1

    Score
    60.75
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Z.AI · Open weights

    GLM-5 (Reasoning)

    Score
    60.58
    Runtime
    Not reported
    Input / 1M
    $1
    Output / 1M
    $3.2
    Context
    200K
    Evidence
    estimated
    Alibaba · Proprietary

    Qwen 3.6 Max (preview)

    Score
    60.41
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Google · Open weights

    Gemma 4 31B

    Score
    60.39
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Alibaba · Open weights

    Qwen3.5 397B (Reasoning)

    Score
    60.30
    Runtime
    Not reported
    Input / 1M
    $0.6
    Output / 1M
    $3.6
    Context
    128K
    Evidence
    estimated
    Moonshot AI · Proprietary

    Kimi K2.5 (Reasoning)

    Score
    60.14
    Runtime
    Not reported
    Input / 1M
    $0.6
    Output / 1M
    $3
    Context
    256K
    Evidence
    estimated
    Moonshot AI · Open weights

    Kimi K2.6

    Score
    60.11
    Runtime
    Not reported
    Input / 1M
    $0.95
    Output / 1M
    $4
    Context
    256K
    Evidence
    estimated
    Google · Proprietary

    Gemini 3 Flash

    Score
    60.09
    Runtime
    Not reported
    Input / 1M
    $0.5
    Output / 1M
    $3
    Context
    1M
    Evidence
    supported
    Alibaba · Open weights

    Qwen3.5-27B

    Score
    60.04
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    xAI · Proprietary

    Grok 4.1 Fast (Reasoning)

    Score
    59.94
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Alibaba · Open weights

    Qwen3.5-122B-A10B

    Score
    59.89
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    xAI · Proprietary

    Grok 4

    Score
    59.84
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    OpenAI · Proprietary

    GPT-5.2 Instant

    Score
    59.80
    Runtime
    Not reported
    Input / 1M
    $1.5
    Output / 1M
    $6
    Context
    128K
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.3 Instant

    Score
    59.69
    Runtime
    Not reported
    Input / 1M
    $1.75
    Output / 1M
    $14
    Context
    128K
    Evidence
    estimated
    Xiaomi · Proprietary

    MiMo-V2.5

    Score
    59.48
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5 (high)

    Score
    59.42
    Runtime
    Not reported
    Input / 1M
    $1.25
    Output / 1M
    $10
    Context
    400K
    Evidence
    estimated
    Moonshot AI · Open weights

    Kimi K2.5

    Score
    59.14
    Runtime
    Not reported
    Input / 1M
    $0.6
    Output / 1M
    $3
    Context
    256K
    Evidence
    supported
    MiniMax · Proprietary

    MiniMax M2.5

    Score
    58.92
    Runtime
    Not reported
    Input / 1M
    $0.3
    Output / 1M
    $1.2
    Context
    128K
    Evidence
    supported
    DeepSeek · Open weights

    DeepSeek V3.2 (Thinking)

    Score
    58.91
    Runtime
    Not reported
    Input / 1M
    $0.55
    Output / 1M
    $2.19
    Context
    128K
    Evidence
    estimated
    Alibaba · Open weights

    Qwen3 235B 2507 (Reasoning)

    Score
    58.80
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.2

    Score
    58.39
    Runtime
    Not reported
    Input / 1M
    $1.75
    Output / 1M
    $14
    Context
    400K
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.2-Codex

    Score
    58.36
    Runtime
    Not reported
    Input / 1M
    $1.75
    Output / 1M
    $14
    Context
    400K
    Evidence
    supported
    Z.AI · Proprietary

    GLM-4.5

    Score
    58.35
    Runtime
    Not reported
    Input / 1M
    $0.6
    Output / 1M
    $2.2
    Context
    128K
    Evidence
    estimated
    Alibaba · Open weights

    Qwen3.5 397B

    Score
    57.78
    Runtime
    Not reported
    Input / 1M
    $0.6
    Output / 1M
    $3.6
    Context
    128K
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.3-Codex-Spark

    Score
    57.68
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Anthropic · Proprietary

    Claude Opus 4.5 Thinking

    Score
    57.56
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Google · Open weights

    Gemma 4 26B A4B

    Score
    57.45
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Anthropic · Proprietary

    Claude Haiku 4.5

    Score
    57.38
    Runtime
    Not reported
    Input / 1M
    $1
    Output / 1M
    $5
    Context
    200K
    Evidence
    estimated
    OpenAI · Proprietary

    GPT-5.4 mini

    Score
    56.96
    Runtime
    Not reported
    Input / 1M
    $0.75
    Output / 1M
    $4.5
    Context
    400K
    Evidence
    estimated
    Google · Proprietary

    Gemini 2.5 Pro

    Score
    56.87
    Runtime
    Not reported
    Input / 1M
    $1.25
    Output / 1M
    $10
    Context
    1M
    Evidence
    supported
    Alibaba · Open weights

    Qwen3 235B 2507

    Score
    56.77
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    estimated
    Arcee AI · Open weights

    Trinity-Large-Preview

    Score
    56.69
    Runtime
    Not reported
    Input / 1M
    $0.25
    Output / 1M
    $1
    Context
    512K
    Evidence
    estimated
    Alibaba · Open weights

    Qwen3.5-35B-A3B

    Score
    56.38
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported
    Google · Proprietary

    Gemini 3.1 Pro

    Score
    56.25
    Runtime
    Not reported
    Input / 1M
    $2
    Output / 1M
    $12
    Context
    1M
    Evidence
    estimated
    xAI · Proprietary

    Grok 4 Fast (Reasoning)

    Score
    56.11
    Runtime
    Not reported
    Input / 1M
    Not reported
    Output / 1M
    Not reported
    Context
    Not reported
    Evidence
    supported

    Lifecycle & retirement watch

    Lifecycle notices and provider retirement dates are not reported by the published directory.

    No verified retirement notices are available. Do not infer a sunset date from absence or record age.

    Recent directory observations

    First-seen timestamps are directory observations, not vendor release announcements.

    1. Hy4 previewFirst observed in the published directory.
    2. Qwen3.8-Flash-NextFirst observed in the published directory.
    3. GLM-5.3-FlashFirst observed in the published directory.
    4. Qwen3.8-27BFirst observed in the published directory.