This is a bounded directory page. More published records are available. Use the page controls; search and access filters query the complete published directory. Your selected models remain in the URL.
Anthropic · Proprietary Compare
Score 83.40
Runtime Not reported
Input / 1M $10
Output / 1M $50
Context 1M
Evidence supported Anthropic · Proprietary Compare
Score 83.15
Runtime Not reported
Input / 1M $10
Output / 1M $50
Context 1M
Evidence supported Anthropic · Proprietary Compare
Score 83.06
Runtime Not reported
Input / 1M $5
Output / 1M $25
Context 1M
Evidence supported OpenAI · Proprietary Compare
Score 82.20
Runtime Not reported
Input / 1M $5
Output / 1M $30
Context 1.1M
Evidence supported Moonshot AI · Unknown Compare
Score 80.61
Runtime Not reported
Input / 1M $3
Output / 1M $15
Context 1.1M
Evidence supported Alibaba · Open weights Compare
Score 79.22
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Tencent · Open weights Compare
Score 79.16
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Meta · Proprietary Compare
Score 77.14
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Anthropic · Proprietary Compare
Score 76.60
Runtime Not reported
Input / 1M $5
Output / 1M $25
Context 1M
Evidence supported Google · Proprietary Compare
Score 75.66
Runtime Not reported
Input / 1M $1.5
Output / 1M $7.5
Context 1M
Evidence supported xAI · Proprietary Compare
Score 75.65
Runtime Not reported
Input / 1M $2
Output / 1M $6
Context 500K
Evidence supported OpenAI · Proprietary Compare
Score 73.56
Runtime Not reported
Input / 1M $2.5
Output / 1M $15
Context 1.1M
Evidence supported OpenAI · Proprietary Compare
Score 72.95
Runtime Not reported
Input / 1M $2.5
Output / 1M $15
Context 1.1M
Evidence estimated OpenAI · Proprietary Compare
Score 72.92
Runtime Not reported
Input / 1M $5
Output / 1M $30
Context 1M
Evidence estimated Anthropic · Proprietary Compare
Score 72.58
Runtime Not reported
Input / 1M $5
Output / 1M $25
Context 1M
Evidence estimated Alibaba · Open weights Compare
Score 72.51
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Anthropic · Proprietary Compare
Score 72.33
Runtime Not reported
Input / 1M $5
Output / 1M $25
Context 1M
Evidence supported Alibaba · Proprietary Compare
Score 71.79
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Meta · Proprietary Compare
Score 71.02
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Xiaomi · Proprietary Compare
Score 69.41
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported MiniMax · Open weights Compare
Score 68.73
Runtime Not reported
Input / 1M $0.3
Output / 1M $1.2
Context 1M
Evidence supported Dots Studio · Open weights Compare
Score 68.66
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Ornith AI · Open weights Compare
Score 68.60
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Anthropic · Proprietary Compare
Score 68.26
Runtime Not reported
Input / 1M $5
Output / 1M $25
Context 1M
Evidence supported Tencent · Open weights Compare
Score 68.16
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Google · Proprietary Compare
Score 67.64
Runtime Not reported
Input / 1M $2
Output / 1M $12
Context 2M
Evidence supported OpenAI · Proprietary Compare
Score 67.35
Runtime Not reported
Input / 1M $1
Output / 1M $6
Context 1.1M
Evidence estimated OpenAI · Proprietary Compare
Score 67.30
Runtime Not reported
Input / 1M $21
Output / 1M $168
Context 400K
Evidence supported Xiaomi · Proprietary Compare
Score 67.15
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Z.AI · Open weights Compare
Score 67.04
Runtime Not reported
Input / 1M $1.4
Output / 1M $4.4
Context 203K
Evidence supported Thinking Machines Lab · Open weights Compare
Score 67.02
Runtime Not reported
Input / 1M $1.87
Output / 1M $4.68
Context 1M
Evidence supported OpenAI · Proprietary Compare
Score 66.78
Runtime Not reported
Input / 1M $0.2
Output / 1M $1.25
Context 400K
Evidence supported Z.AI · Proprietary Compare
Score 66.30
Runtime Not reported
Input / 1M $1.2
Output / 1M $4
Context 200K
Evidence supported OpenAI · Proprietary Compare
Score 66.16
Runtime Not reported
Input / 1M $1.75
Output / 1M $14
Context 400K
Evidence supported Alibaba · Proprietary Compare
Score 65.93
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Z.AI · Open weights Compare
Score 65.89
Runtime Not reported
Input / 1M $1
Output / 1M $3.2
Context 200K
Evidence supported Google · Proprietary Compare
Score 65.50
Runtime Not reported
Input / 1M $0.3
Output / 1M $2.5
Context 1M
Evidence supported Anthropic · Proprietary Compare
Score 65.06
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Anthropic · Proprietary Compare
Score 65.02
Runtime Not reported
Input / 1M $2
Output / 1M $10
Context 1M
Evidence estimated Alibaba · Proprietary Compare
Score 64.99
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Anthropic · Proprietary Compare
Score 64.77
Runtime Not reported
Input / 1M $3
Output / 1M $15
Context 200K
Evidence supported Google · Proprietary Compare
Score 64.73
Runtime Not reported
Input / 1M $1.5
Output / 1M $9
Context 1M
Evidence estimated OpenAI · Proprietary Compare
Score 64.43
Runtime Not reported
Input / 1M $30
Output / 1M $180
Context 1M
Evidence estimated xAI · Proprietary Compare
Score 64.25
Runtime Not reported
Input / 1M $1.25
Output / 1M $2.5
Context 1M
Evidence supported Thinking Machines Lab · Open weights Compare
Score 64.02
Runtime Not reported
Input / 1M $0.58
Output / 1M $1.44
Context 1M
Evidence supported Anthropic · Proprietary Compare
Score 63.91
Runtime Not reported
Input / 1M $5
Output / 1M $25
Context 200K
Evidence supported xAI · Proprietary Compare
Score 63.42
Runtime Not reported
Input / 1M $2
Output / 1M $6
Context 500K
Evidence estimated Z.AI · Open weights Compare
Score 63.40
Runtime Not reported
Input / 1M $1.4
Output / 1M $4.4
Context 1M
Evidence estimated MiniMax · Open weights Compare
Score 63.29
Runtime Not reported
Input / 1M $0.3
Output / 1M $1.2
Context 200K
Evidence supported Z.AI · Open weights Compare
Score 62.84
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Xiaomi · Proprietary Compare
Score 62.61
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Z.AI · Proprietary Compare
Score 62.53
Runtime Not reported
Input / 1M $1.2
Output / 1M $4
Context 200K
Evidence supported Google · Proprietary Compare
Score 62.13
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Meta · Proprietary Compare
Score 61.71
Runtime Not reported
Input / 1M $1.25
Output / 1M $4.25
Context 1M
Evidence estimated OpenAI · Proprietary Compare
Score 61.53
Runtime Not reported
Input / 1M $30
Output / 1M $180
Context 1.1M
Evidence estimated Google · Unknown Compare
Score 61.40
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Alibaba · Open weights Compare
Score 61.32
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Z.AI · Open weights Compare
Score 61.27
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated DeepSeek · Proprietary Compare
Score 61.02
Runtime Not reported
Input / 1M $0.435
Output / 1M $0.87
Context 1M
Evidence estimated InternScience · Open weights Compare
Score 61.02
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated InternScience · Open weights Compare
Score 61.02
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated InternScience · Open weights Compare
Score 61.02
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated InternScience · Open weights Compare
Score 61.02
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated InternScience · Open weights Compare
Score 61.02
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Z.AI · Open weights Compare
Score 60.91
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported xAI · Proprietary Compare
Score 60.75
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Z.AI · Open weights Compare
Score 60.58
Runtime Not reported
Input / 1M $1
Output / 1M $3.2
Context 200K
Evidence estimated Alibaba · Proprietary Compare
Score 60.41
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Google · Open weights Compare
Score 60.39
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Alibaba · Open weights Compare
Score 60.30
Runtime Not reported
Input / 1M $0.6
Output / 1M $3.6
Context 128K
Evidence estimated Moonshot AI · Proprietary Compare
Score 60.14
Runtime Not reported
Input / 1M $0.6
Output / 1M $3
Context 256K
Evidence estimated Moonshot AI · Open weights Compare
Score 60.11
Runtime Not reported
Input / 1M $0.95
Output / 1M $4
Context 256K
Evidence estimated Google · Proprietary Compare
Score 60.09
Runtime Not reported
Input / 1M $0.5
Output / 1M $3
Context 1M
Evidence supported Alibaba · Open weights Compare
Score 60.04
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported xAI · Proprietary Compare
Score 59.94
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Alibaba · Open weights Compare
Score 59.89
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported xAI · Proprietary Compare
Score 59.84
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported OpenAI · Proprietary Compare
Score 59.80
Runtime Not reported
Input / 1M $1.5
Output / 1M $6
Context 128K
Evidence estimated OpenAI · Proprietary Compare
Score 59.69
Runtime Not reported
Input / 1M $1.75
Output / 1M $14
Context 128K
Evidence estimated Xiaomi · Proprietary Compare
Score 59.48
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated OpenAI · Proprietary Compare
Score 59.42
Runtime Not reported
Input / 1M $1.25
Output / 1M $10
Context 400K
Evidence estimated Moonshot AI · Open weights Compare
Score 59.14
Runtime Not reported
Input / 1M $0.6
Output / 1M $3
Context 256K
Evidence supported MiniMax · Proprietary Compare
Score 58.92
Runtime Not reported
Input / 1M $0.3
Output / 1M $1.2
Context 128K
Evidence supported DeepSeek · Open weights Compare
Score 58.91
Runtime Not reported
Input / 1M $0.55
Output / 1M $2.19
Context 128K
Evidence estimated Alibaba · Open weights Compare
Score 58.80
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated OpenAI · Proprietary Compare
Score 58.39
Runtime Not reported
Input / 1M $1.75
Output / 1M $14
Context 400K
Evidence estimated OpenAI · Proprietary Compare
Score 58.36
Runtime Not reported
Input / 1M $1.75
Output / 1M $14
Context 400K
Evidence supported Z.AI · Proprietary Compare
Score 58.35
Runtime Not reported
Input / 1M $0.6
Output / 1M $2.2
Context 128K
Evidence estimated Alibaba · Open weights Compare
Score 57.78
Runtime Not reported
Input / 1M $0.6
Output / 1M $3.6
Context 128K
Evidence estimated OpenAI · Proprietary Compare
Score 57.68
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Anthropic · Proprietary Compare
Score 57.56
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Google · Open weights Compare
Score 57.45
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Anthropic · Proprietary Compare
Score 57.38
Runtime Not reported
Input / 1M $1
Output / 1M $5
Context 200K
Evidence estimated OpenAI · Proprietary Compare
Score 56.96
Runtime Not reported
Input / 1M $0.75
Output / 1M $4.5
Context 400K
Evidence estimated Google · Proprietary Compare
Score 56.87
Runtime Not reported
Input / 1M $1.25
Output / 1M $10
Context 1M
Evidence supported Alibaba · Open weights Compare
Score 56.77
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence estimated Arcee AI · Open weights Compare
Score 56.69
Runtime Not reported
Input / 1M $0.25
Output / 1M $1
Context 512K
Evidence estimated Alibaba · Open weights Compare
Score 56.38
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported Google · Proprietary Compare
Score 56.25
Runtime Not reported
Input / 1M $2
Output / 1M $12
Context 1M
Evidence estimated xAI · Proprietary Compare
Score 56.11
Runtime Not reported
Input / 1M Not reported
Output / 1M Not reported
Context Not reported
Evidence supported