Skip to main content
ModelScale

Multimodal & Grounded benchmark

Design Arena Website leaderboard

Design Arena Website Elo. Every model the catalog carries a published Design Arena Website value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureElo
TasksWebsite generation comparisons
DifficultyDesign and website generation

A display-only Design Arena website-generation Elo score surfaced on OpenRouter model benchmark pages.

Design Arena Website ranking

88 models with a published Design Arena Website value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Design Arena Website Elo value
RankModelProviderElo
1Muse Spark 1.3Meta1364
2Kimi K3Moonshot AI1353
3Claude Fable 5.1Anthropic1322
4Muse Spark 1.2Meta1321
5Claude Opus 5Anthropic1319
6GLM-5.3Z.AI1316
7Gemini 3.7 FlashGoogle1315
7Gemini 3.8 FlashGoogle1315
9Gemini 3.6 FlashGoogle1313
10Claude Fable 5Anthropic1308
11GLM-5.2Z.AI1305
12Grok 4.6xAI1304
13Claude Opus 4.6Anthropic1303
13Claude Opus 4.6 (Adaptive)Anthropic1303
15Claude Sonnet 4.6Anthropic1297
16Grok 4.5xAI1296
17GLM-5.1Z.AI1290
18Claude Sonnet 5Anthropic1289
19GLM-5.3-FlashZ.AI1285
19Qwen3.7 MaxAlibaba1285
21MiMo-V2.5-ProXiaomi1284
22Kimi K2.6Moonshot AI1282
22Qwen3.7 PlusAlibaba1282
24Muse Spark 1.1Meta1281
25Kimi K2.7 CodeMoonshot AI1280
26MiMo-V2.5Xiaomi1279
27GLM-5-TurboZ.AI1278
28Gemini 3.5 FlashGoogle1275
29MiniMax M3MiniMax1271
30GPT-5.5OpenAI1269
31Claude Opus 4.8Anthropic1267
32Gemini 3.1 ProGoogle1265
33Kimi K2.5Moonshot AI1261
33Kimi K2.5 (Reasoning)Moonshot AI1261
35GLM-5Z.AI1260
35GLM-5 (Reasoning)Z.AI1260
37Claude Opus 4.5Anthropic1259
37Claude Opus 4.5 ThinkingAnthropic1259
39DeepSeek V4 Pro 0813DeepSeek1258
40MiniMax M2.7MiniMax1257
41Qwen3.6 PlusAlibaba1253
42GLM-5V-TurboZ.AI1243
43Grok 4.20xAI1242
44GLM-4.7Z.AI1238
45GPT-5.4OpenAI1232
46TMInklingThinking Machines Lab1230
47DeepSeek V4 Flash 0731DeepSeek1220
48Gemini 3 FlashGoogle1208
49GPT-5.2OpenAI1207
49Grok 4.3xAI1207
49Step 3.7 FlashStepFun1207
52Claude Sonnet 4.5Anthropic1202
52Claude Sonnet 4.5 ThinkingAnthropic1202
54GPT-5.1OpenAI1199
55GPT-5 (high)OpenAI1197
55GPT-5 (medium)OpenAI1197
57Hy3Tencent1194
58Claude 4.1 OpusAnthropic1189
59Solar Pro 4Upstage1188
60DeepSeek V3.2DeepSeek1187
60DeepSeek V3.2 (Thinking)DeepSeek1187
62GLM-4.5Z.AI1182
63Gemini 2.5 ProGoogle1179
64GPT-5.3 CodexOpenAI1175
65GPT-5.1-CodexOpenAI1174
66GLM-4.5-AirZ.AI1159
67Claude 4 SonnetAnthropic1158
68Trinity-Large-PreviewArcee AI1149
68Trinity-Large-ThinkingArcee AI1149
70Nemotron 3 UltraNVIDIA1146
71Claude Haiku 4.5Anthropic1135
71Claude Haiku 4.5 ThinkingAnthropic1135
71DeepSeek V3.1DeepSeek1135
71DeepSeek V3.1 (Reasoning)DeepSeek1135
75DeepSeek V3DeepSeek1132
76Qwen3 MaxAlibaba1131
77Gemini 2.5 FlashGoogle1127
78Mistral Medium 3Mistral1091
79Kimi K2Moonshot AI1063
80GPT-4.1OpenAI1051
81o3OpenAI1048
82GPT-4.1 miniOpenAI1010
83GPT-4.1 nanoOpenAI985
84GPT-OSS 120BOpenAI980
85Llama 4 MaverickMeta883
86GPT-OSS 20BOpenAI865
87GPT-4oOpenAI843
88Llama 4 ScoutMeta762

Evidence key: Observed

Rows are ordered by the value OpenRouter Grok 4.3 benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards