Skip to main content
ModelScale

Coding benchmark

OpenHarmony Bench leaderboard

OpenHarmony Bench v1.0. Every model the catalog carries a published OpenHarmony Bench value for, ranked by that value.

CategoryCoding
MeasureTask completion through DevEco Code
Tasks153 app-development and bug-fix tasks
DifficultyEnd-to-end OpenHarmony application development

An app-level coding benchmark that asks DevEco Code configurations to implement observable behavior in buildable OpenHarmony ArkTS applications.

OpenHarmony Bench ranking

12 models with a published OpenHarmony Bench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published OpenHarmony Bench v1.0 value
RankModelProviderTask completion through DevEco Code
1GLM-5.3Z.AI60.8
1Qwen3.8 MaxAlibaba60.8
3DeepSeek V4 Pro 0813DeepSeek59.0
4GLM-5.2Z.AI58.4
5GLM-5.3-FlashZ.AI57.3
5Kimi K3Moonshot AI57.3
7Qwen3.8 Max PreviewAlibaba56.0
8DeepSeek V4 Flash 0731DeepSeek53.8
9Qwen3.7 MaxAlibaba53.4
10GLM-5.1Z.AI52.3
11Kimi K2.7 CodeMoonshot AI52.1
12MiniMax M3MiniMax48.4

Evidence key: Observed

Rows are ordered by the value OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard