Coding benchmark
OpenHarmony Bench leaderboard
OpenHarmony Bench v1.0. Every model the catalog carries a published OpenHarmony Bench value for, ranked by that value.
An app-level coding benchmark that asks DevEco Code configurations to implement observable behavior in buildable OpenHarmony ArkTS applications.
OpenHarmony Bench ranking
12 models with a published OpenHarmony Bench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Task completion through DevEco Code |
|---|---|---|---|
| 1 | Z.AI | 60.8 | |
| 1 | Alibaba | 60.8 | |
| 3 | DeepSeek | 59.0 | |
| 4 | Z.AI | 58.4 | |
| 5 | Z.AI | 57.3 | |
| 5 | Moonshot AI | 57.3 | |
| 7 | Alibaba | 56.0 | |
| 8 | DeepSeek | 53.8 | |
| 9 | Alibaba | 53.4 | |
| 10 | Z.AI | 52.3 | |
| 11 | Moonshot AI | 52.1 | |
| 12 | MiniMax | 48.4 |
Evidence key: Observed
Rows are ordered by the value OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.