Skip to main content
ModelScale

Agentic benchmark

Kimi Claw 24/7 leaderboard

Kimi Claw 24/7 Bench. Every model the catalog carries a published Kimi Claw 24/7 value for, ranked by that value.

CategoryAgentic
MeasureAverage pass rate across repeated OpenClaw runs
Tasks17 professional scenarios, 610 evaluation points
DifficultyLong-horizon agentic work
Published byKimi K2.7 Code

A Moonshot AI internal long-horizon agent benchmark for persistent professional coworking tasks.

Kimi Claw 24/7 ranking

1 model with a published Kimi Claw 24/7 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Kimi Claw 24/7 Bench value
RankModelProviderAverage pass rate across repeated OpenClaw runs
1Kimi K2.7 CodeMoonshot AI46.9

Evidence key: Observed

Rows are ordered by the value Kimi K2.7 Code published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard