Skip to main content
ModelScale

Agentic benchmark

AA Tau3 Banking leaderboard

Artificial Analysis Tau3-Banking. Every model the catalog carries a published AA Tau3 Banking value for, ranked by that value.

CategoryAgentic
MeasureTask success rate
TasksBanking tool-use workflows
DifficultyAgentic banking workflows

An independently evaluated Tau3 banking benchmark from Artificial Analysis.

AA Tau3 Banking ranking

15 models with a published AA Tau3 Banking value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Tau3-Banking value
RankModelProviderTask success rate
1Grok 4.6xAI50.7
2Muse Spark 1.3Meta50.5
3GLM-5.3Z.AI50.3
4Qwen3.8-27BAlibaba48.0
5Claude Fable 5.1Anthropic47.2
5GLM-5.3-FlashZ.AI47.2
7Kimi K3Moonshot AI46.0
8Gemini 3.8 FlashGoogle44.9
9GPT-5.6 SolOpenAI44.3
10Claude Opus 5Anthropic42.1
11GPT-6 AstraOpenAI41.4
12GPT-5.6 TerraOpenAI40.2
13DeepSeek V4 Pro 0813DeepSeek39.6
14Claude Fable 5Anthropic38.1
15Ling 3.0 FlashInclusionAI28.0

Evidence key: Observed

Rows are ordered by the value Artificial Analysis Tau3-Banking Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard