A#

Agent ranking · Coding Agent Performance

Top 10 coding-agent configurations

Snapshot evaluated Jul 18, 2026

A July 18, 2026 snapshot of the ten highest-scoring configurations exposed by the current Artificial Analysis Coding Agent Index. The harness, model, reasoning setting, and fallback behavior are part of each result.

Evaluation date
ScopeCoding Agent Performance
Entries10
Independent resultsOpen source ↗

Dated leaderboard

Ranked configurations

Source retrieved Jul 18, 2026

RankAgentEvaluated configurationScoreEvidence
#1 Codex OpenAI GPT-5.6 Sol · max 61.014Coding Agent Index v1.2 The captured v1.2 leaderboard data places this Codex configuration first at 61.014.[1]
#2 Claude Code Anthropic Claude Fable 5 · max · with fallback 59.216Coding Agent Index v1.2 The captured v1.2 leaderboard data places this Claude Code configuration second at 59.216.[1]
#3 Grok Build SpaceXAI Grok 4.5 · high 57.898Coding Agent Index v1.2 The captured v1.2 leaderboard data places this Grok Build configuration third at 57.898.[1]
#4 Kimi Code Moonshot AI Kimi K3 56.765Coding Agent Index v1.2 The captured v1.2 leaderboard data places Kimi Code CLI with Kimi K3 fourth at 56.765.[1]
#5 Claude Code Anthropic Claude Opus 4.8 · max 54.896Coding Agent Index v1.2 The captured v1.2 leaderboard data places this Claude Code configuration fifth at 54.896.[1]
#6 OpenCode OpenCode Muse Spark 1.1 · xhigh 48.703Coding Agent Index v1.2 The captured v1.2 leaderboard data places this OpenCode configuration sixth at 48.703.[1]
#7 Claude Code Anthropic GLM-5.2 39.907Coding Agent Index v1.2 The captured v1.2 leaderboard data places this Claude Code configuration seventh at 39.907.[1]
#8 Cursor CLI Cursor Composer 2.5 Fast 33.757Coding Agent Index v1.2 The captured v1.2 leaderboard data places Cursor CLI with Composer 2.5 Fast eighth at 33.757.[1]
#9 Gemini CLI Google Gemini 3.1 Pro · high 29.174Coding Agent Index v1.2 The captured v1.2 leaderboard data places this Gemini CLI configuration ninth at 29.174.[1]
#10 Claude Code Anthropic DeepSeek V4 Pro · high 28.483Coding Agent Index v1.2 The captured v1.2 leaderboard data places this Claude Code configuration tenth at 28.483.[1]

How to read this list

Methodology and limits

The cited leaderboard uses Coding Agent Index v1.2, a simple average of pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Each benchmark is run three times. A harness can appear more than once because changing the model or reasoning setting changes the evaluated configuration. Scores below preserve the source values captured on July 18, 2026 and are not a universal product verdict.[2]

Published scope: Top ten configurations in descending order from the source data captured on July 18, 2026.

Version note: Index v1.2 snapshot. The live source can change as new evaluation runs are added.