Agent ranking · Coding Agent Performance
Top 10 coding-agent configurations
Snapshot evaluated Jul 18, 2026
A July 18, 2026 snapshot of the ten highest-scoring configurations exposed by the current Artificial Analysis Coding Agent Index. The harness, model, reasoning setting, and fallback behavior are part of each result.
Dated leaderboard
Ranked configurations
Source retrieved Jul 18, 2026
| Rank | Agent | Evaluated configuration | Score | Evidence |
|---|---|---|---|---|
| #1 | Codex OpenAI | GPT-5.6 Sol · max | 61.014Coding Agent Index v1.2 | The captured v1.2 leaderboard data places this Codex configuration first at 61.014.[1] |
| #2 | Claude Code Anthropic | Claude Fable 5 · max · with fallback | 59.216Coding Agent Index v1.2 | The captured v1.2 leaderboard data places this Claude Code configuration second at 59.216.[1] |
| #3 | Grok Build SpaceXAI | Grok 4.5 · high | 57.898Coding Agent Index v1.2 | The captured v1.2 leaderboard data places this Grok Build configuration third at 57.898.[1] |
| #4 | Kimi Code Moonshot AI | Kimi K3 | 56.765Coding Agent Index v1.2 | The captured v1.2 leaderboard data places Kimi Code CLI with Kimi K3 fourth at 56.765.[1] |
| #5 | Claude Code Anthropic | Claude Opus 4.8 · max | 54.896Coding Agent Index v1.2 | The captured v1.2 leaderboard data places this Claude Code configuration fifth at 54.896.[1] |
| #6 | OpenCode OpenCode | Muse Spark 1.1 · xhigh | 48.703Coding Agent Index v1.2 | The captured v1.2 leaderboard data places this OpenCode configuration sixth at 48.703.[1] |
| #7 | Claude Code Anthropic | GLM-5.2 | 39.907Coding Agent Index v1.2 | The captured v1.2 leaderboard data places this Claude Code configuration seventh at 39.907.[1] |
| #8 | Cursor CLI Cursor | Composer 2.5 Fast | 33.757Coding Agent Index v1.2 | The captured v1.2 leaderboard data places Cursor CLI with Composer 2.5 Fast eighth at 33.757.[1] |
| #9 | Gemini CLI Google | Gemini 3.1 Pro · high | 29.174Coding Agent Index v1.2 | The captured v1.2 leaderboard data places this Gemini CLI configuration ninth at 29.174.[1] |
| #10 | Claude Code Anthropic | DeepSeek V4 Pro · high | 28.483Coding Agent Index v1.2 | The captured v1.2 leaderboard data places this Claude Code configuration tenth at 28.483.[1] |
How to read this list
Methodology and limits
The cited leaderboard uses Coding Agent Index v1.2, a simple average of pass@1 across DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Each benchmark is run three times. A harness can appear more than once because changing the model or reasoning setting changes the evaluated configuration. Scores below preserve the source values captured on July 18, 2026 and are not a universal product verdict.[2]
Published scope: Top ten configurations in descending order from the source data captured on July 18, 2026.
Version note: Index v1.2 snapshot. The live source can change as new evaluation runs are added.