Back to Home

Which LLM Leaderboard Should You Trust? Arena vs AA

Every few weeks another model tops another leaderboard, and the boards almost never agree. Arena (formerly LMArena), the Artificial Analysis Intelligence Index and the independently run SWE-Bench Verified rankings routinely crown different winners from the same crop of frontier models. That is not a flaw: each measures something genuinely different, and knowing which one fits your workload is the difference between a confident pick and a hollow headline.

Three Leaderboards, Three Different Winners

Start with Artificial Analysis. It is the closest thing the field has to an independent measurement lab. It runs every evaluation itself against its own internal dataset copies rather than trusting vendor self-reported numbers, with standardized settings across all models and single-attempt pass@1 scoring. Its Intelligence Index, at version 4.1.1, blends nine evaluations across four weighted categories: agentic work counts for 34 percent, coding 24 percent, scientific reasoning 24 percent and general capability 18 percent. The caveats are stated plainly: the suite is text-only and English-language, and its rows are split by reasoning effort, so Claude Opus 5 appears at 63.1 when pushed to max reasoning yet 58.6 at medium effort. Quoting one number without the tier label says almost nothing.

Arena answers a question measured evals cannot: which model do humans actually prefer? Rebranded from LMArena in January 2026, it pits two anonymous models against the same prompt and converts millions of blind votes into Elo-style ratings. The scientific text board alone has absorbed more than 7.7 million votes. The cost of that signal is that Arena rewards conversational polish and confident formatting as much as correctness, which is exactly why it carries a style-control toggle.

SWE-Bench Verified looks at none of that. It drops a model into a real GitHub repository with a bug report and a test suite, gives it one tool (bash) and grades pass or fail with no partial credit. Run by vals.ai with a deliberately minimal harness, it is the cleanest available measure of finished coding work. And it is why the three boards disagree so sharply.

Who Wins Depends on the Question You Ask

On SWE-Bench, the answer is unambiguous. Claude Opus 5 leads at 97.00 percent, with DeepSeek V4 Pro 0813 second at 96.40 percent for $1.32 per million input tokens, roughly a quarter of Opus 5's list price. GPT-5.6 Sol (96.20 percent), Grok 4.6 (95.60 percent), GPT-5.6 Terra and GLM-5.3 (both 95.40 percent) and Kimi K3 (93.40 percent) round out a top tier where open-weight models now live at the frontier, not near it.

Flip to Arena's Elo board and the order changes completely. Claude Fable 5 leads at 1507, while Opus 5 High sits seventh, and Anthropic holds six of the top ten slots on pure human preference. The same family dominates both lists, but the exact winners are different, which is the cleanest demonstration available that conversational polish and finished code are separate skills.

That split is where the Comparative verdict lands. If you are building a coding agent, ranking by Arena is optimising for the wrong thing; Opus 5, mediocre in blind chat votes, sweeps the scaffolding-heavy task benchmark. If you are shipping human-facing chat or marketing copy, the SWE-Bench table will mislead you just as badly. Pick the instrument that matches the job.

Two numbers put the stakes in perspective. Across the top ten, accuracy spans only 8.2 points while cost spans 129 times. DeepSeek V4 Flash resolves 88.8 percent of tasks at about a cent per try, while Claude Fable 5 hits 95 percent at more than two dollars, roughly 200 times the price for a 6.2 point gain. The right choice depends on what a failed task costs you.

Your Practical Playbook

Use the boards as a shortlist generator, not a verdict. A rank gets you two or three candidates; your own harness decides the winner. Remember that static benchmarks saturate and leak, which is why Hugging Face retired its Open LLM Leaderboard, and that vendor launch numbers are run on scaffolding the vendor chose. Trust the intersection of an independent measured index and a job-matched board, then verify on your own tasks.

  • Agentic coding, quality first: Claude Opus 5 at 97.00 percent, or DeepSeek V4 Pro 0813 at 96.40 percent for a fraction of the cost.
  • Interactivity and latency: GPT-5.6 Terra finishes a task in about 180 seconds, where GLM-5.3 matches its accuracy but takes 841 seconds.
  • Human-facing chat: let Arena be your guide, where Claude Fable 5 and the Opus 4.6 and 4.7 variants dominate.

The headline race is compelling, but it rarely matches your actual workload. Compare the models on the board that reflects what you build, weigh cost against accuracy, and let the leaderboards do their real job: narrowing the field until only your own test matters.

Comments

No comments yet. Be the first to share your thoughts!