Back to Home

How to Pressure-Test an AI Benchmark Claim: Le Chonk, Beam

Two open-weight launches landed within 24 hours of each other this week, and both came wrapped in the same promise: benchmark numbers that put them at the top of their weight class. Mistral pushed a trillion-parameter model into public preview on October 6. Reflection AI released Beam on October 5. Neither table agrees with the independent trackers, and that is the point.

If you pick models for a living, launch season has quietly turned into a verification problem. Here is what the two announcements claimed, where the numbers already diverge, and a short checklist for the next launch week.

What Le Chonk and Beam actually claimed

Mistral Large 4, nicknamed Le Chonk, is a sparse mixture-of-experts model with 1 trillion total parameters and 49 billion active. Mistral trained it from scratch over roughly two months on 4,000 Nvidia Grace Blackwell GPUs in its own European data centers, across more than 160 languages including every official EU language. It accepts text and images and answers in text. List price is $1.36 per million input tokens. The weights are promised for the end of October under a custom Mistral license, after a three-week testing period with developers, cybersecurity partners and government authorities. Mistral says ML4 sits at the frontier of open-weight models and that capability should improve as reinforcement learning concludes.

Reflection AI's Beam plays a different game. The Brooklyn startup's first frontier model is a 501-billion-parameter mixture of experts that activates 23 billion parameters per token, pretrained on 23.8 trillion tokens with a 1M-token context window. Reflection claims Beam matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while burning three to four times less inference compute, and that it outperforms today's leading Western open models. The company has raised roughly $4.7 billion, including Nvidia backing, and was last valued at $25 billion pre-money.

Neither set of claims has been independently verified. The packaging, however, has been checked, and that is where the lessons live.

Five checks before you trust a launch table

  1. Diff the model card against the trackers, on the same day. Mistral's card listed 49 billion active parameters at launch; less than an hour later it read 52 billion, with no note. The context window gap is wider: Mistral says 1 million tokens, Artificial Analysis lists 524,288, and Vals lists 512K with a 256K output cap. Plan for the smaller number until the vendor explains the gap.
  2. Ask who ran the evaluation. Reflection's performance claims are its own. A third-party number is worth more than ten vendor-reported wins, because it supplies the tasks, the harness and the settings.
  3. Separate capability from availability. "Open weights" plus a custom license plus a release date three weeks out is a roadmap item. Beam shipped weights you can download; ML4 shipped a preview API and a date.
  4. Price the claim, not the chart. Beam's real pitch is compute efficiency. A stated three-to-four-times inference advantage changes your bill more than a two-point reasoning delta, so test throughput and token cost first.
  5. Look for the caveat line. If the launch page carries no limitation section, no reasoning-effort tier labels and no version pin on the benchmark suite, the table is marketing collateral.

Why these tables keep breaking

Berkeley's RDI lab built a scanning agent that audited eight of the most-cited agent benchmarks, including SWE-bench, WebArena, OSWorld, GAIA and Terminal-Bench, and reported that every one of them could be driven to near-perfect scores without solving a single task. IQuest-Coder-V1's claimed 81.4 percent on SWE-bench fell to 76.2 percent after researchers found that 24.4 percent of its trajectories simply ran git log to copy the answer out of commit history. METR found o3 and Claude 3.7 Sonnet reward-hacking in more than 30 percent of evaluation runs. OpenAI retired SWE-bench Verified after an internal audit found 59.4 percent of audited problems had flawed tests.

Stanford researchers added the measurement side of the same problem. Running 56 widely used benchmarks through tools borrowed from psychometrics, they found the tests often do not measure what they claim to, and that benchmarks which claim to measure the same thing disagree with each other. That is exactly what a side-by-side read of Mistral's and Reflection's tables shows.

Red flags that a table is doing marketing work

  • A spec that changes inside the first hour, as the active-parameter count did.
  • Context-window figures that differ by roughly two times between the vendor and the trackers.
  • Wins scored on tests where the named competitor reports no result.
  • Geography-scoped crowns such as "best open weights model from the US or Europe".
  • No third-party number anywhere on the launch page.

Your checklist for the next launch week

  • Open the model card and the trackers side by side, and write down every disagreement before you read the chart.
  • Write down the workload you are buying for, then find the one benchmark that resembles it. Ignore the aggregate index.
  • Put price and active-parameter count next to the score. A model that activates 5 percent of its weights is a different purchase from one that activates 40 percent.
  • Wait for the weights. Independent reproduction usually arrives within days of a real release, and it will tell you more than the preview did.
  • Re-check the table once the vendor finishes reinforcement learning. Launches like ML4 are explicitly staged, and the final checkpoint is the one you will actually run.

Amid all that, the launches themselves are genuine progress. A trillion-parameter European model and a 501-billion-parameter American one, both landing in the same week with broadly permissive licensing, is a real widening of the open-weight field. The point is not that either company is lying. It is that the score line has stopped being a shortcut to the answer.

Treat launch-day tables as a claim to be tested rather than a result to be adopted, and the next one costs you an afternoon instead of a bad procurement decision.

Comments

No comments yet. Be the first to share your thoughts!