Back to Home

Mistral Large 4 vs GPT-6 Astra: The Price War Is On

On October 6, 2026, at AI Everything Abu Dhabi, Mistral AI chief executive Arthur Mensch pulled the cover off Mistral Large 4, a 1.05-trillion-parameter model the company nicknamed Le Chonk. This is not a research demo. The API opened the same day in public preview, and the company says the weights themselves land on October 27.

That timing sets up the sharpest head-to-head in the open-weight world right now: Mistral's European flagship against GPT-6 Astra, the Chinese open-weight wave led by Z.ai's GLM-5.3, and Anthropic's Claude Opus 5. Four models, three different business models, and one question. Does open weight finally beat proprietary on the only scorecard enterprises actually read, which is cost per unit of useful capability?

Meet the Contenders

ML4 is a granular Mixture-of-Experts system: 1.05 trillion total parameters but only about 49 billion active per token, roughly 4.7 percent of the stack. It carries a 1.6-billion-parameter vision encoder for native image input, a 1-million-token context window, and a training run that spanned more than 160 languages. Mistral says the model was trained from scratch on roughly 3,800 NVIDIA Grace Blackwell GPUs inside its own European data centers.

The architecture is the pricing story. Because only a sliver of the network fires per token, a trillion-class model can be served at mid-tier rates: $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14. One caveat belongs in the slide deck: the full 1.05 trillion weights still have to sit in memory, so the activation count sets compute cost, not your hardware bill.

  • Mistral Large 4 (Le Chonk): 1.05T total parameters, 49B active, 1M context, weights promised for October 27
  • GPT-6 Astra: the closed frontier reference point, priced by tier rather than by open token economics
  • GLM-5.3 and GLM-5.3-Flash: Z.ai's open-weight challenger, already shipping under permissive licenses
  • Claude Opus 5: Anthropic's premium model, currently the leader on blind human preference testing

Scorecard: Coding, Cyber, and Cost

Mistral's published numbers are strong, and they are also vendor-reported. On agentic coding the company reports 61.7 percent on DeepSWE v1.1, 59.4 percent on SWE-Atlas-QnA, and 28.3 percent on Terminal-Bench 4.0, which combine to a 49.8 percent Artificial Analysis Coding Agent Index. Mistral notes these were evaluated privately before the harness went public, so nobody outside the company has reproduced them yet.

Cybersecurity is where the claims get sharper. Mistral reports 93 percent on Cybench and 82 percent on CyberGym-E2E, placing ML4 in the global top five on the Artificial Analysis Cyber Index. The structural point matters more than the score. Mistral says several closed frontier models score near zero on CyberGym-E2E because they refuse the task outright rather than because they lack the capability. Reproducing a vulnerability to prove it is real is ordinary defensive work, and provider-level refusals block it.

  • Coding, vendor-reported: DeepSWE v1.1 at 61.7 percent, SWE-Atlas-QnA at 59.4 percent, Terminal-Bench 4.0 at 28.3 percent
  • Cyber, vendor-reported: Cybench 93 percent, CyberGym-E2E 82 percent, and a Cyber Index score of 50, level with GLM-5.3-Flash and behind MiMo-V2.6-Pro at 56
  • Human preference, blind Surge AI evaluation: ML4 Preview 3.74 out of 5, second of five, behind Claude Opus 5 at 4.22 and ahead of GLM-5.3 at 3.60 and Kimi K3 at 3.59
  • Price: $1.36 per million input tokens, $4.18 per million output tokens, $0.14 for cached input

The blind human evaluation is the most honest signal in the pack, because it does not depend on a vendor-written harness. Finishing second to Claude Opus 5 and narrowly ahead of GLM-5.3 is a credible place to be, not a marketing fantasy. It also deflates the loudest version of the Le Chonk hype, which is fine. Frontier parity was never the promise here. Frontier-adjacent quality at a fraction of the price tag was.

On price the comparison turns brutal in Mistral's favor. Market analysis of the current field puts the average proprietary model near $6.03 per million tokens, with tier floors around $2 input and $10 output at the top of the closed market. ML4 undercuts that output floor by more than half, and its $0.14 cached input rate changes the economics of long-context agent loops, which is exactly where a 1-million-token window earns its keep.

The Verdict

Score it category by category and there is no single winner. Claude Opus 5 takes human preference. GLM-5.3-Flash matches ML4 on the cyber index at a lower operational bar for teams already inside the Chinese open-weight ecosystem. The best raw cyber score belongs to MiMo-V2.6-Pro. Mistral wins the category that decides most enterprise deployments, which is price per unit of capability at frontier-adjacent quality.

There is an irony worth naming. Mistral markets ML4 as sovereign European AI, backed by contracts with the French military and the Luxembourg Armed Forces and a 3 billion euro Series D led by Samsung. The model was still trained on American silicon, on Grace Blackwell parts that every US competitor also buys. European-led, American-hardware-dependent. Sovereignty in AI is currently a software story, not a supply-chain one.

  • Choose ML4 if you want frontier-adjacent throughput at the lowest advertised output price and can live with vendor-reported benchmarks until independent harnesses catch up
  • Choose Claude Opus 5 if human preference scores and refusal-safe behavior outweigh cost
  • Choose GLM-5.3 if you are already committed to a Chinese open-weight stack and want a comparable cyber score
  • Wait for October 27 if self-hosting matters, because until the weights ship ML4 is an API product, not an open model

The bottom line: Mistral Large 4 does not beat the closed frontier on quality, and it does not need to. It resets the price of near-frontier capability, forces every proprietary vendor to defend a premium that is now harder to justify, and arrives with open weights promised inside three weeks. On the scorecard that enterprises actually fill in, that is the round that matters.

Comments

No comments yet. Be the first to share your thoughts!