Back to Home

Zhipu GLM-5.3 Doubles SWE-Marathon, Tops CyberGym

Zhipu AI's founder just lobbed a grenade at the industry's favorite religion. Trillion-parameter foundation models, Tang Jie argues in a striking new take on scaling laws, were something of a detour. The real intelligence leap, he insists, is hiding in what happens after the base model is trained: post-training. In case anyone needed proof, the lab's GLM-5.3 delivered a thunderclap on August 14 that turned the statement from philosophy into a benchmark sheet.

The Beijing lab that gave the world GLM-5.2 now finds itself in an extraordinary position. Its newest model doubled its SWE-Marathon coding score, topped CyberGym, tied the best open-source challenger Kimi K3, and stood toe-to-toe with closed-source flagships like Claude Fable 5 and GPT-5.6 Sol. Each model generation no longer feels like an incremental step. It feels like a re-declaration of where China's most aggressive AI bet is heading.

The Detour That Changed the Playbook

For years, the entire industry ran on one rule: bigger equals smarter. Every cycle added parameters, tokens, and megawatt-hours in a race to pile compute onto the pretraining run. It worked, spectacularly, until it stopped being the whole story.

Tang Jie's mid-August comments reframe that narrative. Instead of scaling pretraining without end, he argues the decisive gains are now coming from the fine-tuning, reinforcement-learning, and alignment stage that shapes a raw model into a truly capable assistant. The base model matters, but the intelligence leap is increasingly manufactured in post-training, where a model learns to reason, follow instructions, and actually complete hard tasks.

GLM-5.3 is being held up as the walking proof. It is built as what Zhipu calls a coding and security model, with post-training doing the heavy lifting on the two capabilities customers care about most. That is exactly the bet Tang Jie has been making since DeepSeek's explosive arrival rewrote the rulebook on cheap, powerful reasoning models.

GLM-5.3 by the Numbers

The response to the August 14 drop was not a trickle. Zhipu's sales team reported being flooded with calls and WeChat messages within the first hour, with enterprise clients racing to lock in inference compute before launch. Demand grew so intense that the API rollout, originally scheduled for August 18, was pushed back to keep the maths honest.

Here is what the breakthrough actually looks like in concrete, market-moving terms:

  • Doubled its SWE-Marathon coding score and climbed to the top of the CyberGym security benchmark.
  • Tied Kimi K3 for first place among open-source models, matching the level of closed-source leaders Claude Fable 5 and GPT-5.6 Sol.
  • Hitched Zhipu's annual recurring revenue to roughly $1 billion by July, up from an estimated $500-600 million just two months earlier.
  • Surged to the fastest growth in model invocation volume since 2026 on Vercel, outpacing even DeepSeek V4.

The momentum is not confined to Zhipu's own marketing. Cloud partners including Alibaba Cloud, Volcano Engine, and even Xiaomi and Kingsoft Cloud have been pushing GLM aggressively, a shareholder told 36Kr, because enterprises keep asking for it by name.

From Coding Bet to Revenue Breakthrough

This did not happen overnight. In May 2025, with DeepSeek hijacking the headlines and poaching nearly a third of Zhipu's enterprise clients, Tang Jie gathered a small group of core executives and shareholders to pick the company's next move. The verdict: abandon a scattered collection of vertical models and bet everything on a single all-in-one model built around reasoning, coding, and agentic tasks.

That gamble became GLM 4.5 in July 2025, Zhipu's first real coding hit and the moment it became China's first large-model company to own the coding track. Shares listed in Hong Kong in January 2026, making Zhipu the world's first publicly traded large-model company, with a market capitalization that has since peaked near HK$1.3 trillion.

"The chatbot story ended after DeepSeek came out," Tang Jie said at an event in early 2025. "What we should think about is what the next bet is." With GLM-5.3 and a post-training thesis to back it, that bet is paying off faster than almost anyone predicted.

The euphoria is justified, at least for now. Zhipu has turned a contrarian view of scaling into real revenue, real benchmarks, and real enterprise adoption. If Tang Jie is right that post-training is the next frontier, the rest of the industry is going to have to answer the question he has already bet the company on: what happens when the smartest models are the ones that stop training first and start reasoning harder.

Comments

No comments yet. Be the first to share your thoughts!