Back to Home

Who Built Ox Alpha? The Mystery Is the Marketing

A model named stealth/ox-alpha appeared on OpenRouter on August 21 under the provider handle "stealth," and it has already become the most talked-about release of the week. Not because of a press event, and not because of a benchmark paper. Because nobody can confirm who built it, even as community tests put it on par with the most advanced coding models in the world.

What We Actually Know About Ox Alpha

The listing is unusually candid for a model with no owner attached. OpenRouter describes Ox Alpha as a reasoning model aimed at coding, sustained agentic work, and production workloads, built for long-horizon software engineering and jobs that mix text with visual context. It ships with a 1,048,576-token context window, an output cap of 131,072 tokens, and native multimodality. Whoever made it clearly designed it for developers who live inside an IDE all day.

OpenCode, the open-source terminal coding agent, is running a matching promotion. In a post on X it said Ox Alpha would be free for the next week, with zero data retention and rate limits generous enough to call "near unlimited." Together the two platforms claim capacity for a staggering 100 trillion tokens a day, the kind of number a provider throws around when it wants engineers to stress-test a model rather than poke at it politely.

  • Appeared on OpenRouter on August 21 under the provider handle "stealth".
  • Free for one week on OpenRouter and OpenCode, with roughly 100 trillion daily tokens offered.
  • A 1M-token context window alongside native vision and a coding focus.
  • Community trials rate it at par with, and at least once ahead of, GPT-5.6 Sol and Claude Fable 5 on coding.

Stealth launches have quietly become a normal part of how frontier labs test new models before a public rollout, and OpenRouter has hosted its share of them. The pattern is simple: a lab ships a model anonymously, watches how it performs against real developer traffic instead of curated benchmarks, and then either walks it back or unmasks it as the real product once it has proved itself. Ox Alpha fits that routine so perfectly that the developer community immediately started trying to name the lab.

A Mystery With a Method

Let me be direct about what I think is happening: the anonymity is the point. This is not a leak and it is not an accident. When a lab hides a genuinely strong model behind a fake name and floods the market with free tokens, the unanswered question is not a bug in the release. It is the release.

The leading theory is Z.AI, the company formerly known as Zhipu AI, and it is the most technically grounded. A widely shared tokenizer probe ran eleven probes across ten frontier models and found Ox Alpha matches the GLM family on eleven of eleven, while no other lab gets past four. The same thread notes DeepSeek burns 98 tokens on a digit probe where GLM uses 29, and that Ox Alpha reads images even though GLM-5.3 has no vision endpoint at all, suggesting an unreleased flagship. The jailbreaker account Pliny the Liberator reached the same conclusion independently.

The counter-theories are just as plausible on their own terms. Xiaomi has run this exact playbook before, hiding MiMo-V2-Pro on OpenRouter under the alias Hunter Alpha, and Ox Alpha's spec sheet matches the shape of what Xiaomi has been building.

Robert Lukoszko, the CEO of Stormy, read the tokenizer differently and concluded the cl100k_base lineage rules out every Chinese frontier lab and points to Microsoft's unreleased MAI-2, with IBM Granite as an outside bet. Some whisper a delayed Gemini, though a failed vision test undercuts that one. Others point to ByteDance, which has the compute to actually give away trillions of tokens a day, and the numbers seem to back them: over 3.8 trillion tokens served in a single day, several times what GLM-5.3 drew across its entire OpenRouter run.

Here is the part that matters more than the answer. There is a real story underneath the whodunit, and it is about market share. On OpenRouter, the share of tokens processed by US models has collapsed from roughly 70 percent a year ago to about 30 percent today, with Chinese labs such as DeepSeek, Tencent, Xiaomi, and Z.AI eating into that share month after month. A free, high-context, multimodal stealth drop is exactly the kind of move that accelerates that shift. Developers get to push a frontier-class model on real workloads for nothing, learn to depend on it, and quietly retrain their muscle memory on the winner, whoever that turns out to be.

Whether Ox Alpha turns out to be Z.AI's unreleased GLM, a Xiaomi derivative, or a Microsoft flagship served anonymously, the strategy has already earned what no launch event could buy: days of nonstop headline attention, an army of developers stress-testing the model for free, and a permanent seat in the conversation. When the mystery itself becomes the marketing, we can stop pretending the frontier belongs to any single country. It belongs to whoever can put a million tokens in your context window and make you fall in love before you ever ask for the name.

Comments

No comments yet. Be the first to share your thoughts!