Back to Home

Grok 4.6 Ties GPT-5.6 Sol, SpaceX Stock Soars

On August 12, xAI, now operating as SpaceXAI after a merger that made even aerospace analysts do a double take, released Grok 4.6. What happened next was the strangest part: a model update performed like a rocket launch. SpaceX stock jumped roughly 11 percent within days and about 36 percent across five sessions, adding more than $500 billion to the company's market cap. The catalyst was a chatbot upgrade. Wall Street said yes.

Grok 4.6 is not a fresh pretrain. It is a hefty post-training and reinforcement-learning upgrade on Grok 4.5, built partly with the RL expertise xAI absorbed when it acquired Cursor. The mission: long-running agents, multi-step reliability, sub-agents, self-verification, and the ability to stay on a task long enough to finish it instead of bailing at the first merge conflict. It is a model designed to babysit your codebase.

Benchmarks: Tied at the Top, Priced Like a Bargain

The headline number is 61 on the Artificial Analysis Intelligence Index, which ties GPT-5.6 Sol Max and trails only Claude Fable 5 Max at 62 and Claude Opus 5 at 63. On the GDPval-AA v2 knowledge-work benchmark, Grok 4.6 ranks second with an Elo of 1,753 and finishes complex tasks in about 53 steps. Claude Opus 5 needs roughly 103. It literally finishes the job twice as fast.

  • 61 Intelligence Index score, tied with GPT-5.6 Sol Max
  • 500K context window, same as Grok 4.5
  • $2 per million input tokens, $6 per million output, more than 60 percent below GPT-5.6 Sol and Claude Opus 5
  • Real-world cost per task around $0.84 in independent testing
  • Live in Cursor, Grok Build, the API, OpenRouter, Vercel, and Cloudflare, with GitHub Copilot support landing August 14

The Agentic Coding Story: Built to Babysit

For agentic coding specifically, early independent testing is genuinely strong. One developer and founder walked Grok 4.6 through a security audit of a real project, a migration of a Cursor integration from one adapter to the official SDK, and a roughly 1,000-line pull request that the model planned, implemented, filed, and babysat until it merged. Multi-step, multi-context work is exactly where it shines.

xAI says the model was trained on agentic RL tasks spanning kernel optimization, web development, computer-aided design, and general coding, and that it started showing real self-verification behavior on longer trajectories. It checks its own work before moving on, which is more than some human juniors manage on a Friday afternoon.

The catch is that the efficiency magic is gone. Independent testing shows Grok 4.6 burns more than 30 percent extra tokens compared with Grok 4.5, and cache reads got slightly pricier. Design and 3D work remain the embarrassing uncle at Thanksgiving: one reviewer called the outputs old AI slop, and a 3D aquarium game port failed outright with a black screen on the first attempt, then delivered inverted axes and bad geometry after a fix. Open-weight models did better. Aquarium games: the final frontier.

The Stock Market Side Quest

Then there is the market reaction, which is the most absurd chapter. Financial wires put the SpaceX surge at around 11 percent on Wednesday and about 36 percent over five sessions, adding more than $500 billion in market capitalization. The trigger was Grok 4.6, an upgrade of a software product. Wall Street decided that a model good at long-running coding tasks makes a rocket company's AI ambitions worth half a trillion dollars more.

Never mind that the AI unit's losses were reported at $1.26 billion. In 2026, vibes are a growth asset, and Grok 4.6 has excellent vibes.

If you are wondering whether to adopt Grok 4.6 or wait, Elon Musk has already teased that Grok 4.7 is three to four weeks out and significantly better. They are shipping model versions faster than phone makers ship colorways. By the time your team finishes evaluating 4.6, 4.7 will be on stage holding a microphone and asking why you waited.

For developers, the practical news is still good. Frontier-adjacent agentic coding at $2 and $6 per million tokens, double usage in Cursor and Grok Build for the first week, and GitHub Copilot support already live. If you have been paying GPT-5.6 Sol prices for the same benchmark tier, the market finally noticed the price tag was a bit much.

The joke writes itself. The lab that put a chatbot on a rocket stock now wants to put your codebase on autopilot. Grok 4.6 ties the best models in the world, undercuts them by more than half, and lets a rocket company's stock take the victory lap. Whether Grok 4.7 keeps the streak alive in a month is the next episode of a genuinely entertaining show.

Comments

No comments yet. Be the first to share your thoughts!