Back to Home

DeepSeek Just Quadrupled Prices. Still a Steal.

DeepSeek, the lab that turned "cheap AI" into a whole personality, has officially done the thing it spent the last year insisting it would never need to do: it raised the prices. By a lot. Up to 1,100 percent on some rates. And somehow, it is still the cheapest serious model in the building.

The twist is worth savoring. Last week DeepSeek warned developers that API prices would rise "significantly," and the internet reacted the way you react when your landlord says "we should talk about rent soon." This week the actual bill arrived. V4-Pro-0813, the flagship's general-availability build, launched alongside a new peak and off-peak pricing scheme that roughly quadruples headline rates and raises some access tiers by as much as 1,100 percent.

The new pricing went live this weekend: Sunday afternoon UTC, which is Monday morning in Beijing, which is exactly the sort of timing designed to ruin nobody's weekend until they check their API dashboard at work.

The Receipt Arrives: V4-Pro Finally Leaves Preview

The model itself is a monster. V4-Pro is a mixture-of-experts system with 1.6 trillion total parameters, 49 billion active per token, and a one-million-token context window, roughly 1,500 pages of text in a single prompt. It wrapped up a preview period that ran since late April, and the production build landed on the API and OpenRouter on August 12.

Under the hood, DeepSeek leans on a hybrid attention architecture that it says cuts single-token inference compute to 27 percent and KV cache to 10 percent of what its V3.2 generation needed at the million-token context. That is the difference between a context window that looks good on a spec sheet and one you can actually afford to run.

You also get a reasoning dial: low for quick answers, high for daily agent work, and max for the sort of problems that make you cancel your evening plans. In max mode the model does not just think. It composes a novel about thinking. Artificial Analysis logged 130 million output tokens during its evaluation of the max variant, and the evaluation itself cost $604.51 to run. The model talked so much that testing it cost more than most startups spend on coffee in a quarter.

The numbers hold up. V4-Pro resolves 80.6 percent of SWE-bench Verified issues, a fraction behind Claude Opus 4.6 at 80.8, and scores 90.1 on GPQA Diamond. Artificial Analysis gives the max variant an Intelligence Index of 53 against a class median of 27, while the API serves up 81 tokens per second with a snappy 1.85-second time to first token.

Surge Pricing, But for Tokens

The pricing story is the funnier one. DeepSeek has introduced peak and off-peak tiers, which is a sentence nobody expected from the lab that started the race to zero. Peak hours run 9 a.m. to noon and 2 p.m. to 6 p.m. Beijing time, and off-peak traffic gets a 50 percent discount. Here is the new scorecard for V4-Pro:

  • Peak: $1.32 per million input tokens, $3.96 per million output tokens
  • Off-peak: $0.66 and $1.98, exactly half price
  • Cache-hit input: from 0.36 cents per million tokens to 2.2 cents off-peak and 4.4 cents peak, the sixfold and twelvefold jump behind those 1,100 percent headlines

Context makes the numbers land. At $3.96 per million output tokens, V4-Pro costs about 14 times as much as V4 Flash, and roughly four times its own preview pricing. The coverage wrote itself: "quadruples prices," "up to 1,100 percent," "the cheap era is over." DeepSeek is now second only to Anthropic in token consumption across major API platforms, which means the world is mainlining the discount model and complaining about the new menu at the same time.

But here is the punchline: it is still absurdly cheap. Decrypt framed it as "Claude Fable is only 5 percent better at 4,500 percent the price." DeepSeek is the only company on Earth that can raise prices by four figures in percentage terms and still be the value option. If you are a developer, this is your Uber surge moment. Schedule the heavy batch jobs, the nightly codebase analysis, and the big retrieval runs into the off-peak window like a responsible vampire, and your bill barely moves.

Also, There Is a Claude Code Rival Now

Buried under the price drama was the release that may matter more in the long run. DeepSeek Harness, an open-source framework for turning models into autonomous agents, landed in developer preview on Thursday, and the design philosophy is best summarized as: everything is a plugin. Every core element is modular and swappable, in sharp contrast to Anthropic's integrated, ready-to-use Claude Code.

Harness ships with four operational settings, which is four more than most of us have for our own work schedules:

  • Standard mode for general tasks
  • A code-focused mode that lets the agent drive multiple applications at once
  • A creative mode for experimenting with custom tools
  • A minimal mode for isolated testing

DeepSeek has been staffing up for this since March, when it hired former Jane Street engineer Cui Tianyu into its harness group. The strategic read from SCMP is that competition is shifting from raw model intelligence to how seamlessly an agent connects to real-world software, and DeepSeek is building the scaffolding for that fight in the open.

So the story of the week is a delightful contradiction. The lab that made cheap AI its brand just installed dynamic pricing, quadrupled its flagship rates, and still ended up as the best deal on the market, all while quietly shipping a Claude Code rival on the side. The race to zero hit a speed bump, and it is shaped like a pricing table that takes effect at 16:00 UTC. Bring your tokens, and maybe schedule your agents for 3 a.m.

Comments

No comments yet. Be the first to share your thoughts!