Back to Home

3 Cents Per Test: DeepSeek V4-Flash vs Gemini 3.6 Flash

On August 3, research firm Artificial Analysis dropped a number that quietly reset the terms of the AI price war: DeepSeek's V4-Flash completes a benchmark test for an average of three cents. On the firm's Intelligence Index, the same model scores 50 out of 100, exactly matching Google's Gemini 3.6 Flash. The cheapest well-known model to run is no longer a curiosity. It is now a pricing reference point for the entire market.

The number matters because it measures something different from list price. Three cents is the average cost of completing each test in the benchmark suite, not the cost of a single customer query. The Intelligence Index bundles nine coding, reasoning, and workplace tasks into one score, so the per-test figure reflects how much compute a model actually burns to reach an answer, including reasoning tokens, retries, and output length. That makes it a far more honest comparison than per-token sticker prices, which ignore how verbose a model is.

The 3-Cent Metric, Decoded

DeepSeek's published token pricing is already aggressive: $0.14 per million input tokens and $0.28 per million output tokens. But token price alone understates the efficiency story, because a cheaper-per-token model can still be expensive if it generates far more tokens per task. The per-test metric folds all of that in.

Here is what Artificial Analysis' benchmark-cost comparison looks like across the field:

  • Anthropic Claude Fable 5: $3.15 per benchmark test
  • OpenAI GPT-5.6 Solu: $1.86 per test
  • Moonshot AI Kimi K3: $0.86 per test
  • DeepSeek V4-Flash: $0.03 per test

At that spread, V4-Flash is more than 100 times cheaper to operate than Claude Fable 5. The gap is not a rounding artifact; it is structural. V4-Flash is a 284-billion-parameter mixture-of-experts model that activates only about 13 billion parameters per token, an architecture designed around keeping inference cost low. Frontier rivals burn comparatively huge token budgets per task, which is exactly what the per-test number exposes.

Parity With Gemini 3.6 Flash, and Where It Stops

The Intelligence Index score is the other half of the story. V4-Flash lands at 50, matching Gemini 3.6 Flash, the model Google positioned as its low-cost workhorse. But parity here is parity at the bottom of the chart. Moonshot's Kimi K3 scores 57, and the frontier models from OpenAI and Anthropic clear V4-Flash by at least seven points.

Read that combination correctly: V4-Flash is not the smartest model on the board. It is the cheapest model at a specific, increasingly usable capability level. For routine enterprise workloads, summarization, classification, code generation, and agent scaffolding, that level is enough. The market for high-volume commodity inference is where the 3-cent price does its damage, not in the research-preview tier.

One caveat belongs in any serious reading of these numbers. Artificial Analysis runs a fixed battery of nine tasks, and a benchmark average is not a workload guarantee. The 3-cent figure is a fleet average across its suite; real pipelines with long reasoning chains, image inputs, or heavy tool use will bill differently. Treat it as a directional signal, not a contract.

Why Alibaba Turned It Into a Price War

The most interesting reaction came from a company that does not even own DeepSeek. Alibaba Cloud's AI Gateway already supports DeepSeek V4 APIs and can route workloads between DeepSeek and its own Qwen models. Alibaba gets to monetize the price pressure either way: when DeepSeek's low prices pull developers into its cloud, Alibaba sells them the surrounding services.

On the same day, August 3, Alibaba unveiled Qwen3.8-Max, a 2.4-trillion-parameter model, and its Hong Kong shares rose 7 percent. The timing frames the strategy: flood the low end with cheap open-weight options, then make money on infrastructure. The cloud numbers show the bet working so far. Cloud revenue grew 38 percent in the March quarter, external cloud revenue rose 40 percent, and AI products now make up 30 percent of external cloud sales. The catch is that group revenue grew only 3 percent, and Alibaba plans to exceed its earlier 380 billion yuan three-year AI commitment, treating margins as secondary.

That creates a demanding unit-economics test. Lower inference prices pull customers into Alibaba Cloud, but revenue compounds only if workload volume grows faster than prices fall, and if those users also buy storage, networking, and databases. Cutting the price of intelligence is a customer-acquisition strategy, not a business model by itself.

Google's defense shows the alternative path. Google Cloud revenue jumped 82 percent to $24.8 billion in the second quarter of 2026, with operating margin expanding to 35.6 percent. Customers still pay for distribution, infrastructure, and integrated products even when standalone models become commodities. DeepSeek attacks model pricing, not Google's entire moat.

None of this is lost on DeepSeek itself. Days after the 3-cent benchmark went public, the company warned developers of a significant API price hike, a quiet admission that its own economics are not designed to stay at zero forever. The 3-cent test is best read as a floor that keeps lowering, and as a ceiling on what any vendor can charge for routine inference. For engineers choosing a stack, cost per benchmark test is now the metric to watch.

Comments

No comments yet. Be the first to share your thoughts!