Back to Home

China H100 Rentals Surge 30% as Token Demand Hits 1,000x

While the Western narrative about NVIDIA's Hopper generation has been dominated by collapsing prices, a very different story has been playing out behind China's export-control firewall. Chinese media outlet Calian Press reports that NVIDIA H100 rental prices in China have climbed roughly 20-30% since October 2025, and one-year lease contracts may have surged close to 40% over that same window. The driver is not hype. It is token economics, and the numbers are staggering: China's National Data Administration puts daily token calls at over 140 trillion by March 2026, up from 100 billion in early 2024. That is a more than 1,000x increase in roughly two years.

This is the rare GPU market story where the secondary market is telling us more than the roadmap announcements. H100s are four years old, Blackwell is shipping, and Rubin lands in the second half of 2026. Yet in China, the H-series is not depreciating - it is appreciating. The divergence between the global market and the Chinese market is now so extreme that it deserves a closer technical look.

What the Price Data Actually Shows

Let's get concrete about the numbers, because the magnitude matters. According to sources cited by Calian Press, typical H100 rental rates in China previously sat in the RMB 50,000-60,000 range, with occasional lows of RMB 40,000-50,000. Including rack costs, current pricing has now climbed to roughly RMB 80,000-90,000. That is a 30-50% jump from the low end of the historical band, depending on the contract structure.

The appreciation is not confined to rentals. An H200 system purchased for RMB 2.45 million in February 2025 is now valued at around RMB 3 million after more than a year of operation - a used AI server selling for more than it cost new. Berlin Cloud Servers, the source cited in the report, attributes the rise to broad-based cost inflation across memory, storage, GPUs, CPUs, and optical modules. Memory is seeing the most pronounced increases, which makes sense: HBM supply is the industry's binding constraint, and every token served on an H100 requires high-bandwidth memory bandwidth that is increasingly expensive to produce.

This is the counterintuitive part of the Hopper economics. Globally, H100 rental prices cratered 64-75% from their 2023 peak, with open-market rates hovering around $2.29-$3.12 per hour. China is insulated from that deflation because export controls cap the supply of H100-class silicon entering the country, and the sanctioned gray market that feeds it carries a risk premium. When demand for compute rises faster than the smuggled supply can grow, prices only go one direction.

The Token Boom Behind the Surge

So where is the demand coming from? The 1,000x token surge is being driven by two workload classes that did not exist at scale when Hopper launched. First, large-scale image and video generation via tools like Seedance and Nano Banana, which consume tens of thousands of tokens per output and are now mainstream consumer features in China. Second, high-concurrency, multi-step workflow agents such as OpenClaw, which require continuous iteration loops that keep GPUs occupied for minutes rather than milliseconds. The result, as one Chinese cloud provider put it, is that compute can no longer "stay idle" - utilization is being driven toward saturation.

Three forces compound the pressure:

  • HBM memory inflation: The cost of HBM3e and the packaging capacity around it (CoWoS) keeps climbing, raising the floor under every server price.
  • Supply-side ceiling: Export controls and the 75K-unit cap on H200 sales to China mean the sanctioned supply of Hopper silicon is finite, and gray-market channels add 15-25% premiums.
  • Agent-driven concurrency: Multi-step agent workloads are more compute-hungry than simple chat inference, pushing effective tokens-per-GPU-hour well above what 2024-era demand models assumed.

Cloud Providers Are Passing the Costs On

The rental surge is now propagating into end-user pricing. Tencent Cloud announced on April 9 that it will revise list prices for AI compute, container services, and Elastic MapReduce (EMR) effective May 9, 2026 - with all three categories rising 5%. Alibaba Cloud moved first on March 18, raising prices on compute-card offerings, including T-Head Zhenwu 810E, by 5% to 34%, with CPFS storage up 30%. Notably, JD Cloud is holding its core lineup flat, a strategic choice that will be tested if HBM prices keep climbing.

The strategic read is more interesting than the price chart. NVIDIA's official China market share is collapsing - Bernstein projects it will fall to around 8% in 2026 as Huawei's Ascend line climbs toward 50%. Yet the H100 rental market is booming. Both can be true, and the tension between them is the real story: the sanctioned, policy-approved market is being rebuilt around domestic silicon, while the gray-market Hopper fleet remains the workhorse for the token-hungry workloads that Chinese labs and startups actually run today.

The technical takeaway for anyone planning AI infrastructure: treat regional GPU markets as separate liquidity pools. Capacity planning based on global spot prices will badly misprice compute in export-controlled markets, where scarcity premiums, memory inflation, and token growth compound each other. For the hyperscalers and neoclouds watching from outside China, the signal is that token demand growth - not model announcements - is the fundamental driver of GPU economics. When daily token calls go from 100 billion to 140 trillion, even a four-year-old GPU becomes an appreciating asset.

The bottom line: Hopper's second act in China is not the quiet depreciation story the rest of the world is seeing. It is a scarcity-driven bull market, fueled by a 1,000x token explosion and a supply ceiling that export controls will not lift anytime soon. For engineers and operators, the lesson is straightforward - watch the token counters, not the keynote slides, because they are the leading indicator for where compute prices go next.

Comments

No comments yet. Be the first to share your thoughts!