The Neocloud Price Collapse Nobody Announced
If you have rented NVIDIA H100 compute at any point since 2023, the numbers that matter most in mid-2026 are not the teraflops. They are the dollars per hour. On the neocloud spot market, an H100 SXM5 now rents for $1.46 per GPU per hour, down from $1.66 in May 2026, while on-demand rates hold around $2.54.
That 43 percent spot-to-on-demand gap is not a glitch in a pricing dashboard; it is a structural signal about where the AI compute market has moved. Supply from tier-2 and tier-3 data centers has finally caught up with demand, and for the first time since the Hopper generation shipped, the people who bought GPUs at the top of the hype cycle are competing on price rather than allocation.
The contrast with the hyperscalers makes the shift even sharper. AWS still lists the p5.48xlarge, an 8x H100 node with 192 vCPUs and 2048 GiB of RAM, at roughly $55 per hour on-demand, or about $6.88 per GPU. Google Cloud's A3 Mega (H100 SXM5) sits around $12.36 per GPU per hour, and the A3 Ultra, powered by the H200 with its 141 GB of HBM3e, went generally available in Q2 2026 at approximately $15.20 per GPU per hour on-demand.
A team running the same long H100 training job on a neocloud at $2.54 per hour versus AWS at $6.88 is looking at roughly 2.7x cost savings with zero architectural changes. That is the kind of delta that moves real budgets, and it is why neocloud capacity keeps getting absorbed as fast as it comes online.
HBM3e Supply Math Explains the Paradox
Here is the counterintuitive part: spot prices are falling for H100 while the supply chain that feeds its successor is tighter than ever. The H200, which pairs the same SXM5 form factor with 141 GB of HBM3e instead of 80 GB of HBM3, commands a persistent premium on the spot market at $1.77 per hour. Custom H200 and GB200 NVL72 configurations are running 36 to 52 week lead times as of May 2026, because HBM3e packaging capacity is the binding constraint across the entire AI hardware pipeline. TSMC, Samsung, and SK Hynix have all reallocated wafer and packaging capacity toward HBM3e and HBM4, which feed exclusively into data-center accelerators. Every wafer that becomes HBM is one that does not become GDDR7 for a consumer card, which is exactly why the retail GPU market and the AI rental market are moving in opposite directions this year.
What this means in practice is a market where memory capacity, not raw compute, sets the price tier. The 141 GB H200 is the workhorse for production inference workloads that do not fit in an 80 GB H100, and teams that cannot get custom builds are absorbing neocloud spot capacity instead.
H100 SXM5 availability, by contrast, looks stable through H2 2026 because its older memory generation faces no comparable supply constraint. The result is a widening price ladder that rewards knowing exactly where your workload sits on the memory wall:
- H100 SXM5: $1.46/hr spot, $2.54/hr on-demand, 80 GB HBM3
- H200 SXM5: $1.77/hr spot, $4.84/hr on-demand, 141 GB HBM3e
- B200 SXM6: $2.71/hr spot, $7.37/hr on-demand, 192 GB HBM3e
- B300 SXM6: $3.29/hr spot, $9.02/hr on-demand, 288 GB HBM3e
- AWS p5.48xlarge: ~$6.88/hr per GPU on-demand, 8x H100 node
Spot pricing for the Blackwell Ultra B300 opened at $3.29 per hour per GPU in early June, which makes the newest Blackwell part only 2.25x the price of an H100 spot hour while carrying 3.6x the memory. For fault-tolerant batch inference and fine-tuning jobs, the price-performance math on newer parts keeps improving, but the H-series retains one decisive advantage: it runs in air-coolable 20-45 kW racks that ordinary data centers can host without liquid-cooling retrofits. The Blackwell NVL72 class demands 120-132 kW liquid-cooled racks, a facility-level forklift that most mid-size operators simply will not undertake this cycle.
What HBM4 and Vera Rubin Mean for the Rental Market
The next structural shift is already visible in the supply chain. SK Hynix and Samsung are targeting H2 2026 for HBM4 volume production, which will feed the Vera Rubin NVL72 platform and next-generation AMD Instinct builds. NVIDIA confirmed in June that Vera Rubin NVL72 enters production ramp in Q3 2026, pairing 72 Rubin GPUs per rack with roughly 20.7 TB of HBM4 total memory, or 288 GB per GPU. Early yields are reportedly satisfactory, but volume availability will depend on how NVIDIA allocates initial HGX shipments. Cloud providers expecting Q4 2026 Rubin availability are already in allocation queues, and independent neoclouds will follow three to six months behind. Custom NVL72 rack configurations are tracking at 40-plus weeks from order.
For engineers running inference today, the software stack is quietly making the H-series cheaper to use as well. vLLM 0.21.0 shipped KV cache offload with a hybrid memory allocator that tiers KV cache between GPU and CPU memory, making 128K-context runs on a single H100 economically viable without paying for a second GPU just to hold the cache. SGLang 0.7 added speculative decoding improvements and multi-node tensor parallel support on H100 and H200, with radix attention cache changes that cut KV fragmentation by 15 to 30 percent on shared-prefix chatbot workloads. Older hardware plus smarter memory management is a real economic combination, not a marketing one.
The bottom line for anyone buying GPU time in 2026: spot is now the default for fault-tolerant workloads, the H-series is the value tier of the rental market, and the pricing pressure that pushed H100 to $1.46 per hour is only going to accelerate as HBM4 capacity comes online and Vera Rubin begins shipping to the hyperscalers. The compute gold rush is over; the efficiency era has begun, and the H-series is its workhorse.
Comments