H200 Rental
What an H200 Costs to Rent in 2026, and When It Beats the H100 and B200
The H200 rents for a median $4.48 an hour and carries 76 percent more memory than an H100. The same GPU runs from $3.59 to $13.78 depending on the cloud. We get it quoted across providers.
Get H200 Quotes Through UsH200 rental costs a median of $4.48 per GPU-hour on demand across the providers that list it, up about 26 percent in a year, according to GetDeploying, with specialist clouds at $3.59 to $4.59, marketplace spot from $1.98, and the hyperscalers at $10 to $13.78 for the same GPU. The NVIDIA H200 is the H100's Hopper chip with 141 GB of HBM3e memory instead of 80 and 4.8 TB/s of bandwidth instead of 3.35, so it runs the same software with no changes and holds models, batches and contexts that would need two H100s. That is why it has become the standard GPU for large-model inference: at about a dollar an hour more than an H100, it delivers roughly 32 GB of memory per dollar-hour against 21, the best in the data center tier. Prices are not falling with age: a transaction-based index of neocloud trades settled at $5.16 in late September, up 21 percent in a month.
Prices checked September 2026. Sources: GetDeploying, Ornn OCPI-H200 index, Thunder Compute, Spheron, GPU Finder. Free to cite with a link to this page.
Once you have asked a provider for pricing directly, it will only work with you on its own terms. Send it to us first and you'll hear back within 24 hours from the person who will run your search, with every provider that fits quoted at once and benchmarked against the figures on this page. The provider you choose pays us, so it costs you nothing.
How Much Does It Cost to Rent an H200? H200 Cloud Pricing in 2026
| Figure | Price | Source |
|---|---|---|
| On-demand median, all providers | $4.48 per GPU-hour, up 3 percent in 90 days and 26 percent in a year | GetDeploying, September 29, 2026 |
| On-demand median, specialist GPU clouds | $4.40 per GPU-hour | AIMultiple, September 2026 |
| Volume-weighted average, 20 providers | $4.43 per GPU-hour | Mercatus GPU Index, September 13, 2026 |
| Transacted neocloud index | $5.16 per GPU-hour, up 21 percent in 30 days | Ornn OCPI-H200, September 24, 2026 |
| Specialist cloud range | $3.59 to $4.59; single-GPU access available | Thunder Compute, Shattered |
| Hyperscalers, on demand | AWS about $4.98; Azure about $10.60 to $13.78; Google Cloud about $10.87; eight-GPU nodes only | Spheron, AIMultiple |
| Cheapest confirmed in stock | $1.98 per GPU-hour on a marketplace; eight-GPU node floor $2.00 per GPU | GPU Finder, September 16, 2026 |
| Spot | About $1.99 to $2.88 per GPU-hour | Jarvislabs, Spheron |
| One GPU, one month, on demand | About $3,270 at the median | MCA calculation |
| One eight-GPU node, one month, on demand | About $26,000 at the median, with 1,128 GB of pooled memory | MCA calculation |
| Buying instead | $30,000 to $40,000 per GPU | Jarvislabs, Thunder Compute |
H200 vs H100 vs B200: Which One Should You Rent?
The H200 sits between the two on price and above both on memory per dollar. The right choice depends on one question: does the model, its batch and its context fit?
| GPU | On-demand median, Sept 2026 | Memory | Bandwidth | Where it wins |
|---|---|---|---|---|
| H100Hopper, 2023 | $3.25 | 80 GB HBM3 | 3.35 TB/s | Anything that fits in 80 GB: fine-tuning, smaller-model inference, cost-sensitive training |
| H200Hopper, 2024 | $4.48 | 141 GB HBM3e | 4.8 TB/s | Memory-bound inference: 70B to 400B-class models at FP8, long contexts, high concurrency; same code as the H100 |
| B200Blackwell, 2025 | $6.49 | 180 GB HBM3e | 8 TB/s | Training throughput and FP4 inference at scale; roughly 2.5 to 3 times the H100's training work per GPU |
Sources: GetDeploying, AIMultiple, Thunder Compute, Jarvislabs.
Rent the H200 when
A model needs more than 80 GB but fits in 141, or you are serving long contexts and large batches where the KV cache, not the compute, is the limit. One H200 often replaces two H100s on inference, which makes $4.48 cheaper than $6.50. It also suits training runs where a 70B-class model fits on fewer GPUs.
Stay on the H100 when
Everything fits in 80 GB. Same architecture, same speed per GPU on compute-bound work, a dollar less per hour. There is no reason to pay for memory you will not use.
Step up to the B200 when
You are training at scale or serving at FP4, and time to result matters more than the rate. A B200 costs about $2 more per hour than an H200 but does far more training work, so on large runs it is cheaper per result.
Look at the B300 when
A single-node inference deployment needs more than 141 GB per GPU. The B300's 288 GB is the only way to serve some models on one card; see B200 and B300 rental.
H200 SXM vs H200 NVL
| Specification | H200 SXM | H200 NVL (PCIe) |
|---|---|---|
| Memory and bandwidth | 141 GB HBM3e, 4.8 TB/s | 141 GB HBM3e, 4.8 TB/s, the same |
| GPU power | Up to 700 W | Up to 600 W |
| Interconnect | NVLink and NVSwitch across all eight GPUs at 900 GB/s | NVLink bridges between two or four cards; PCIe otherwise |
| Typical server | Eight-GPU HGX H200, about 10 kW | Two to eight cards in a standard server, about 4 to 8 kW |
| Rental use | Multi-GPU inference and training; the standard rental node | Single or paired GPUs, cheaper servers, air-cooled enterprise racks |
Single-GPU performance is the same; the SXM version scales far better past one GPU because every GPU talks to every other at full NVLink speed. For multi-node clusters of H200s, see GPU cluster rental.
On Demand, Reserved or Spot
On demand
By the hour, cancel anytime. Median $4.48; specialist clouds from $3.59. Right for variable work and for finding out what you need.
Reserved
One to twelve months or longer at a lower rate with guaranteed capacity. Discounts vary widely by provider and size; on committed terms for 16 GPUs and up, 25 to 40 percent off on demand is typical, and it is where negotiation earns its keep. With the transacted index rising, a reserved rate is also insurance.
Spot and marketplace
From about $1.98 to $2.88, interruptible. Fine for batch inference and fault-tolerant jobs; wrong for a production endpoint.
What the Hourly Rate Leaves Out
| Item | What to check |
|---|---|
| Node minimums | AWS, Azure, Google Cloud and Oracle sell the H200 only in eight-GPU nodes; specialist clouds offer single GPUs. A team that needs two H200s pays for eight at a hyperscaler |
| Egress | Free at most specialist clouds, metered at the hyperscalers; inference services move a lot of data out |
| Storage | Model weights and datasets at per-GB-month rates; check what is included |
| Billing increment | Per-second, per-minute or per-hour; minimum runtimes of 20 minutes are common |
| Availability | Listed is not in stock; several trackers now report which listings are actually available |
Rent or Buy an H200 Server?
At $30,000 to $40,000 per GPU, an H200 rented at $4.48 an hour matches its purchase price after about 7,000 to 9,000 hours, roughly a year of full use. Owning then means hosting a 10 kW server.
MCA calculation from the figures on this page; excludes power consumption, staff and the residual value of the hardware. Free to cite with a link to this page.
Rent when
The deployment is under a year, the load is variable, or you want the option to move to Blackwell without owning a depreciating asset.
Own and colocate when
A production inference service will run for three years at high utilization and you can secure a facility for 10 kW servers, about 40 kW a rack air-cooled. The AI and GPU colocation guide covers which facilities can; we place those deployments too.
Bare metal in between
A dedicated H200 server on a monthly or annual term with no capital outlay. Closer to owning on price, closer to renting on commitment. See bare metal.
How We Get You a Better H200 Rate
You tell us the requirement: GPU count, term, the model or workload, storage, egress and where the data lives. We take it to every specialist GPU cloud (the neoclouds) and bare metal provider that fits, at the same time, so they compete; we benchmark the quotes against the figures on this page, including the transacted index rather than rate cards; and we negotiate the rate, the term, egress and availability before you sign. If the better answer is an H100 or a B200, or the provider you already use, we say so. The provider you choose pays us from the channel budget it would otherwise spend on its own sales team, so the service costs you nothing and the rate is not marked up. How it works.
Send the requirement first
Before you ask any provider directly; once you have, that provider will not work through an advisor on your deal.
Every provider that fits, quoted at once
Specialist clouds and bare metal providers, for your GPU count and term, single GPUs included.
Benchmarked and negotiated
Rate, term, egress, storage, node minimums and availability, compared as a full bill.
You choose
No obligation, and no cost to you at any stage.
Frequently Asked Questions
How much does it cost to rent an H200?
In September 2026 the median on-demand rate is $4.48 per GPU-hour across 50 tracked providers, according to GetDeploying, with specialist clouds at $3.59 to $4.59, spot from about $1.98, and the hyperscalers at $4.98 (AWS) to $13.78 (Azure). One H200 costs about $3,270 a month at the median; an eight-GPU node about $26,000. A transaction-based index of neocloud trades settled higher, at $5.16.
Is the H200 worth it over the H100?
Yes when the model, batch or context does not fit in 80 GB. The H200 has the same Hopper compute with 141 GB of memory and 4.8 TB/s of bandwidth, so one H200 often does the work of two H100s on inference, and at $4.48 against $3.25 that is cheaper per result. When everything fits in 80 GB, the H100 is the better price.
Should I rent an H200 or a B200?
For inference on models that fit in 141 GB, the H200 at $4.48 is usually cheaper than the B200 at $6.49. For training at scale, or FP4 inference where throughput matters more than the rate, the B200 does roughly two and a half to three times the work per GPU and wins per result. The B300, with 288 GB, is the answer when a single card needs more than 141 GB.
What is the difference between H200 SXM and H200 NVL?
Same memory and bandwidth. The SXM version draws up to 700 W and links all eight GPUs in an HGX server over NVLink and NVSwitch at 900 GB/s; it is the standard rental node. The NVL is a PCIe card at up to 600 W with NVLink bridges between two or four cards, for standard servers and single or paired GPUs.
Can I rent a single H200?
Yes, from specialist GPU clouds, several of which offer single-GPU H200 access on demand at $3.59 to $4.59 an hour. The hyperscalers sell the H200 only in eight-GPU nodes.
How much does an H200 cost on AWS?
AWS's P5e instance works out to about $4.98 per H200 per hour on demand after the 15 percent price increase of January 2026, in eight-GPU nodes only. That is close to specialist-cloud pricing and well below Azure and Google Cloud, which run about $10.60 to $13.78 per GPU-hour for the same chip.
How much memory does an H200 have and what can it run?
141 GB of HBM3e per GPU, 1,128 GB pooled across an eight-GPU node. A single H200 serves models up to roughly 70 billion parameters at FP16 or about 140 billion at FP8 with room for context; an eight-GPU node serves 400B-class models at FP8 on one node.
How much power does an H200 server need?
An eight-GPU HGX H200 server draws about 10 kW, the same as an HGX H100; four in a rack is about 40 kW, air-cooled but above what most facilities built for 5 to 15 kW racks can serve. The NVL card at up to 600 W fits standard enterprise servers.
Is it cheaper to rent or buy an H200 server?
Over three years at full utilization, owning an eight-GPU H200 server and hosting it costs about $390,000, against about $610,000 to rent on a reserved term and $940,000 on demand. Owning wins on long, steady production workloads with somewhere to host it; renting wins for shorter or variable work.
Are H200 prices going up or down?
Listed on-demand prices are flat to rising slowly (up 26 percent in the year to September 2026), while a transaction-based index of neocloud trades rose 21 percent in a single month to $5.16. Some trackers expect H200 rates to soften as Blackwell supply grows, but the trades say capacity is tight now, which favors reserving at a specialist cloud rather than waiting.
What does it cost to get H200 quotes through Metro Colo Advisory?
Nothing. The provider you choose pays us from its channel budget, the same way it pays its own sales team, and the rate is not marked up. If the best answer is a different GPU or the provider you already use, we say so.
Get H200 Quotes Through Us
Tell us the GPU count, the term, the model or workload, and when you need it. You'll hear back within 24 hours from the person who will run your search, with every provider that fits quoted at once and benchmarked against the rates on this page. If the best answer is a different GPU, or the provider you already use, we say so. No cost, no obligation, and the provider you choose pays us.
That is what keeps every provider competing for it.
GPU rental: the GPU rental hub, GPU cluster rental, B200 and B300 rental, H100 rental, GB200 and GB300 NVL72 rental, monthly GPU rental prices, neocloud providers, CoreWeave competitors and bare metal.
What comes next: NVIDIA Rubin and the GPU roadmap, and what it means for Hopper prices.
Owning and hosting: AI and GPU colocation, direct to chip cooling and data center cost.