H200 Rental
What an H200 Costs to Rent in 2026, and When It Beats the H100 and B200

The H200 rents for a median $4.48 an hour and carries 76 percent more memory than an H100. The same GPU runs from $3.59 to $13.78 depending on the cloud. We get it quoted across providers.

Get H200 Quotes Through Us
Single GPUs to eight-GPU nodes and clusters Every provider that fits, quoted at once Paid by the provider, never by you

H200 rental costs a median of $4.48 per GPU-hour on demand across the providers that list it, up about 26 percent in a year, according to GetDeploying, with specialist clouds at $3.59 to $4.59, marketplace spot from $1.98, and the hyperscalers at $10 to $13.78 for the same GPU. The NVIDIA H200 is the H100's Hopper chip with 141 GB of HBM3e memory instead of 80 and 4.8 TB/s of bandwidth instead of 3.35, so it runs the same software with no changes and holds models, batches and contexts that would need two H100s. That is why it has become the standard GPU for large-model inference: at about a dollar an hour more than an H100, it delivers roughly 32 GB of memory per dollar-hour against 21, the best in the data center tier. Prices are not falling with age: a transaction-based index of neocloud trades settled at $5.16 in late September, up 21 percent in a month.

ON-DEMAND MEDIAN
$4.48
per GPU-hour across 50 tracked providers, up 26 percent in a year (GetDeploying, Sept 29, 2026)
SPECIALIST CLOUDS
$3.59 to $4.59
per GPU-hour on demand; the benchmark to negotiate against (Thunder Compute, Spheron)
TRANSACTED INDEX
$5.16
per GPU-hour on real neocloud trades, up 21 percent in 30 days (Ornn, Sept 24, 2026)
THE SPREAD
$1.98 to $13.78
from marketplace spot to hyperscaler on demand for the same 141 GB GPU
What an H200 costs per hour, depending on where you rent it
US dollars per GPU-hour, September 2026
Azure and Google Cloud, on demandeight-GPU nodes only
$10.60 to $13.78
Transacted neocloud indexreal trades, not rate cards
$5.16
AWS P5e, on demandafter the 15 percent increase in January 2026
$4.98
Tracker median, all providers50 providers
$4.48
Specialist GPU clouds, on demandsingle GPUs available
$3.59 to $4.59
Marketplace and spotinterruptible; capacity can be reclaimed
$1.98 and up

Prices checked September 2026. Sources: GetDeploying, Ornn OCPI-H200 index, Thunder Compute, Spheron, GPU Finder. Free to cite with a link to this page.

Send us the requirement before you request quotes.
Once you have asked a provider for pricing directly, it will only work with you on its own terms. Send it to us first and you'll hear back within 24 hours from the person who will run your search, with every provider that fits quoted at once and benchmarked against the figures on this page. The provider you choose pays us, so it costs you nothing.
Get H200 Quotes Through Us

How Much Does It Cost to Rent an H200? H200 Cloud Pricing in 2026

FigurePriceSource
On-demand median, all providers$4.48 per GPU-hour, up 3 percent in 90 days and 26 percent in a yearGetDeploying, September 29, 2026
On-demand median, specialist GPU clouds$4.40 per GPU-hourAIMultiple, September 2026
Volume-weighted average, 20 providers$4.43 per GPU-hourMercatus GPU Index, September 13, 2026
Transacted neocloud index$5.16 per GPU-hour, up 21 percent in 30 daysOrnn OCPI-H200, September 24, 2026
Specialist cloud range$3.59 to $4.59; single-GPU access availableThunder Compute, Shattered
Hyperscalers, on demandAWS about $4.98; Azure about $10.60 to $13.78; Google Cloud about $10.87; eight-GPU nodes onlySpheron, AIMultiple
Cheapest confirmed in stock$1.98 per GPU-hour on a marketplace; eight-GPU node floor $2.00 per GPUGPU Finder, September 16, 2026
SpotAbout $1.99 to $2.88 per GPU-hourJarvislabs, Spheron
One GPU, one month, on demandAbout $3,270 at the medianMCA calculation
One eight-GPU node, one month, on demandAbout $26,000 at the median, with 1,128 GB of pooled memoryMCA calculation
Buying instead$30,000 to $40,000 per GPUJarvislabs, Thunder Compute
WHAT THAT MEANS
Two things stand out. First, the specialist clouds and AWS sit within a dollar of each other, while Azure and Google Cloud charge more than double; on the H200 the hyperscaler premium is the biggest single cost decision. Second, the trackers disagree on direction: listed prices are flat to rising slowly, while the transacted index jumped 21 percent in a month, which usually means capacity is tightening faster than the rate cards show. Both argue for locking a reserved rate at a specialist cloud now rather than waiting. H100 pricing is on H100 rental, Blackwell on B200 and B300 rental, and every GPU's current rate on GPU rental prices, updated monthly.

H200 vs H100 vs B200: Which One Should You Rent?

The H200 sits between the two on price and above both on memory per dollar. The right choice depends on one question: does the model, its batch and its context fit?

GPUOn-demand median, Sept 2026MemoryBandwidthWhere it wins
H100Hopper, 2023$3.2580 GB HBM33.35 TB/sAnything that fits in 80 GB: fine-tuning, smaller-model inference, cost-sensitive training
H200Hopper, 2024$4.48141 GB HBM3e4.8 TB/sMemory-bound inference: 70B to 400B-class models at FP8, long contexts, high concurrency; same code as the H100
B200Blackwell, 2025$6.49180 GB HBM3e8 TB/sTraining throughput and FP4 inference at scale; roughly 2.5 to 3 times the H100's training work per GPU

Sources: GetDeploying, AIMultiple, Thunder Compute, Jarvislabs.

Rent the H200 when

A model needs more than 80 GB but fits in 141, or you are serving long contexts and large batches where the KV cache, not the compute, is the limit. One H200 often replaces two H100s on inference, which makes $4.48 cheaper than $6.50. It also suits training runs where a 70B-class model fits on fewer GPUs.

Stay on the H100 when

Everything fits in 80 GB. Same architecture, same speed per GPU on compute-bound work, a dollar less per hour. There is no reason to pay for memory you will not use.

Step up to the B200 when

You are training at scale or serving at FP4, and time to result matters more than the rate. A B200 costs about $2 more per hour than an H200 but does far more training work, so on large runs it is cheaper per result.

Look at the B300 when

A single-node inference deployment needs more than 141 GB per GPU. The B300's 288 GB is the only way to serve some models on one card; see B200 and B300 rental.

WHY THE H200 IS THE INFERENCE WORKHORSE
Inference on large models is limited by memory: the model weights plus the KV cache for every concurrent request. An eight-GPU H200 node pools 1,128 GB, enough to serve a 400B-parameter model at FP8 on a single node, where an H100 node would need two nodes and the interconnect between them. Fewer GPUs, one node, no cross-node traffic: that is why the H200 is usually the cheapest way to serve a big model, even at a higher hourly rate.

H200 SXM vs H200 NVL

SpecificationH200 SXMH200 NVL (PCIe)
Memory and bandwidth141 GB HBM3e, 4.8 TB/s141 GB HBM3e, 4.8 TB/s, the same
GPU powerUp to 700 WUp to 600 W
InterconnectNVLink and NVSwitch across all eight GPUs at 900 GB/sNVLink bridges between two or four cards; PCIe otherwise
Typical serverEight-GPU HGX H200, about 10 kWTwo to eight cards in a standard server, about 4 to 8 kW
Rental useMulti-GPU inference and training; the standard rental nodeSingle or paired GPUs, cheaper servers, air-cooled enterprise racks

Single-GPU performance is the same; the SXM version scales far better past one GPU because every GPU talks to every other at full NVLink speed. For multi-node clusters of H200s, see GPU cluster rental.

On Demand, Reserved or Spot

On demand

By the hour, cancel anytime. Median $4.48; specialist clouds from $3.59. Right for variable work and for finding out what you need.

Reserved

One to twelve months or longer at a lower rate with guaranteed capacity. Discounts vary widely by provider and size; on committed terms for 16 GPUs and up, 25 to 40 percent off on demand is typical, and it is where negotiation earns its keep. With the transacted index rising, a reserved rate is also insurance.

Spot and marketplace

From about $1.98 to $2.88, interruptible. Fine for batch inference and fault-tolerant jobs; wrong for a production endpoint.

What the Hourly Rate Leaves Out

ItemWhat to check
Node minimumsAWS, Azure, Google Cloud and Oracle sell the H200 only in eight-GPU nodes; specialist clouds offer single GPUs. A team that needs two H200s pays for eight at a hyperscaler
EgressFree at most specialist clouds, metered at the hyperscalers; inference services move a lot of data out
StorageModel weights and datasets at per-GB-month rates; check what is included
Billing incrementPer-second, per-minute or per-hour; minimum runtimes of 20 minutes are common
AvailabilityListed is not in stock; several trackers now report which listings are actually available

Rent or Buy an H200 Server?

At $30,000 to $40,000 per GPU, an H200 rented at $4.48 an hour matches its purchase price after about 7,000 to 9,000 hours, roughly a year of full use. Owning then means hosting a 10 kW server.

One eight-GPU H200 server over three years
Running around the clock; owning uses a $320,000 HGX H200 server plus about $2,000 a month to host a 10 kW server in colocation
Rent on demand$4.48 per GPU-hour
About $940,000
Rent reservedabout 35 percent off on demand
About $610,000
Own and colocateserver plus hosting
About $390,000

MCA calculation from the figures on this page; excludes power consumption, staff and the residual value of the hardware. Free to cite with a link to this page.

Rent when

The deployment is under a year, the load is variable, or you want the option to move to Blackwell without owning a depreciating asset.

Own and colocate when

A production inference service will run for three years at high utilization and you can secure a facility for 10 kW servers, about 40 kW a rack air-cooled. The AI and GPU colocation guide covers which facilities can; we place those deployments too.

Bare metal in between

A dedicated H200 server on a monthly or annual term with no capital outlay. Closer to owning on price, closer to renting on commitment. See bare metal.

How We Get You a Better H200 Rate

You tell us the requirement: GPU count, term, the model or workload, storage, egress and where the data lives. We take it to every specialist GPU cloud (the neoclouds) and bare metal provider that fits, at the same time, so they compete; we benchmark the quotes against the figures on this page, including the transacted index rather than rate cards; and we negotiate the rate, the term, egress and availability before you sign. If the better answer is an H100 or a B200, or the provider you already use, we say so. The provider you choose pays us from the channel budget it would otherwise spend on its own sales team, so the service costs you nothing and the rate is not marked up. How it works.

1

Send the requirement first

Before you ask any provider directly; once you have, that provider will not work through an advisor on your deal.

2

Every provider that fits, quoted at once

Specialist clouds and bare metal providers, for your GPU count and term, single GPUs included.

3

Benchmarked and negotiated

Rate, term, egress, storage, node minimums and availability, compared as a full bill.

4

You choose

No obligation, and no cost to you at any stage.

Frequently Asked Questions

How much does it cost to rent an H200?

In September 2026 the median on-demand rate is $4.48 per GPU-hour across 50 tracked providers, according to GetDeploying, with specialist clouds at $3.59 to $4.59, spot from about $1.98, and the hyperscalers at $4.98 (AWS) to $13.78 (Azure). One H200 costs about $3,270 a month at the median; an eight-GPU node about $26,000. A transaction-based index of neocloud trades settled higher, at $5.16.

Is the H200 worth it over the H100?

Yes when the model, batch or context does not fit in 80 GB. The H200 has the same Hopper compute with 141 GB of memory and 4.8 TB/s of bandwidth, so one H200 often does the work of two H100s on inference, and at $4.48 against $3.25 that is cheaper per result. When everything fits in 80 GB, the H100 is the better price.

Should I rent an H200 or a B200?

For inference on models that fit in 141 GB, the H200 at $4.48 is usually cheaper than the B200 at $6.49. For training at scale, or FP4 inference where throughput matters more than the rate, the B200 does roughly two and a half to three times the work per GPU and wins per result. The B300, with 288 GB, is the answer when a single card needs more than 141 GB.

What is the difference between H200 SXM and H200 NVL?

Same memory and bandwidth. The SXM version draws up to 700 W and links all eight GPUs in an HGX server over NVLink and NVSwitch at 900 GB/s; it is the standard rental node. The NVL is a PCIe card at up to 600 W with NVLink bridges between two or four cards, for standard servers and single or paired GPUs.

Can I rent a single H200?

Yes, from specialist GPU clouds, several of which offer single-GPU H200 access on demand at $3.59 to $4.59 an hour. The hyperscalers sell the H200 only in eight-GPU nodes.

How much does an H200 cost on AWS?

AWS's P5e instance works out to about $4.98 per H200 per hour on demand after the 15 percent price increase of January 2026, in eight-GPU nodes only. That is close to specialist-cloud pricing and well below Azure and Google Cloud, which run about $10.60 to $13.78 per GPU-hour for the same chip.

How much memory does an H200 have and what can it run?

141 GB of HBM3e per GPU, 1,128 GB pooled across an eight-GPU node. A single H200 serves models up to roughly 70 billion parameters at FP16 or about 140 billion at FP8 with room for context; an eight-GPU node serves 400B-class models at FP8 on one node.

How much power does an H200 server need?

An eight-GPU HGX H200 server draws about 10 kW, the same as an HGX H100; four in a rack is about 40 kW, air-cooled but above what most facilities built for 5 to 15 kW racks can serve. The NVL card at up to 600 W fits standard enterprise servers.

Is it cheaper to rent or buy an H200 server?

Over three years at full utilization, owning an eight-GPU H200 server and hosting it costs about $390,000, against about $610,000 to rent on a reserved term and $940,000 on demand. Owning wins on long, steady production workloads with somewhere to host it; renting wins for shorter or variable work.

Are H200 prices going up or down?

Listed on-demand prices are flat to rising slowly (up 26 percent in the year to September 2026), while a transaction-based index of neocloud trades rose 21 percent in a single month to $5.16. Some trackers expect H200 rates to soften as Blackwell supply grows, but the trades say capacity is tight now, which favors reserving at a specialist cloud rather than waiting.

What does it cost to get H200 quotes through Metro Colo Advisory?

Nothing. The provider you choose pays us from its channel budget, the same way it pays its own sales team, and the rate is not marked up. If the best answer is a different GPU or the provider you already use, we say so.

Get H200 Quotes Through Us

Tell us the GPU count, the term, the model or workload, and when you need it. You'll hear back within 24 hours from the person who will run your search, with every provider that fits quoted at once and benchmarked against the rates on this page. If the best answer is a different GPU, or the provider you already use, we say so. No cost, no obligation, and the provider you choose pays us.

Send the requirement before you request quotes.
That is what keeps every provider competing for it.
Get H200 Quotes Through Us