GPU Cluster Rental
What H100, H200 and B200 Clusters Cost in 2026, How the Terms Work, and What to Negotiate Before You Commit

A 64-GPU H100 cluster runs about $1.8 million a year on demand, and a reserved term cuts that by a quarter to nearly half. We negotiate the rate, the delivery date and the exit terms.

Price a GPU Cluster Through Us
16 to 1,000-plus GPUs, reserved or on demand Every provider that can deliver, quoted at once Paid by the provider, never by you

GPU cluster rental means renting many GPU servers, usually eight GPUs each, joined by a high-speed InfiniBand or equivalent fabric so they train or serve as one machine, typically on a reserved term of three months to three years. It is where the money in AI infrastructure is: a 64-GPU H100 cluster costs about $5,000 a day on demand at the current $3.25 median, roughly $1.8 million a year, and reserved cluster terms run 25 to 45 percent below on-demand rates. A B200 cluster costs about double per hour and delivers two and a half to three times the training throughput, and the newest 36-month Blackwell commitments have been quoted as low as $2.25 per GPU-hour. The difference between a good cluster deal and a bad one is rarely the headline rate. It is whether the capacity exists on the date promised, what the interconnect actually is, what happens when nodes fail, and what you owe if your needs change.

64 H100s, ON DEMAND
~$5,000
a day at the $3.25 median, about $152,000 a month (MCA calculation)
RESERVED DISCOUNT
25 to 45%
below on demand on 3 to 12 month cluster terms of 16 GPUs and up (GPUPerHour)
B200 MEDIAN
$6.52
per GPU-hour on demand; 36-month reserved quoted as low as $2.25 (AIMultiple, IntuitionLabs)
THE RANGE
$1.49 to $12
per H100-hour across marketplaces, specialist clouds and hyperscalers (IntuitionLabs)
What a 64-GPU cluster costs per year, by GPU and term
Running around the clock; rental figures use September 2026 medians and typical reserved discounts, before storage and networking
B200, on demand$6.52 per GPU-hour median
About $3.7 million
H100, on demand$3.25 per GPU-hour median
About $1.8 million
B200, 36-month reservedas low as $2.25 per GPU-hour
About $1.3 million
H100, reservedabout 35 percent off on demand
About $1.2 million

MCA calculation; prices checked September 2026. Sources: AIMultiple GPU index, IntuitionLabs, GPUPerHour. A B200 does roughly two and a half to three times the training work of an H100, so compare per result, not per hour. Free to cite with a link to this page.

Send us the requirement before you request quotes.
Once you have asked a provider for cluster pricing directly, it will only work with you on its own terms. Send it to us first and you'll hear back within 24 hours from the person who will run your search, with every provider that can deliver quoted at once and benchmarked against the figures on this page. The provider you choose pays us, so it costs you nothing.
Price a GPU Cluster Through Us

GPU Cluster Pricing in 2026

Cluster pricing is quoted per GPU-hour, with the reserved rate depending on the GPU, the size and the term. These are the current benchmarks to negotiate against.

GPUOn demand, medianReserved cluster termsNotes
H100 SXM80 GB HBM3$3.25 per GPU-hourAbout $2.10 to $2.45 on 3 to 12 month terms at 25 to 45 percent off; 1-year contracts priced around $2.35 in March 2026Widest supply and the most providers to negotiate between; reserved discounts thin at the median but deep at cluster scale
H200141 GB HBM3e$4.40 per GPU-hourTypically 25 to 40 percent off on committed termsSame architecture as the H100 with 76 percent more memory; the choice for memory-bound inference clusters
B200180 GB HBM3e$6.52 per GPU-hour; $4.99 to $6.04 at specialist providers36-month contracts quoted as low as $2.25 per GPU-hourSupply-constrained and volatile: spot rose 48 percent between February and April 2026. Getting allocation matters more than the rate
B300288 GB HBM3eAbout $7.87 per GPU-hourReserved median about $5.71The current top of the HGX range; the largest reserved discount of the group
GB200 NVL7272 GPUs per rack, liquid cooledAbout $10.50 per GPU-hour at specialist clouds; $16 to $27 at hyperscalersQuoted case by case; rack-scale commitments onlySold by the rack, not the node; the domain for frontier-scale training

Sources: AIMultiple, September 2026; IntuitionLabs, July 2026; GPUPerHour, June 2026; ColoPrice, August 2026. Reserved ranges for the H100 are MCA calculations from the on-demand median and the published discount range.

WHAT THAT MEANS
Per hour, the H100 cluster is the cheapest; per result, the B200 cluster often wins, because each B200 does the work of roughly two and a half to three H100s. At $6.52 against $3.25 that is close to break-even on cost per token, and any reserved discount tips it toward Blackwell. The catch is allocation: B200 capacity is scarce enough that the question is usually not the rate but whether a provider can deliver 64 or 256 of them on your date. That is what a competitive search finds out. Model-by-model prices are on H100 rental, B200 and B300 rental and H200 rental, and every GPU's current rate is on GPU rental prices, updated monthly.

What Makes a Cluster, and What to Check in the Quote

Sixty-four GPUs on the same invoice are not a cluster. A cluster is nodes joined by a fabric fast enough that the GPUs spend their time computing rather than waiting for each other, and the quote should specify every layer.

LayerWhat a real cluster hasWhat to ask
The nodeEight GPUs in an HGX server with NVLink and NVSwitch between them, about 10 kW for H100 and higher for BlackwellGPU model and memory, CPU and system memory, local NVMe per node
The fabricDedicated InfiniBand (400 Gbps NDR is the current standard for H100 and Blackwell clusters) or RDMA-capable Ethernet, non-blocking, with one port per GPUFabric type, bandwidth per GPU, oversubscription ratio, and whether it is shared with other tenants
StorageParallel storage fast enough to feed the cluster and hold checkpoints, priced per GB-monthThroughput, capacity included, price per GB beyond it, and whether checkpoints count
OrchestrationKubernetes or Slurm, with the GPU drivers and communication libraries tuned to the fabricWhat is managed, what you run, and who fixes it at 3 a.m.
CapacityNodes that exist and are energized on the delivery date, plus sparesAre the nodes built? Where? What is the delivery date in the contract, and what is the remedy if it is missed?
THE NETWORK IS THE CLUSTER
A poorly built fabric can add 40 percent to the cost of a training run, because the GPUs sit idle while data moves. Two quotes at the same per-GPU rate, one with a dedicated non-blocking InfiniBand fabric and one with shared Ethernet, are not the same price. Compare cost per training run, not cost per GPU-hour.

Cluster Sizes and What They Are For

SizeTypical useMonthly cost, H100 on demandNotes
8 GPUsone nodeFine-tuning, inference serving, models to about 13B parameters trained from scratchAbout $19,000No fabric needed; NVLink inside the node does the work
16 to 32 GPUs2 to 4 nodesSerious fine-tuning, multi-model inference, training to about 30BAbout $38,000 to $76,000The smallest size most providers will reserve; the first point where the fabric matters
64 GPUs8 nodesPre-training models to about 70B; the most common reserved clusterAbout $152,000Reserved terms of 3 to 12 months typical; about 100 kW of power
256 GPUs32 nodesLarge-scale pre-training; the point past which returns diminish for most model sizesAbout $607,000Multi-year terms common; allocation, not price, is the constraint on Blackwell
1,000 and up125-plus nodesFrontier training$2.4 million and upUsually a dedicated build or rack-scale GB200 and GB300 NVL72 systems; often closer to a colocation decision than a rental

Monthly figures are MCA calculations at the $3.25 on-demand median, running around the clock, before storage and networking; reserved terms reduce them by 25 to 45 percent.

The Terms That Decide a Cluster Contract

On a $1 million-plus commitment, these move more money than the rate. Every one is negotiable before signature and almost none after.

Capacity guarantee and delivery date

The contract should name the GPU count, the fabric, the location and the date, with a remedy (credits, the right to walk) if the provider misses it. "Subject to availability" is not a guarantee.

Term and exit

Shorter terms cost more per hour but protect you from a chip generation change; a three-year H100 commitment signed today will overlap with cheaper Blackwell capacity, and Rubin follows it. Negotiate an upgrade path or a mid-term exit rather than the longest term for the lowest rate.

Ramp

Pay for GPUs as they come online and as you need them, not for the full cluster from day one. A ramp schedule in the contract is standard at 64 GPUs and up.

SLA and node replacement

GPU failures are routine at scale. Specify uptime, replacement time for a failed node, and credits that mean something, not a token percentage of one month's bill.

Storage, egress and networking charges

Confirm what is included, what is metered, and the price of moving your checkpoints and models out at the end. Most specialist clouds waive egress; the hyperscalers do not.

Expansion rights and price protection

A right of first refusal on adjacent capacity, and a cap on the rate for any expansion during the term, so growth does not reset the negotiation.

Rent a Cluster, or Own One and Colocate It?

At 64 GPUs the arithmetic starts to favor owning, if the cluster will run for three years and you can house it. Owning means buying eight HGX servers, a fabric and storage, and placing roughly 100 kW of load in a facility that can serve it.

A 64-GPU H100 cluster over three years
Running around the clock; owning includes eight HGX H100 servers at about $285,000 each, roughly $300,000 of fabric and storage, and 100 kW of colocation at about $204 per kW per month
Rent on demand$3.25 per GPU-hour
About $5.5 million
Rent reservedabout 35 percent off
About $3.6 million
Own and colocatehardware, fabric, storage and hosting
About $3.3 million

MCA calculation from the figures on this page and the data center cost page; excludes staff, power consumption and the residual value of the hardware. Free to cite with a link to this page.

Rent when

The project is under two years, the GPU generation is likely to change, or you cannot commit capital. Renting also avoids the two-to-four-year wait for new power that owning at scale can involve.

Own and colocate when

The cluster will run for three years or more at high utilization, you have the team to operate it, and you can secure a facility for 40 kW-plus air-cooled racks, or liquid cooling for rack-scale systems. The AI and GPU colocation guide covers which facilities can, and we place those deployments as well.

The hybrid

Own a base cluster for steady work and rent reserved capacity for training peaks. Increasingly the shape of mature AI infrastructure, and the one that keeps the most negotiating leverage with every provider.

Six Mistakes in GPU Cluster Deals

1

Comparing quotes on the per-GPU rate

Two clusters at $2.40 an hour with different fabrics can differ by 40 percent in cost per training run. Compare the whole configuration and the whole bill.

2

Signing the longest term for the lowest rate

Three years of H100 capacity at 2026 prices will look expensive in 2028. Take shorter terms, or a term with an upgrade path.

3

Accepting "subject to availability"

Without a dated delivery obligation and a remedy, you have a quote, not capacity. On Blackwell, this is the clause that matters most.

4

Paying for the whole cluster from day one

Ramp schedules are standard at scale. If a provider will not offer one, another will.

5

Ignoring the exit

Egress on petabytes of checkpoints, data deletion terms and end-of-term renewals should be settled before signature, when you still have alternatives.

6

Going to one provider first

Cluster pricing is negotiated, not listed, and providers price differently when they know they are competing. Once you have asked one directly, it will not work through an advisor on your deal, and the competitive tension is gone.

BUILDING YOUR OWN CLUSTER INSTEAD?
For a cluster that will run for years at high utilization, owning the hardware and placing it in colocation can beat renting, once you have the capital and the team to run it. The constraint is finding a hall with the power and cooling on your date. We place that too, at no cost to you: see AI and GPU colocation and wholesale colocation.

How We Run a Cluster Search

You give us the requirement: GPU model and count, the training or inference workload, term, start date, storage, networking and where the data lives. We take it to every specialist GPU cloud (the neoclouds) and bare metal provider that can deliver it, at the same time, so they compete; we verify that the capacity exists rather than taking a rate card's word for it; we benchmark the quotes against the figures on this page; and we negotiate the rate, the capacity guarantee, the ramp, the SLA and the exit before you sign. If owning and colocating is the better answer, we say so and run that search instead. The provider you choose pays us from the channel budget it would otherwise spend on its own sales team, so the service costs you nothing and the rate is not marked up. How it works.

Frequently Asked Questions

How much does it cost to rent a GPU cluster?

A 64-GPU H100 cluster costs about $5,000 a day on demand at the September 2026 median of $3.25 per GPU-hour, roughly $152,000 a month or $1.8 million a year, before storage and networking. Reserved cluster terms of three to twelve months run 25 to 45 percent below that. A B200 cluster costs about double per hour on demand, with 36-month commitments quoted as low as $2.25 per GPU-hour.

What is the minimum size for a reserved GPU cluster?

Most providers reserve from 16 GPUs (two nodes) upward, and the meaningful discounts start around 64 GPUs on terms of three months or more. Below 16 GPUs, on-demand or single-node bare metal is usually the better structure.

How much cheaper is a reserved GPU cluster than on demand?

Typically 25 to 45 percent per GPU-hour on three to twelve month terms, according to GPUPerHour, with deeper discounts on multi-year Blackwell commitments. The trade is flexibility: a long term at today's H100 rate will overlap with cheaper Blackwell capacity later, so shorter terms or upgrade rights are often worth more than the last few percent.

Should I rent an H100 cluster or a B200 cluster?

Per hour the H100 is half the price; per result the B200 often wins, because each B200 does roughly two and a half to three times the training work. At current medians the two are close on cost per token, and any reserved discount favors the B200. The real constraint is allocation: B200 capacity is scarce, so whether a provider can deliver on your date matters more than the rate.

What is InfiniBand and does my cluster need it?

InfiniBand is the high-bandwidth, low-latency network that links GPU nodes so they can train as one machine; 400 Gbps NDR is the current standard for H100 and Blackwell clusters. Any multi-node training workload needs it or an RDMA-capable Ethernet equivalent. Without it, GPUs sit idle waiting for data, which can add 40 percent to the cost of a run.

How much power does a GPU cluster need?

An eight-GPU HGX H100 server draws about 10 kW, so a 64-GPU cluster is about 100 kW including networking and storage, and Blackwell nodes draw more. A rented cluster is the provider's problem; an owned one needs a facility that can serve 40 kW-plus air-cooled racks, or liquid cooling for rack-scale GB200 and GB300 systems.

Is it cheaper to rent a GPU cluster or build one?

Over three years at high utilization, owning a 64-GPU H100 cluster and colocating it costs about $3.3 million against about $3.6 million to rent on a reserved term and $5.5 million on demand, before staff and power consumption. Owning wins on long, steady workloads with a team to run it; renting wins for shorter projects, uncertain demand, or when the GPU generation is about to change.

What should a GPU cluster contract include?

A named GPU count, fabric and location with a dated delivery obligation and a remedy for missing it; a ramp schedule; uptime and node replacement commitments with meaningful credits; the included storage and the price beyond it; egress terms for the end of the contract; expansion rights with price protection; and an exit or upgrade path if the term is longer than a year.

How long does it take to get a GPU cluster?

On-demand nodes are available in minutes to days for small sizes. Reserved H100 clusters of 64 GPUs are typically delivered in days to a few weeks where capacity is built. Blackwell clusters and anything above a few hundred GPUs depend on allocation and can take months, which is why the delivery date belongs in the contract.

Can Metro Colo Advisory get cluster pricing that I cannot get myself?

Usually, yes, because cluster pricing is negotiated rather than listed, and providers price differently when several are competing for the same requirement at once. We also negotiate the capacity guarantee, ramp, SLA and exit, which most buyers accept as drafted. The provider pays us, so the rate you get is not marked up.

What does it cost to run a cluster search through Metro Colo Advisory?

Nothing. The provider you choose pays us from its channel budget, the same way it pays its own sales team. If the best answer is to own and colocate, or to stay with your current provider, we say so.

Price a GPU Cluster Through Us

Tell us the GPU model and count, the workload, the term, the start date, and the storage and networking you need. You'll hear back within 24 hours from the person who will run your search, and every provider that can deliver will be quoted at once and benchmarked against the figures on this page. If the better answer is a different GPU, a different size, or owning and colocating, we say so. No cost, no obligation, and the provider you choose pays us.

Send the requirement before you request quotes.
That is what keeps every provider competing for it, and it is the only way to negotiate the terms as well as the rate.
Price a GPU Cluster Through Us

GPU rental: the GPU rental hub, B200 and B300 rental, H100 rental, H200 rental, GB200 and GB300 NVL72 rental, monthly GPU rental prices, neocloud providers, CoreWeave competitors and bare metal.

What comes next: NVIDIA Rubin and the GPU roadmap, and what it means for cluster terms signed today.

Owning and hosting: AI and GPU colocation, direct to chip cooling, wholesale colocation and data center cost.

The contract: the data center lease guide covers the capacity, delivery and exit terms that carry over to cluster contracts.