GPU Cluster Rental
What H100, H200 and B200 Clusters Cost in 2026, How the Terms Work, and What to Negotiate Before You Commit
A 64-GPU H100 cluster runs about $1.8 million a year on demand, and a reserved term cuts that by a quarter to nearly half. We negotiate the rate, the delivery date and the exit terms.
Price a GPU Cluster Through UsGPU cluster rental means renting many GPU servers, usually eight GPUs each, joined by a high-speed InfiniBand or equivalent fabric so they train or serve as one machine, typically on a reserved term of three months to three years. It is where the money in AI infrastructure is: a 64-GPU H100 cluster costs about $5,000 a day on demand at the current $3.25 median, roughly $1.8 million a year, and reserved cluster terms run 25 to 45 percent below on-demand rates. A B200 cluster costs about double per hour and delivers two and a half to three times the training throughput, and the newest 36-month Blackwell commitments have been quoted as low as $2.25 per GPU-hour. The difference between a good cluster deal and a bad one is rarely the headline rate. It is whether the capacity exists on the date promised, what the interconnect actually is, what happens when nodes fail, and what you owe if your needs change.
MCA calculation; prices checked September 2026. Sources: AIMultiple GPU index, IntuitionLabs, GPUPerHour. A B200 does roughly two and a half to three times the training work of an H100, so compare per result, not per hour. Free to cite with a link to this page.
Once you have asked a provider for cluster pricing directly, it will only work with you on its own terms. Send it to us first and you'll hear back within 24 hours from the person who will run your search, with every provider that can deliver quoted at once and benchmarked against the figures on this page. The provider you choose pays us, so it costs you nothing.
GPU Cluster Pricing in 2026
Cluster pricing is quoted per GPU-hour, with the reserved rate depending on the GPU, the size and the term. These are the current benchmarks to negotiate against.
| GPU | On demand, median | Reserved cluster terms | Notes |
|---|---|---|---|
| H100 SXM80 GB HBM3 | $3.25 per GPU-hour | About $2.10 to $2.45 on 3 to 12 month terms at 25 to 45 percent off; 1-year contracts priced around $2.35 in March 2026 | Widest supply and the most providers to negotiate between; reserved discounts thin at the median but deep at cluster scale |
| H200141 GB HBM3e | $4.40 per GPU-hour | Typically 25 to 40 percent off on committed terms | Same architecture as the H100 with 76 percent more memory; the choice for memory-bound inference clusters |
| B200180 GB HBM3e | $6.52 per GPU-hour; $4.99 to $6.04 at specialist providers | 36-month contracts quoted as low as $2.25 per GPU-hour | Supply-constrained and volatile: spot rose 48 percent between February and April 2026. Getting allocation matters more than the rate |
| B300288 GB HBM3e | About $7.87 per GPU-hour | Reserved median about $5.71 | The current top of the HGX range; the largest reserved discount of the group |
| GB200 NVL7272 GPUs per rack, liquid cooled | About $10.50 per GPU-hour at specialist clouds; $16 to $27 at hyperscalers | Quoted case by case; rack-scale commitments only | Sold by the rack, not the node; the domain for frontier-scale training |
Sources: AIMultiple, September 2026; IntuitionLabs, July 2026; GPUPerHour, June 2026; ColoPrice, August 2026. Reserved ranges for the H100 are MCA calculations from the on-demand median and the published discount range.
What Makes a Cluster, and What to Check in the Quote
Sixty-four GPUs on the same invoice are not a cluster. A cluster is nodes joined by a fabric fast enough that the GPUs spend their time computing rather than waiting for each other, and the quote should specify every layer.
| Layer | What a real cluster has | What to ask |
|---|---|---|
| The node | Eight GPUs in an HGX server with NVLink and NVSwitch between them, about 10 kW for H100 and higher for Blackwell | GPU model and memory, CPU and system memory, local NVMe per node |
| The fabric | Dedicated InfiniBand (400 Gbps NDR is the current standard for H100 and Blackwell clusters) or RDMA-capable Ethernet, non-blocking, with one port per GPU | Fabric type, bandwidth per GPU, oversubscription ratio, and whether it is shared with other tenants |
| Storage | Parallel storage fast enough to feed the cluster and hold checkpoints, priced per GB-month | Throughput, capacity included, price per GB beyond it, and whether checkpoints count |
| Orchestration | Kubernetes or Slurm, with the GPU drivers and communication libraries tuned to the fabric | What is managed, what you run, and who fixes it at 3 a.m. |
| Capacity | Nodes that exist and are energized on the delivery date, plus spares | Are the nodes built? Where? What is the delivery date in the contract, and what is the remedy if it is missed? |
Cluster Sizes and What They Are For
| Size | Typical use | Monthly cost, H100 on demand | Notes |
|---|---|---|---|
| 8 GPUsone node | Fine-tuning, inference serving, models to about 13B parameters trained from scratch | About $19,000 | No fabric needed; NVLink inside the node does the work |
| 16 to 32 GPUs2 to 4 nodes | Serious fine-tuning, multi-model inference, training to about 30B | About $38,000 to $76,000 | The smallest size most providers will reserve; the first point where the fabric matters |
| 64 GPUs8 nodes | Pre-training models to about 70B; the most common reserved cluster | About $152,000 | Reserved terms of 3 to 12 months typical; about 100 kW of power |
| 256 GPUs32 nodes | Large-scale pre-training; the point past which returns diminish for most model sizes | About $607,000 | Multi-year terms common; allocation, not price, is the constraint on Blackwell |
| 1,000 and up125-plus nodes | Frontier training | $2.4 million and up | Usually a dedicated build or rack-scale GB200 and GB300 NVL72 systems; often closer to a colocation decision than a rental |
Monthly figures are MCA calculations at the $3.25 on-demand median, running around the clock, before storage and networking; reserved terms reduce them by 25 to 45 percent.
The Terms That Decide a Cluster Contract
On a $1 million-plus commitment, these move more money than the rate. Every one is negotiable before signature and almost none after.
Capacity guarantee and delivery date
The contract should name the GPU count, the fabric, the location and the date, with a remedy (credits, the right to walk) if the provider misses it. "Subject to availability" is not a guarantee.
Term and exit
Shorter terms cost more per hour but protect you from a chip generation change; a three-year H100 commitment signed today will overlap with cheaper Blackwell capacity, and Rubin follows it. Negotiate an upgrade path or a mid-term exit rather than the longest term for the lowest rate.
Ramp
Pay for GPUs as they come online and as you need them, not for the full cluster from day one. A ramp schedule in the contract is standard at 64 GPUs and up.
SLA and node replacement
GPU failures are routine at scale. Specify uptime, replacement time for a failed node, and credits that mean something, not a token percentage of one month's bill.
Storage, egress and networking charges
Confirm what is included, what is metered, and the price of moving your checkpoints and models out at the end. Most specialist clouds waive egress; the hyperscalers do not.
Expansion rights and price protection
A right of first refusal on adjacent capacity, and a cap on the rate for any expansion during the term, so growth does not reset the negotiation.
Rent a Cluster, or Own One and Colocate It?
At 64 GPUs the arithmetic starts to favor owning, if the cluster will run for three years and you can house it. Owning means buying eight HGX servers, a fabric and storage, and placing roughly 100 kW of load in a facility that can serve it.
MCA calculation from the figures on this page and the data center cost page; excludes staff, power consumption and the residual value of the hardware. Free to cite with a link to this page.
Rent when
The project is under two years, the GPU generation is likely to change, or you cannot commit capital. Renting also avoids the two-to-four-year wait for new power that owning at scale can involve.
Own and colocate when
The cluster will run for three years or more at high utilization, you have the team to operate it, and you can secure a facility for 40 kW-plus air-cooled racks, or liquid cooling for rack-scale systems. The AI and GPU colocation guide covers which facilities can, and we place those deployments as well.
The hybrid
Own a base cluster for steady work and rent reserved capacity for training peaks. Increasingly the shape of mature AI infrastructure, and the one that keeps the most negotiating leverage with every provider.
Six Mistakes in GPU Cluster Deals
Comparing quotes on the per-GPU rate
Two clusters at $2.40 an hour with different fabrics can differ by 40 percent in cost per training run. Compare the whole configuration and the whole bill.
Signing the longest term for the lowest rate
Three years of H100 capacity at 2026 prices will look expensive in 2028. Take shorter terms, or a term with an upgrade path.
Accepting "subject to availability"
Without a dated delivery obligation and a remedy, you have a quote, not capacity. On Blackwell, this is the clause that matters most.
Paying for the whole cluster from day one
Ramp schedules are standard at scale. If a provider will not offer one, another will.
Ignoring the exit
Egress on petabytes of checkpoints, data deletion terms and end-of-term renewals should be settled before signature, when you still have alternatives.
Going to one provider first
Cluster pricing is negotiated, not listed, and providers price differently when they know they are competing. Once you have asked one directly, it will not work through an advisor on your deal, and the competitive tension is gone.
For a cluster that will run for years at high utilization, owning the hardware and placing it in colocation can beat renting, once you have the capital and the team to run it. The constraint is finding a hall with the power and cooling on your date. We place that too, at no cost to you: see AI and GPU colocation and wholesale colocation.
How We Run a Cluster Search
You give us the requirement: GPU model and count, the training or inference workload, term, start date, storage, networking and where the data lives. We take it to every specialist GPU cloud (the neoclouds) and bare metal provider that can deliver it, at the same time, so they compete; we verify that the capacity exists rather than taking a rate card's word for it; we benchmark the quotes against the figures on this page; and we negotiate the rate, the capacity guarantee, the ramp, the SLA and the exit before you sign. If owning and colocating is the better answer, we say so and run that search instead. The provider you choose pays us from the channel budget it would otherwise spend on its own sales team, so the service costs you nothing and the rate is not marked up. How it works.
Frequently Asked Questions
How much does it cost to rent a GPU cluster?
A 64-GPU H100 cluster costs about $5,000 a day on demand at the September 2026 median of $3.25 per GPU-hour, roughly $152,000 a month or $1.8 million a year, before storage and networking. Reserved cluster terms of three to twelve months run 25 to 45 percent below that. A B200 cluster costs about double per hour on demand, with 36-month commitments quoted as low as $2.25 per GPU-hour.
What is the minimum size for a reserved GPU cluster?
Most providers reserve from 16 GPUs (two nodes) upward, and the meaningful discounts start around 64 GPUs on terms of three months or more. Below 16 GPUs, on-demand or single-node bare metal is usually the better structure.
How much cheaper is a reserved GPU cluster than on demand?
Typically 25 to 45 percent per GPU-hour on three to twelve month terms, according to GPUPerHour, with deeper discounts on multi-year Blackwell commitments. The trade is flexibility: a long term at today's H100 rate will overlap with cheaper Blackwell capacity later, so shorter terms or upgrade rights are often worth more than the last few percent.
Should I rent an H100 cluster or a B200 cluster?
Per hour the H100 is half the price; per result the B200 often wins, because each B200 does roughly two and a half to three times the training work. At current medians the two are close on cost per token, and any reserved discount favors the B200. The real constraint is allocation: B200 capacity is scarce, so whether a provider can deliver on your date matters more than the rate.
What is InfiniBand and does my cluster need it?
InfiniBand is the high-bandwidth, low-latency network that links GPU nodes so they can train as one machine; 400 Gbps NDR is the current standard for H100 and Blackwell clusters. Any multi-node training workload needs it or an RDMA-capable Ethernet equivalent. Without it, GPUs sit idle waiting for data, which can add 40 percent to the cost of a run.
How much power does a GPU cluster need?
An eight-GPU HGX H100 server draws about 10 kW, so a 64-GPU cluster is about 100 kW including networking and storage, and Blackwell nodes draw more. A rented cluster is the provider's problem; an owned one needs a facility that can serve 40 kW-plus air-cooled racks, or liquid cooling for rack-scale GB200 and GB300 systems.
Is it cheaper to rent a GPU cluster or build one?
Over three years at high utilization, owning a 64-GPU H100 cluster and colocating it costs about $3.3 million against about $3.6 million to rent on a reserved term and $5.5 million on demand, before staff and power consumption. Owning wins on long, steady workloads with a team to run it; renting wins for shorter projects, uncertain demand, or when the GPU generation is about to change.
What should a GPU cluster contract include?
A named GPU count, fabric and location with a dated delivery obligation and a remedy for missing it; a ramp schedule; uptime and node replacement commitments with meaningful credits; the included storage and the price beyond it; egress terms for the end of the contract; expansion rights with price protection; and an exit or upgrade path if the term is longer than a year.
How long does it take to get a GPU cluster?
On-demand nodes are available in minutes to days for small sizes. Reserved H100 clusters of 64 GPUs are typically delivered in days to a few weeks where capacity is built. Blackwell clusters and anything above a few hundred GPUs depend on allocation and can take months, which is why the delivery date belongs in the contract.
Can Metro Colo Advisory get cluster pricing that I cannot get myself?
Usually, yes, because cluster pricing is negotiated rather than listed, and providers price differently when several are competing for the same requirement at once. We also negotiate the capacity guarantee, ramp, SLA and exit, which most buyers accept as drafted. The provider pays us, so the rate you get is not marked up.
What does it cost to run a cluster search through Metro Colo Advisory?
Nothing. The provider you choose pays us from its channel budget, the same way it pays its own sales team. If the best answer is to own and colocate, or to stay with your current provider, we say so.
Price a GPU Cluster Through Us
Tell us the GPU model and count, the workload, the term, the start date, and the storage and networking you need. You'll hear back within 24 hours from the person who will run your search, and every provider that can deliver will be quoted at once and benchmarked against the figures on this page. If the better answer is a different GPU, a different size, or owning and colocating, we say so. No cost, no obligation, and the provider you choose pays us.
That is what keeps every provider competing for it, and it is the only way to negotiate the terms as well as the rate.
GPU rental: the GPU rental hub, B200 and B300 rental, H100 rental, H200 rental, GB200 and GB300 NVL72 rental, monthly GPU rental prices, neocloud providers, CoreWeave competitors and bare metal.
What comes next: NVIDIA Rubin and the GPU roadmap, and what it means for cluster terms signed today.
Owning and hosting: AI and GPU colocation, direct to chip cooling, wholesale colocation and data center cost.
The contract: the data center lease guide covers the capacity, delivery and exit terms that carry over to cluster contracts.