H100 Rental
What an NVIDIA H100 Costs to Rent in 2026, On Demand and Reserved, and How to Get the Best Rate
The same H100 rents for $1.49 an hour on one cloud and $6.98 on another; the specialist-cloud median is $3.25. We benchmark your requirement and negotiate the terms.
Get H100 Quotes Through UsH100 rental costs a median of $3.25 per GPU-hour on demand across specialist GPU clouds, with reserved terms at about $3.10 and published rates running from $1.49 on marketplaces to $6.98 at the hyperscalers, according to AIMultiple and CloudZero. That spread is the whole story of renting an H100: the chip is identical everywhere, so the price you pay is decided by which provider you ask, what term you commit to, and how well you negotiate. Two years after launch the H100 is the price-performance workhorse of AI infrastructure: a mature software stack, the widest supply of any data center GPU, and rates that have held or risen through 2026 because demand for any working GPU capacity still outruns supply.
Prices checked September 2026. Sources: AIMultiple GPU index, GetDeploying, CloudZero. Free to cite with a link to this page.
Once you have asked a provider for pricing directly, it will only work with you on its own terms. Send it to us first and you'll hear back within 24 hours from the person who will run your search, with every provider that fits quoted at once and benchmarked against the medians above. The provider you choose pays us, so it costs you nothing.
How Much Does It Cost to Rent an H100?
Hourly rates are the headline, but the bill depends on the term, the provider tier and what surrounds the GPU. These are the September 2026 figures.
| Figure | Price | Source |
|---|---|---|
| On-demand median, specialist GPU clouds | $3.25 per GPU-hour | AIMultiple, September 2026 |
| Reserved median, specialist GPU clouds | $3.10 per GPU-hour, a $0.13 discount to on demand | AIMultiple, September 2026 |
| Median across 58 tracked providers | $3.49 per GPU-hour, flat over 90 days, up about 12 percent in a year | GetDeploying, September 29, 2026 |
| Published on-demand range | $1.49 to $6.98 per GPU-hour depending on provider | CloudZero, August 2026 |
| AWS, reserved effective rate | About $5.19 per H100 per hour in US regions on the P5 instance | IntuitionLabs, August 2026 |
| Spot rental benchmark | About $2.77 to $2.80 per hour, up roughly a third in a year | Bank of America research, via Shattered, August 2026 |
| One GPU, one month, on demand | About $2,370 at the $3.25 median; about $2,900 at $3.99 | MCA calculation; CloudZero |
| One eight-GPU server, one month, on demand | About $19,000 at the median, running around the clock | MCA calculation |
| Buying instead: one H100 | $25,000 to $40,000; an eight-GPU HGX server $250,000 to $320,000 | GPU Smith, CloudZero |
Every GPU's price, updated monthly, is on GPU rental prices; the specialist clouds are compared on CoreWeave competitors and explained on neocloud providers; and the GPU rental hub covers how renting through us works.
H100 vs H200 vs B200: Which One Should You Rent?
The H100 is no longer the top of the range, and that is exactly why it is often the right rental. Newer chips cost more per hour; whether they cost less per job depends on the workload.
| GPU | On-demand median, Sept 2026 | Memory | Where it wins |
|---|---|---|---|
| H100Hopper, 2023 | $3.25 per GPU-hour | 80 GB HBM3 | Cost-sensitive training and inference with a mature toolchain; the widest supply and the most providers to negotiate between |
| H200Hopper, 2024 | $4.40 per GPU-hour | 141 GB HBM3e | Memory-bound inference on large models; same architecture as the H100, so code moves without changes |
| B200Blackwell, 2025 | About $4 to $5.50 on dedicated GPU clouds; more at hyperscalers | 180 GB HBM3e | Training throughput per result; roughly the work of two and a half to three H100s per GPU on large training runs |
| A100Ampere, 2020 | $1.76 to $1.79 per GPU-hour; from $1.09 | 40 or 80 GB | Fine-tuning, inference under about 30B parameters, and development work, at half the H100 price |
Sources: AIMultiple, GetDeploying, CloudZero, Thunder Compute.
Rent the H100 when
Your code already runs on Hopper, the models fit in 80 GB, and cost per hour matters more than the fastest possible run. At $3.25 against $4.40 for the H200, the H100 is cheaper per job for most inference and fine-tuning.
Step up to the H200 when
The model or batch does not fit in 80 GB and you are paying for a second GPU just for memory. The H200's 141 GB often lets one GPU do the work of two H100s on inference, which pays for the premium. See H200 rental.
Step up to the B200 when
You are training at scale and the metric is time to result. Each B200 displaces several H100s, and B200 rental at $4 to $5.50 on the dedicated clouds is no longer far above the H100. See B200 and B300 rental for current rates, and GB200 NVL72 for rack-scale systems.
Step down to the A100 when
The work is LoRA fine-tuning, sub-30B inference or development. The A100 lacks FP8 and the Transformer Engine, so the H100 pulls ahead on large-model inference, but for everything else the A100 at about half the price is usually cheaper per job.
H100 SXM vs H100 PCIe, and What an H100 Server Needs
Multi-node deployments of 16 GPUs and up are covered on GPU cluster rental. Most single-server rentals are the SXM version in an eight-GPU HGX server; the PCIe card appears in smaller and cheaper configurations. The difference matters for performance and for what a facility can host.
| Specification | H100 SXM | H100 PCIe |
|---|---|---|
| Memory | 80 GB HBM3, about 3.35 TB/s | 80 GB HBM2e, about 2 TB/s |
| GPU power | Up to 700 W | Up to 350 W |
| Interconnect | NVLink and NVSwitch across all eight GPUs in an HGX server | PCIe, with NVLink bridges between pairs |
| Typical server | Eight-GPU HGX H100, about 10 kW per server | Two to eight cards, about 3 to 6 kW per server |
| Rental use | Multi-GPU training and high-throughput inference; the standard rental unit | Single-GPU inference, development, cost-sensitive work |
On Demand, Reserved or Spot: How the Term Changes the Price
On demand
Pay by the hour, cancel anytime, no commitment. The right choice for variable or short work, and the anchor every other price is measured against. Median $3.25 on the specialist clouds.
Reserved
A one to twelve month (or longer) commitment for a lower rate and guaranteed capacity. The median discount is small this year, about 4 percent, but individual providers go much further for larger or longer commitments, which is where negotiation earns its keep.
Spot and marketplace
Interruptible capacity at roughly half the on-demand rate, sometimes less. Fine for fault-tolerant batch work; wrong for anything that cannot be paused, since the provider can reclaim it with little warning.
What the Hourly Rate Leaves Out
Two quotes at the same rate can differ by a third once the surrounding costs are added. Compare the whole bill.
| Item | What to check |
|---|---|
| Storage | Per-GB-month charges for datasets and checkpoints; high-performance storage can cost as much as the GPUs on data-heavy work |
| Data transfer | Egress fees for moving results and models out; the hyperscalers charge, most specialist clouds do not |
| Networking between servers | Whether multi-node training gets InfiniBand or fast Ethernet, or shares a general network that stalls large jobs |
| Minimums and billing increments | Per-minute, per-hour or per-day billing; minimum GPU counts; minimum terms |
| Availability | Whether the quoted capacity exists now; several trackers report which listings are actually in stock |
| Support and managed services | What is included, what is billed hourly, and response times in the contract |
Rent or Buy an H100?
At full utilization, renting costs more than owning within about a year: an H100 rented at $3.99 an hour costs about $35,000 a year, more than the card itself. But most hardware does not run at full utilization, and owning means finding somewhere to put it.
MCA calculation from the figures on this page. Owning wins when the hardware runs most of the time and you can host it; renting wins for short or variable work, and reserved terms sit in between. Free to cite with a link to this page.
Rent when
The work is under a year, variable, or you are not sure the workload will hold. You pay a premium per hour for not being wrong at scale.
Buy and colocate when
The hardware will run most of the time for three years or more and you have, or can get, a facility that serves 40 kW racks. The data center cost page has the full arithmetic.
Bare metal in between
A dedicated GPU server on a monthly or annual term with no capital outlay: closer to owning on price, closer to renting on commitment. See bare metal.
If the workload will run for three years or more at high utilization, owning H100 servers and hosting them can cost less than renting. H100 servers draw about 10 kW each, around 40 kW a rack air-cooled, which many colocation halls can take. We place that too, at no cost to you: see AI and GPU colocation.
How We Get You a Better H100 Rate
The spread between $1.49 and $6.98 for the same GPU exists because most buyers ask one or two providers and take the answer. We run it differently. You tell us the requirement: GPU count, term, storage, networking and where the data lives. We take it to every specialist cloud and bare metal provider that fits, at the same time, so they compete; we benchmark the quotes against the medians on this page; and we negotiate the rate, the term, the storage and egress charges and the availability guarantee before you sign. The provider you choose pays us from the channel budget it would otherwise spend on its own sales team, so the service costs you nothing and the rate is not marked up. How it works.
Send the requirement first
Before you ask any provider directly. Once you have, that provider will not work through an advisor on your deal, and you negotiate alone.
Every provider that fits, quoted at once
Specialist GPU clouds and bare metal providers, for your GPU count and term, within days for standard configurations.
Benchmarked and negotiated
Rate, term, storage, egress, networking and availability, compared as a full bill, not an hourly headline.
You choose
If the best answer is the provider you already use, or a different GPU, we say so.
Frequently Asked Questions
How much does it cost to rent an H100?
In September 2026 the median on-demand rate across specialist GPU clouds is $3.25 per GPU-hour, with reserved terms at about $3.10 and published rates ranging from $1.49 on marketplaces to $6.98 at the hyperscalers. Running one H100 around the clock for a month costs about $2,370 at the median; an eight-GPU server about $19,000.
How much does an H100 cost per hour on AWS?
AWS's P5 instance works out to about $5.19 per H100 per hour as an effective reserved rate in US regions, according to IntuitionLabs's August 2026 comparison, after AWS cut GPU instance pricing by up to 45 percent in mid-2025. On demand it is higher. Specialist GPU clouds price the same GPU at roughly half.
Is it cheaper to rent or buy an H100?
For hardware that runs most of the time, buying wins: one eight-GPU server costs about $360,000 over three years to own and host in colocation, against about $685,000 to rent on demand at the median. For short or variable work, or without a facility that can host 10 kW servers, renting wins.
Should I rent an H100 or an H200?
The H100 costs $3.25 per GPU-hour at the median against $4.40 for the H200. If your models fit in 80 GB, the H100 is cheaper per job. If they do not, the H200's 141 GB often lets one GPU do the work of two H100s on inference, which more than pays for the premium.
Should I rent an H100 or a B200?
For training at scale, the B200: each one does the work of roughly two and a half to three H100s, and at $4 to $5.50 per hour on dedicated GPU clouds it is often cheaper per result. For inference, fine-tuning and anything with a mature Hopper toolchain, the H100 is the price-performance choice.
Is the A100 still worth renting?
Yes, for the right work. At a median of about $1.76 to $1.79 per GPU-hour, and from $1.09 at the cheapest providers, the A100 is often the cheapest option for LoRA fine-tuning, inference on models under about 30 billion parameters, and development. It lacks FP8 and the Transformer Engine, so the H100 wins on large-model inference.
What is the difference between H100 SXM and H100 PCIe?
The SXM version has faster HBM3 memory, up to 700 W of power and NVLink across all eight GPUs in an HGX server; it is the standard rental unit for multi-GPU work. The PCIe card runs at up to 350 W with slower memory and NVLink only between pairs, and suits single-GPU inference and development at a lower price.
How much power does an H100 server need?
An eight-GPU HGX H100 server draws about 10 kW. Four in a rack is about 40 kW, which is air-cooled but above what most facilities built for 5 to 15 kW racks can serve, so buyers who own H100s need a facility designed for that density.
How much cheaper is reserved H100 capacity?
At the median, only about 4 percent in September 2026 ($3.10 against $3.25), because on-demand supply is loose. Individual providers discount far more for larger or longer commitments; the market rule of thumb is about 25 percent off on demand for one year and about 45 percent for three, and which providers offer that for your size is the negotiation.
Why do H100 rental prices vary so much?
The GPU is identical; the provider is not. Hyperscalers bundle networking, storage and their own margins; specialist GPU clouds compete on the hourly rate; marketplaces sell interruptible spare capacity. The same silicon runs from $1.49 to $6.98 an hour, which is why quoting several providers at once matters more for GPUs than for almost anything else in infrastructure.
Will H100 prices keep falling?
Not in 2026 so far. Spot rates rose about a third in the year to August 2026 as demand for any working GPU outran supply, according to Bank of America research. Analysts expect marketplace rates to drift down as Blackwell capacity absorbs premium demand and older reservations expire, but the timing is uncertain, which is why we update this page monthly. The Rubin roadmap covers what arrives next.
What does it cost to get quotes through Metro Colo Advisory?
Nothing. The provider you choose pays us from its channel budget, the same way it pays its own sales team, and the rate is not marked up. If the best answer is to stay with your current provider or rent a different GPU, we say so.
Get H100 Quotes Through Us
Tell us the GPU count, the term, the storage and networking you need, and when you need it. You'll hear back within 24 hours from the person who will run your search, and every provider that fits will be quoted at once and benchmarked against the rates on this page. If the best answer is a different GPU, or the provider you already use, we say so. No cost, no obligation, and the provider you choose pays us.
That is what keeps every provider competing for it.
GPU rental: the GPU rental hub, GPU cluster rental, B200 and B300 rental, H200 rental, GB200 and GB300 NVL72 rental, monthly GPU rental prices, neocloud providers, CoreWeave competitors and bare metal.
What comes next: NVIDIA Rubin and the GPU roadmap, and what it means for H100 prices.
Owning and hosting: AI and GPU colocation, direct to chip cooling and data center cost.
The cloud bill: cloud repatriation, FinOps consulting and the cloud vs colocation calculator.