GPU Rental
GPU as a Service (GPUaaS): Providers, Pricing and Better Terms
Reserved H100, H200, B200 and B300 servers and clusters, compared across providers and negotiated for you. Tell us the GPU, count and term, and you'll have a first read within 24 hours, at no cost to you.
Price My GPU RequirementUpdated September 2026
GPU as a service is renting graphics processors for AI and other accelerated computing instead of buying them: by the hour, by the month, or reserved for a term, from a hyperscale cloud, a specialist GPU cloud known as a neocloud, or a provider of dedicated GPU servers. For most AI companies it is the fastest way to get capacity for inference, fine-tuning and training without committing capital to hardware the next generation will outrun.
What Each GPU Costs to Rent
Per GPU-hour unless noted, September 2026
| GPU | Specialist cloud, on demand | Hyperscaler, on demand | Reserved, 3 to 12 months (est.) | 8-GPU server per month, on demand | Memory per GPU | Minimum term | Time to live |
|---|---|---|---|---|---|---|---|
| H100Training and inference workhorse | Median about $3.25 | About $6.88 | About $1.80 to $2.45 | About $19,000 | 80 GB | 1 month | 1 to 2 weeks |
| H200Large models, long context | About $3.00 to $4.50 | Varies by cloud | About $2.05 to $2.80 | About $21,900 | 141 GB | 1 month | 4 to 8 weeks |
| B200Current-generation training | About $4.00 to $6.50 | From about $11.28 | About $2.90 to $3.95 | About $30,700 | 180 GB | About 3 months | 4 to 8 weeks |
| B300Largest models, most memory | Median about $7.89 | From about $17.80 | About $4.35 to $5.90 | About $46,100 | 288 GB | About 6 months | 4 to 8 weeks |
| GB200 and GB300 NVL72Rack-scale, liquid-cooled | Quoted per rack | Quoted per rack | Negotiated per rack | Rented as 72-GPU racks | Up to 288 GB | 6 to 12 months | Phased delivery |
Sources: GetDeploying (H100, September 2026) and GetDeploying (B300); CloudZero (AWS H100, August 2026); GPUPerHour (H200, B200 and reserved discounts); hyperscaler B200 and B300 floors from FastGPU's September 2026 tracker. Reserved figures are Metro Colo Advisory estimates applying the typical 25 to 45 percent discount to the on-demand midpoint. Monthly figures assume 8 GPUs running around the clock. Prices move weekly; treat them as a starting point, not a quote.
The hard part is not finding a price list. It is finding capacity that exists, on the hardware you need, where you need it, on terms that hold up. Going to a provider on your own means negotiating against its rate card alone. We benchmark your requirement against what the market pays, negotiate the price and terms, and stay on the account after you sign. The provider you choose pays us; you pay nothing. For rack-scale, liquid-cooled deployments you own, the AI and GPU colocation guide covers that path.
Send the GPU model, count, term, location and workload. Within 24 hours you'll have a first read on who can deliver it and what it should cost. No signup, no call needed, one contact instead of a sales team.
Why rent GPUs through us instead of going direct
Going direct gets you a rate card, a sales call and terms you negotiate alone. Here is the difference.
| Going direct | Through us | |
|---|---|---|
| Price | The provider's rate card | Negotiated against what the market actually pays |
| Cost to you | Nothing extra | Nothing extra; the provider pays us |
| Options compared | Only the ones you find and call | The options that fit your requirement, compared for you |
| Contract terms | Negotiated alone, against a deal team | Negotiated by someone who knows what providers accept |
| Time to an answer | Weeks of calls | A first read within 24 hours, often the same day |
| Sales contact | A sales team for every provider you call | One contact |
| After you sign | An account rep who works for the provider | Someone on your side through delivery and renewal |
1. You pay what the market pays, not the rate card.
Prices for the same GPU vary widely from one provider to the next, and reserved terms vary even more. We know what providers actually sign at, compare the options that fit your requirement, and negotiate against the market rather than accepting the first number. Where more than one provider can serve it, they compete for it.
2. You do not pay more for it.
The provider you choose pays us through its channel program, from the budget it would otherwise spend on its own sales team. Going direct does not get you a lower price; it only means nobody is comparing for you.
3. We negotiate the terms that decide whether the capacity is worth having.
Delivery dates with remedies if they slip, the right to move to newer hardware during the term, network and storage performance written into the contract, bandwidth charges, support commitments and exit rights. On a reserved GPU contract those terms are worth more than a few cents an hour.
4. We know where capacity is, not just what is listed.
Published availability and delivered availability are different things, especially for current-generation GPUs. We take live AI requirements to operators and providers across North America, and Data Center Knowledge led its August 2026 reporting on usable AI capacity with the practice’s own search. That is the gap between a provider that can deliver and one that can quote.
5. One contact, before and after you sign.
Providers quote through us rather than chasing you. And because we are paid over the life of the contract, we stay on the account: if delivery slips or a server fails, you have one person to call who knows who to escalate to, and before renewal we price the alternatives again. If the right answer is a provider we do not work with, or owning hardware instead of renting, we will tell you.
Metro Colo Advisory is an independent infrastructure advisory practice working across North America. We do not own GPUs, run a cloud, or operate data centers, and we work for the buyer.
If a launch or a training run depends on capacity arriving on time, tell us the date. The earlier we start, the more providers can still meet it.
What Does GPU as a Service Cost?
GPU cloud pricing follows three models, and the gap between them is large, whether it is GPU server rental by the month or capacity by the hour. Current prices for each GPU are in the table at the top of this page.
| Model | How it works | Typical price versus on demand | Fits |
|---|---|---|---|
| On demand | By the hour or month, no commitment | Full price | Experiments, bursts, short projects |
| Reserved | A fixed amount of capacity for a term, commonly 3 to 36 months | Typically 25 to 45 percent less | Steady inference and planned training |
| Spot or interruptible | Spare capacity that can be reclaimed at short notice | Often around half | Work that can stop and resume from checkpoints |
What the hourly rate leaves out
On a large reservation, those lines can move the total as much as the headline rate, which is why every quote we bring back shows them side by side.
Send us the GPU, count and term, and within 24 hours you'll have a read on what it should cost and who can deliver it, at no cost to you.
A Worked Example: One Inference Fleet, Four Ways to Pay
An AI company runs steady inference around the clock on eight servers, each with eight H100 GPUs: 64 GPUs in total. The figures below are illustrative, using September 2026 market rates.
| Hyperscaler, on demand | GPU cloud, on demand | Reserved, 12 months | Owned, in colocation | |
|---|---|---|---|---|
| Price basis | About $6.88 per GPU-hour | Median about $3.25 per GPU-hour | About $2.28 per GPU-hour, 30% below on demand | About $285,000 per server, plus space and power |
| Monthly cost | About $321,000 | About $152,000 | About $106,000 | About $80,000 over 3 years, before power and financing |
| Commitment | None | None | 12 months | Hardware life, typically 3 to 5 years |
Reserving is where the risk moves to you. A year of commitment is worth it only if the delivery date, the network performance and the right to upgrade are written into the contract.
At steady scale over several years, owning can win. Eight H100 servers at about $285,000 each cost about $2.3 million. Spread over three years, plus about $16,000 a month of colocation, that is roughly $80,000 a month, about a quarter below the reserved rate, before metered power and financing, and with the hardware still yours at the end (Metro Colo Advisory calculation from published prices; see the data center cost page). That path is covered in the AI and GPU colocation guide, and for fleets planned beyond three years it is worth pricing alongside the reservation.
Sources: CloudZero (AWS H100, August 2026); GetDeploying median, September 2026. Monthly figures assume 64 GPUs running around the clock.
Send the GPU type, server count and how the workload runs. Within 24 hours you'll know what reserving could save and who can deliver it.
How we get you a better price
Free to you does not mean hands-off. Five things move the number, and we work all of them.
Market knowledge.
A provider quoting a buyer on their own quotes its rate card. A buyer who knows what comparable deployments actually sign for, and where more than one provider can serve the requirement, gets a different number. You do not pay more for this: the provider pays us from the budget it would spend on its own sales team.
Term structure.
Prices step down at set term lengths, commonly 3, 6 and 12 months. We quote each so you can see where the breaks are, and whether a slightly longer term costs less overall.
Size tiers.
Cluster pricing also steps down with size. One major provider’s published reserved H100 cluster rates fall from $6.16 to $5.85 to $5.54 per GPU-hour as clusters pass 64 and 256 GPUs, according to GPUPerHour. Sometimes a few more GPUs cost almost nothing extra.
Idle inventory.
Providers with capacity sitting unused price it more sharply than providers who are sold out. We know who has what this month.
The terms that cost money later.
Data transfer charges, storage, ramp schedules that bill servers only as they are delivered, and upgrade rights to newer GPUs. These are often worth more than the headline rate, and they are where going direct loses the most.
Which GPU Fits the Workload?
| GPU | Typical use | What to know |
|---|---|---|
| L40S and RTX Pro 6000 | Inference for small and mid-sized models, vision, graphics | Lower cost per hour and widely available |
| H100 | Training and inference, the most widely deployed data center GPU | 80 GB of memory per GPU; eight-GPU servers are the standard unit, and supply is the deepest of any current GPU |
| H200 | Inference for large models and long context | 141 GB of memory per GPU, so larger models run with less splitting across GPUs |
| B200 and B300 | Current-generation training and high-throughput inference | The highest performance per server; B300's 288 GB fits the largest models on fewer GPUs. Tighter supply and longer minimum terms |
| GB200 and GB300 NVL72 | Rack-scale training and inference | 72 GPUs working as one system in a liquid-cooled rack; rented as whole racks by a few providers, or owned in liquid-cooled colocation |
Current rental prices for each are in the table at the top of this page. The newest hardware is not always the right buy: a model that fits comfortably on H100s may cost less to serve there than on B200s, and the older generation is often easier to get. For rack-scale systems, the direct to chip cooling guide explains what the facility needs.
Rental ranges from independent price trackers, September 2026: GPUPerHour for H200 and B200, and GetDeploying for B300.
The newest hardware is not always the right buy. A model that fits comfortably on H100s may cost less to serve there than on B200s, and the older generation is often easier to get. Rack-scale systems such as NVIDIA’s GB200 and GB300 are a different category: they are usually deployed in liquid-cooled colocation rather than rented by the server, which the direct to chip cooling guide explains.
GPU Cluster Rental: How Scale Changes the Market
A node is one server, usually with eight GPUs linked inside it. A GPU cluster is many nodes joined by a high-speed network, commonly InfiniBand, so they work as one machine for training. That network is what separates the two ways companies rent GPUs at scale, and they are priced very differently.
What an H100 cluster costs
Per GPU-hour and per month running continuously, September 2026
| Size | Setup | Typical price per GPU-hour | Per month |
|---|---|---|---|
| 8 GPUs1 server | Single server, on demand | Median about $3.25 | About $19,000 |
| 64 GPUs8 servers | Inference fleet, reserved 3 to 12 months (est.) | About $1.80 to $2.45 | About $84,000 to $114,000 |
| 64+ GPUs8+ servers | Training cluster with InfiniBand, reserved (published rate) | From about $5.85 | About $273,000 for 64 GPUs |
| 256+ GPUs32+ servers | Training cluster with InfiniBand, reserved (published rate) | From about $5.54 | About $1.04 million for 256 GPUs |
Sources: GetDeploying median, September 2026; published reserved cluster rates from one major provider via GPUPerHour. The inference fleet row is a Metro Colo Advisory estimate applying the typical 25 to 45 percent reserved discount. Monthly figures assume every GPU running around the clock.
Five things that move a cluster quote
Who can serve you at each size
| Scale | Typical use | The market |
|---|---|---|
| 1 to 4 servers8 to 32 GPUs | Inference, fine-tuning, development | Many providers, often available within days |
| 5 to 32 servers40 to 256 GPUs | Production inference fleets, mid-sized training | A solid field of neoclouds and dedicated providers; delivery dates vary |
| Hundreds of servers1,000+ GPUs | Frontier training | A handful of providers, long lead times, large commitments |
| Rack-scale systemsGB200 and GB300 NVL72 | Next-generation training and inference | Rented as whole racks by a few providers, or owned in liquid-cooled colocation |
Most AI companies sit in the middle two rows, where competition is strongest and a well-run comparison makes the most difference. At that size, the gap between the first quote you receive and a negotiated one is often tens of thousands of dollars a month.
Tell us the GPU, cluster size, network and storage needs, term and target date. We'll tell you who can deliver it, when, and what it should cost, within 24 hours.
What should a GPU rental contract include?
| Term | Why it matters |
|---|---|
| A delivery date, with remedies if it slips | Capacity promised is not capacity delivered |
| The exact hardware, and upgrade rights during the term | A year is a long time in GPU generations |
| Network and storage performance | Training stalls on slow interconnect and slow storage |
| Bandwidth and data transfer charges | Moving datasets and serving traffic can rival the GPU bill |
| Uptime and hardware replacement commitments | A failed GPU in an eight-GPU server can idle the whole machine |
| Data location and security attestations | Customer data and model weights may carry obligations |
| Term, renewal price and exit | The next contract is negotiated from this one |
The data center lease guide covers how we approach contract terms, and the compliance guide covers attestations.
Neocloud providers or hyperscalers?
A neocloud is a cloud built mainly to rent GPUs for AI, rather than a general-purpose cloud with GPUs added. Neocloud providers such as CoreWeave, Lambda, Crusoe and Nebius typically offer current-generation GPUs sooner, at lower prices for committed capacity, with networking built for training. Hyperscalers offer the surrounding services, managed AI platforms, databases and security tooling, and the convenience of one bill.
Many AI companies use both: a hyperscaler for the application and a neocloud or dedicated provider for the GPUs, joined by a private connection. The decision is less about which is better than about which part of the workload belongs where.
For a GPU cloud for inference specifically, location matters as much as price: serving users from GPUs close to them cuts latency, which is why many LLM hosting providers now spread capacity across several metros rather than one large campus. For a side-by-side of the main providers, see CoreWeave competitors and alternatives.
When is renting GPUs the wrong answer?
When a large fleet will run steadily for three years or more, owning the hardware and placing it in colocation often costs less over its life. When a workload depends on one cloud’s managed AI services, moving the GPUs elsewhere can cost more in rework than it saves. And for occasional experiments, hourly capacity or cloud credits are simpler than any contract. In each case, we will say so. The cloud repatriation guide covers the economics of moving steady workloads off public cloud.
What this page cannot tell you
It cannot tell you which provider has your GPUs available this month, because that changes weekly. It cannot tell you what a reservation will cost for your requirement, because committed pricing is negotiated rather than published. It cannot tell you whether renting or owning is cheaper for your fleet over three years, because that depends on your workload and your financing. And it cannot tell you which contract terms a provider will accept, because that is the negotiation.
Those are the work. This page is the framework it uses.
How the practice works on GPU capacity
Placement, paid by the provider. We build the requirement, run the comparison, negotiate the contract and set up the connectivity. The provider you choose pays us through its channel program, and you are under no obligation.
Consulting, paid by you, only if you want it. A fixed-fee rent-versus-own analysis for companies planning large fleets.
See how it works for the placement process and data center consulting for the wider practice, including data center due diligence for investors and VMware alternatives for enterprises replacing their own platforms.
Frequently Asked Questions
What is GPU as a service?
GPU as a service is renting GPUs for AI and accelerated computing instead of buying them, by the hour, by the month or reserved for a term. It is sold by hyperscale clouds, specialist GPU clouds known as neoclouds, dedicated GPU server providers and marketplaces. It lets AI companies get capacity quickly without committing capital to hardware. Metro Colo Advisory helps you compare the options that fit, at no cost to you.
How much does it cost to rent an H100?
H100 rental is usually priced per server of eight GPUs. Independent price trackers put the median specialist GPU cloud at about $3.25 per GPU-hour on demand in September 2026, about $26 an hour or $19,000 a month for an eight-GPU server running continuously, while hyperscalers charge roughly twice that. Reserved capacity on 3 to 12 month terms typically runs 25 to 45 percent lower. Storage, bandwidth and cluster networking are usually extra. If you want to rent H100 capacity, tell us the count and term and we will price it for you.
How much does it cost to rent an H200, B200 or B300?
On demand, H200 rental runs about $24 to $36 an hour per eight-GPU server, B200 about $4.00 to $6.50 per GPU-hour, and B300 at a market median near $7.89 per GPU-hour, according to independent price trackers in September 2026. Reserved terms cost less, and availability for B200 and B300 is tighter, so delivery dates matter as much as price. Tell us the GPU, count and term and we will price B200, B300 or H200 rental for you.
What is a neocloud?
A neocloud is a cloud provider built mainly to rent GPUs for AI, such as CoreWeave, Lambda, Crusoe or Nebius, rather than a general-purpose cloud. Neoclouds often offer current-generation GPUs sooner and cheaper for committed capacity, with networking built for training. Many companies pair a neocloud for GPUs with a hyperscaler for everything else. We can help you weigh neocloud providers against the other options.
What is the difference between on-demand and reserved GPU capacity?
On-demand capacity is billed by the hour or month with no commitment and costs the most. Reserved capacity commits you to a fixed amount for a term, commonly 3 to 36 months, in exchange for a discount that is typically 25 to 45 percent. Reserved suits steady workloads; on-demand suits experiments and bursts. We negotiate reserved terms so the commitment is protected.
What is a dedicated GPU server?
A dedicated GPU server, also called bare metal GPU, is a whole physical machine with its GPUs assigned to you alone, with no virtualization layer and no other tenants. It suits steady inference and fine-tuning where performance consistency and control matter. It is usually rented monthly or reserved for a term. Metro Colo Advisory can compare dedicated GPU server providers for your workload.
Should we rent GPUs or buy them?
Rent when the workload is uncertain, bursty or short-lived, or when you need the newest hardware without a capital commitment. Owning, placed in colocation, often wins for large fleets that will run steadily for three years or more. The answer depends on fleet size, utilization and financing. We can price both for your fleet.
How do I find GPU capacity when providers are sold out?
Take the requirement to more providers and more provider types at once, stay flexible on GPU generation and location where you can, and ask for committed delivery dates rather than listed availability. Older generations and secondary locations often have capacity when the newest hardware does not. Running that search is what we do, at no cost to you.
What should a GPU rental contract include?
A delivery date with remedies, the exact hardware and upgrade rights, network and storage performance, bandwidth charges, uptime and replacement commitments, data location and security attestations, and renewal and exit terms. Those terms decide whether a reservation is worth having. We negotiate them on your behalf.
Can you help with GPU clusters for training?
Yes. For training clusters we compare providers on the cluster network, storage throughput and delivery date as well as price, since a cluster is only as fast as its slowest link. Very large clusters have fewer providers and longer lead times, so starting early matters. Tell us the size and timeline and we will say what is realistic.
How do you get a better price than going direct?
Five ways: knowing what comparable deployments actually sign for, quoting each term length to find the price breaks, sizing to the provider’s volume tiers, finding providers with idle inventory, and negotiating the terms that cost money later, such as data transfer, storage and upgrade rights. You do not pay more for it, because the provider pays us from its own sales budget.
Do you share my details with providers?
Only with the providers quoting your requirement, and only what they need to price it. They quote through us rather than chasing you, and none of them will contact you unless you ask to speak with them. We never list or sell your details.
How quickly do I get pricing, and do I need a call?
A first read within 24 hours for standard configurations, often the same day, by email. You do not need a call to see it. Large clusters and rack-scale Blackwell systems can take a little longer because fewer providers can deliver them.
Which GPU providers do you work with?
We work through channel partnerships with providers of dedicated GPU servers, fully managed GPU hosting, and GPU capacity inside managed private clouds, including options audited for regulated data. Which provider fits depends on your GPU, size, region and how much you want managed for you. If a requirement needs something our partners cannot supply, we will tell you rather than stretch the fit.
Do I pay anything to use Metro Colo Advisory for GPU capacity?
No. The provider you choose pays us through its channel program, and you are under no obligation to use us. If you want a formal rent-versus-own analysis for a large fleet, that is a separate fixed-fee engagement, and only if you ask for it. The first conversation is free either way.
Related reading
- AI infrastructure: AI and GPU colocation, direct to chip cooling, wholesale colocation for large owned deployments, and data center tiers.
- Other ways to run steady workloads: CoreWeave competitors and alternatives for the providers compared, bare metal, private cloud, the data center cost guide, what colocation is, the colocation pricing guide, and a comparison of all providers.
- The cloud decision: the cloud repatriation guide, FinOps consulting, and the cloud versus colocation calculator.
- Contracting: the data center lease guide, how it works, and why companies use an independent advisor.
Tell us what you need to run
The GPU model or performance target, how many, for how long, where, and what the workload is. That is enough for a first read on which providers can deliver it and roughly what it should cost, and you will have it within 24 hours. If renting is the wrong answer, we will tell you that.
North American coverage. Independent of any single GPU provider. No cost to buyers.