Cloud GPU vs on-premise GPU server for steady inference: rent or buy, and how to compare
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- For inference that runs most hours of the month, compare cost per productive GPU-hour: the rented instance hours you are billed under the pricing model you need, against the owned server’s life-cycle cost divided by the GPU-hours it serves requests
- Idle hours cost money on both sides unless rented capacity is released; as of October 2026, AWS bills an On-Demand Capacity Reservation at the On-Demand rate whether instances run in it or not, and its Savings Plans reserve no capacity
- As of October 2026, AWS’s P5 page lists the H200 only in 8-GPU instances, p5e.48xlarge and p5en.48xlarge with 1,128 GB of HBM3e each, and AWS states that restarting a stopped instance can fail when it lacks On-Demand capacity
- A month of vLLM’s running-request gauge and DCGM’s SM activity per GPU, measured before you decide, gives the productive share and the hours in which it falls
- A hybrid keeps base load on owned servers and rents planned peaks; as of October 2026, AWS Capacity Blocks reserve GPU instances to start up to eight weeks ahead for 1 to 182 days, charged up front, and Google’s Flex-start serves up to seven days
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Cloud GPU rental or own GPU servers: how to compare for steady inference
Whether a cloud GPU or an on-premise GPU server costs less for inference that runs most hours of the month depends on the cost of one productive GPU-hour on each side. On the rented side, that is what you are billed for instance hours under the pricing model your workload needs, divided by the GPU-hours that served requests. On the owned side, it is the server’s life-cycle cost divided by its productive GPU-hours over the years in service.
This article compares GPU capacity, the instances or servers you run a model on. Paying per token for a model API is a different calculation, set out in our comparison of a private LLM and a cloud API, and moving other workloads out of public cloud is the subject of our guide to cloud repatriation. The method takes rates from your own invoices and quotes.
Cost per productive GPU-hour on both sides
A productive GPU-hour is an hour in which a GPU serves at least one request or runs a batch job. An owned GPU costs the same in every hour of its service life. With five years in service, for example, that is 43,800 hours per GPU (5 × 8,760), and the owned cost per productive GPU-hour is the life-cycle cost of the server divided by the number of GPUs, by those hours and by the productive share. The life-cycle cost covers the purchase, licences, power in kWh, rack space or colocation, support and operating time.
On the rented side, an 8-GPU instance bills eight GPU-hours for every hour it is allocated, plus storage, data transfer and any support plan, divided by the productive GPU-hours in the same period. A rented GPU that is stopped or terminated costs nothing for compute, so a low productive share favours renting only if the capacity is released in idle hours and is there again when users return.
When rented capacity has to stay allocated in every hour of the month, the productive share appears on both sides and cancels out. The comparison then reduces to the owned cost per GPU-hour of service life against the rented rate per GPU-hour under the commitment or reservation you would hold.
Cloud GPU pricing models: on-demand, commitments, reservations and spot
The pricing model decides which hours are billed and whether capacity is held for you.
| PRICING MODEL | CAPACITY HELD | BILLING | TERM |
|---|---|---|---|
| AWS On-Demand | no; launch depends on available capacity | per second while running | none |
| AWS Savings Plans | no capacity reservation | the hourly commitment, every hour | one or three years |
| AWS Capacity Reservation | yes, in one Availability Zone | On-Demand rate, used or not | none for immediate use; cancel at any time |
| AWS Capacity Blocks for ML | yes, GPU instances on future dates | up front at purchase | 1 to 182 days |
| AWS Spot | no; reclaimed with a 2-minute notice | per second while running | none |
| Google Flex-start | best effort, not preemptible | lower rates on A3 and other series | up to 7 days |
| Google calendar mode | yes, up to 80 VMs | lower rates | up to 90 days |
| Google Spot VMs | no; preempted at any time | lower rates | none |
AWS EC2 pricing page, Savings Plans user guide and FAQ, EC2 user guide (Capacity Reservations, Capacity Blocks, Spot interruptions, troubleshooting); Google Cloud “Consumption options for AI Hypercomputer”, updated 8 October 2026. All read on 10 October 2026.
A Savings Plan lowers the rate in exchange for a commitment “for a one or three year period”. AWS states that “Each hour’s commitment can only be used within that hour and cannot be carried over”, that usage beyond it is “charged at regular On Demand rates”, and that Savings Plans do not provide a capacity reservation. Capacity Reservations “are charged at the equivalent On-Demand rate whether you run instances in reserved capacity or not”, and Savings Plans discounts apply to them. Google offers resource-based committed use discounts for GPUs among other resources, for one or three years.
As of 10 October 2026, AWS’s P5 page lists two instance sizes for the H200, p5e.48xlarge and p5en.48xlarge, each with eight GPUs and 1,128 GB of HBM3e; the only one-GPU size on its P5 page is the H100 in p5.4xlarge. Restarting a stopped instance can fail with InsufficientInstanceCapacity, which AWS explains as not having “enough available On-Demand capacity to fulfill your request”.
Measuring GPU utilisation and busy hours before you compare
GPU utilisation, as nvidia-smi reports it, counts the time in which any kernel runs, so it overstates how busy a serving GPU is. DCGM’s SM activity is “the fraction of time at least one warp was active on a multiprocessor, averaged over all multiprocessors”, and NVIDIA writes that “A value of 0.8 or greater is necessary, but not sufficient, for effective use of the GPU.” Our guide to DCGM metrics and XID errors lists the fields and the exporter settings. On the serving side, vLLM’s gauge vllm:num_requests_running gives the “Number of requests currently running”, on the server’s /metrics endpoint.
- Export a month of
vllm:num_requests_runningand DCGM’s SM activity per GPU, from the cloud instances you run today or from a pilot server. - Count, per model server and the GPUs it runs on, the hours with at least one running request or a batch job; their share of the month is the productive share.
- Record when those hours fall (working days, evenings, nights, month-end and seasonal peaks) and the SM activity within them.
- Decide in which hours the model must answer without delay, because that sets the pricing model on the rented side.
- Enter the rented side’s billed hours and the owned side’s life-cycle cost into the formula above.
Cost items on rented and owned GPU servers
| COST ITEM | RENTED GPUS | OWN GPU SERVERS |
|---|---|---|
| GPU capacity | instance hours, or a commitment billed every hour | purchase, spread over the years in service |
| Capacity when needed | depends on the pricing model, see above | the server is there; a second one for failover |
| Model weights and data | persistent storage; instance store is erased on stop | NVMe in the server |
| Data transfer | data transfer out billed as its own item | your own uplink |
| Power, cooling, space | part of the provider’s rate | kWh, cooling, rack or colocation |
| Operations | you run the OS, drivers and model server | the same, plus hardware and firmware |
| Failed hardware | the provider’s; a start typically moves to a new host | manufacturer warranty, spare parts |
AWS EC2 user guide (stop and start, Capacity Reservations) and EC2 On-Demand pricing page, read on 10 October 2026; the other rows are our reading.
When you compare H200 offers, note that NVIDIA lists the H200 as an SXM module for HGX boards “with 4 or 8 GPUs” and as the H200 NVL, a PCIe card at up to 600 W that “2- or 4-way NVIDIA NVLink bridge” joins at 900 GB/s per GPU. Both have 141 GB at 4.8 TB/s. AWS’s P5 page does not name the module but lists “900 GB/s NVSwitch” between the GPUs of each H200 instance. H200 NVL cards join in NVLink domains of two or four, so a model that needs eight GPUs in one NVLink domain needs an SXM system, while eight NVL cards form two such domains. A rented 8-GPU instance bills all eight, while an owned server can start with four cards and grow. For the owned side, our guide to comparing GPU server quotes lists the line items and the five-year cost categories.
Worked example: a company of 1,000 staff on eight GPUs
As an example, take a company of 1,000 staff whose chat and RAG assistant runs on eight GPUs from 07:00 to 19:00 on 21 working days, 252 hours. A nightly batch indexes new documents and summarises tickets for four hours on the same 21 days, another 84 hours. With the 730 hours of an average month (8,760 / 12), the GPUs are productive for 336 hours, a share of 46 per cent. Your own export replaces these example figures.
On an 8-GPU instance kept allocated all month, under On-Demand or a Capacity Reservation, the bill counts 730 instance hours, or 5,840 GPU-hours, for 2,688 productive GPU-hours. Stopped outside the 336 hours, the bill counts those hours plus each start and model load, two a day in this example, but every start depends on available On-Demand capacity, and the weights load from persistent storage because the instance store is erased on stop.
Over the example’s five years, eight owned GPUs serve 350,400 GPU-hours, 161,280 of them productive at the same share. If the model must answer from 07:00 every working day, the rented capacity is held for all 730 hours of the month, and the comparison is owned cost per GPU-hour against the rented rate under a reservation. If a late start on some mornings is acceptable, renting only for the busy hours and accepting the capacity risk is the case in which the rented side can come out lower, depending on the rates in your quotes.
Our Private AI/ML service includes a TCO calculation against cloud GPUs before the purchase. Send us a month of your GPU and request metrics through the form below, with the instance types you rent today.
Which workloads fit rented GPUs and which fit owned servers
| WORKLOAD PATTERN | BETTER FIT | WHY |
|---|---|---|
| Chat and RAG on working days | owned, or rented with a reservation | capacity is needed every morning, so idle hours are paid either way |
| Inference most of the day | owned | high productive share over the service life |
| Planned peaks on known dates | rented block on top of owned base | reserved for fixed dates, released after |
| Overflow on busy days | on-demand or Flex-start, if the data may go there | short and occasional, capacity not assured |
| Restartable batch jobs | Spot, or owned GPUs at night | interruptions are tolerated |
| Occasional fine-tuning | rented block | large and short |
| Pilot with unknown volume | on-demand, or an API | no commitment until the volume is measured |
Our reading of the pricing models in the table above and of the worked example; Google Cloud recommends Spot VMs for “Batch processing and data analytics”.
The hybrid pattern keeps the base load on owned servers and rents the peaks. It needs the same model and serving engine version on both sides, a gateway that sends overflow to the rented instances when the queue on the owned servers grows, which vLLM reports as vllm:num_requests_waiting, and a data class that may be processed at the provider. Planned peaks suit reserved blocks. AWS offers both H200 sizes as Capacity Blocks in selected Regions; a block can be booked to start up to eight weeks ahead and runs in 1-day steps up to 14 days and 7-day steps up to 182 days, at a price that “depends on available supply and demand” at purchase.
We build AI servers to order, sized by model and concurrent users. Tell us the base load you would move in-house and the peaks you would keep renting, and the configuration and quote follow within one business day.
Data transfer and leaving a cloud GPU provider
AWS lists data transfer out to the internet as its own item on the EC2 pricing page, with rate tiers across services. Streamed answers are small; document sets, vector indexes, logs, backups and model copies are the larger transfers, for the hybrid pattern and for the move itself.
The EU Data Act, Regulation (EU) 2023/2854, in its text as published in 2023, includes a move to on-premises ICT infrastructure in its definition of switching (Article 2(34)) and counts data egress charges among switching charges (Article 2(36)). Until 12 January 2027 providers may impose reduced switching charges, capped at their costs directly linked to the switch, and from that date none (Article 29). Standard service fees and early termination penalties are not switching charges (Article 2(36)), so check the remaining term of any Savings Plan or committed use discount before you fix a date. These rules concern the switch itself; for the hybrid pattern, which keeps using the provider in parallel, ask the provider in writing which charges apply. How these rules apply to a contract is a legal assessment for the company’s legal department, and our guide to Data Act cloud switching sets out the periods, the exceptions and the status of the Digital Omnibus proposal.
What we supply
We build AI servers to order for inference and RAG, with 2 to 8 GPUs per node, assembled and burn-in tested, with manufacturer warranty and delivery anywhere in the EU, on one EU contract and invoice. The GPUs for steady inference include the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges and the L40S, with NVIDIA AI Enterprise and vGPU licences on the same invoice. We check the rack, power and airflow before we quote. Our Private AI/ML service, with engineering by our partner Vixen.UNO, runs models on-premise or on dedicated hardware in a Tier-3 data centre in Lithuania.
FAQ
Is a cloud GPU cheaper than an on-premise GPU server?
Should I rent or buy a GPU server for LLM inference?
How do I calculate GPU utilisation for a break-even?
Can I rent a single H200 GPU on AWS?
Do Savings Plans or committed use discounts reserve GPU capacity?
Does the EU Data Act remove egress fees when moving GPU workloads on-premise?
Send us a month of GPU and request metrics from your cloud instances or pilot, the instance types and pricing models you use today and the hours your users need the model. We reply within one business day with a configuration and quote for that load, which you can compare with your cloud GPUs before buying.
Talk to an expertWe reply within one business day