BLOG · COMPARISON ·

Cloud GPU vs on-premise GPU server for steady inference: rent or buy, and how to compare

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • For inference that runs most hours of the month, compare cost per productive GPU-hour: the rented instance hours you are billed under the pricing model you need, against the owned server’s life-cycle cost divided by the GPU-hours it serves requests
  • Idle hours cost money on both sides unless rented capacity is released; as of October 2026, AWS bills an On-Demand Capacity Reservation at the On-Demand rate whether instances run in it or not, and its Savings Plans reserve no capacity
  • As of October 2026, AWS’s P5 page lists the H200 only in 8-GPU instances, p5e.48xlarge and p5en.48xlarge with 1,128 GB of HBM3e each, and AWS states that restarting a stopped instance can fail when it lacks On-Demand capacity
  • A month of vLLM’s running-request gauge and DCGM’s SM activity per GPU, measured before you decide, gives the productive share and the hours in which it falls
  • A hybrid keeps base load on owned servers and rents planned peaks; as of October 2026, AWS Capacity Blocks reserve GPU instances to start up to eight weeks ahead for 1 to 182 days, charged up front, and Google’s Flex-start serves up to seven days

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Cloud GPU rental or own GPU servers: how to compare for steady inference

Whether a cloud GPU or an on-premise GPU server costs less for inference that runs most hours of the month depends on the cost of one productive GPU-hour on each side. On the rented side, that is what you are billed for instance hours under the pricing model your workload needs, divided by the GPU-hours that served requests. On the owned side, it is the server’s life-cycle cost divided by its productive GPU-hours over the years in service.

This article compares GPU capacity, the instances or servers you run a model on. Paying per token for a model API is a different calculation, set out in our comparison of a private LLM and a cloud API, and moving other workloads out of public cloud is the subject of our guide to cloud repatriation. The method takes rates from your own invoices and quotes.

Cost per productive GPU-hour on both sides

A productive GPU-hour is an hour in which a GPU serves at least one request or runs a batch job. An owned GPU costs the same in every hour of its service life. With five years in service, for example, that is 43,800 hours per GPU (5 × 8,760), and the owned cost per productive GPU-hour is the life-cycle cost of the server divided by the number of GPUs, by those hours and by the productive share. The life-cycle cost covers the purchase, licences, power in kWh, rack space or colocation, support and operating time.

On the rented side, an 8-GPU instance bills eight GPU-hours for every hour it is allocated, plus storage, data transfer and any support plan, divided by the productive GPU-hours in the same period. A rented GPU that is stopped or terminated costs nothing for compute, so a low productive share favours renting only if the capacity is released in idle hours and is there again when users return.

When rented capacity has to stay allocated in every hour of the month, the productive share appears on both sides and cancels out. The comparison then reduces to the owned cost per GPU-hour of service life against the rented rate per GPU-hour under the commitment or reservation you would hold.

Cloud GPU pricing models: on-demand, commitments, reservations and spot

The pricing model decides which hours are billed and whether capacity is held for you.

PRICING MODELCAPACITY HELDBILLINGTERM
AWS On-Demandno; launch depends on available capacityper second while runningnone
AWS Savings Plansno capacity reservationthe hourly commitment, every hourone or three years
AWS Capacity Reservationyes, in one Availability ZoneOn-Demand rate, used or notnone for immediate use; cancel at any time
AWS Capacity Blocks for MLyes, GPU instances on future datesup front at purchase1 to 182 days
AWS Spotno; reclaimed with a 2-minute noticeper second while runningnone
Google Flex-startbest effort, not preemptiblelower rates on A3 and other seriesup to 7 days
Google calendar modeyes, up to 80 VMslower ratesup to 90 days
Google Spot VMsno; preempted at any timelower ratesnone

AWS EC2 pricing page, Savings Plans user guide and FAQ, EC2 user guide (Capacity Reservations, Capacity Blocks, Spot interruptions, troubleshooting); Google Cloud “Consumption options for AI Hypercomputer”, updated 8 October 2026. All read on 10 October 2026.

A Savings Plan lowers the rate in exchange for a commitment “for a one or three year period”. AWS states that “Each hour’s commitment can only be used within that hour and cannot be carried over”, that usage beyond it is “charged at regular On Demand rates”, and that Savings Plans do not provide a capacity reservation. Capacity Reservations “are charged at the equivalent On-Demand rate whether you run instances in reserved capacity or not”, and Savings Plans discounts apply to them. Google offers resource-based committed use discounts for GPUs among other resources, for one or three years.

As of 10 October 2026, AWS’s P5 page lists two instance sizes for the H200, p5e.48xlarge and p5en.48xlarge, each with eight GPUs and 1,128 GB of HBM3e; the only one-GPU size on its P5 page is the H100 in p5.4xlarge. Restarting a stopped instance can fail with InsufficientInstanceCapacity, which AWS explains as not having “enough available On-Demand capacity to fulfill your request”.

Measuring GPU utilisation and busy hours before you compare

GPU utilisation, as nvidia-smi reports it, counts the time in which any kernel runs, so it overstates how busy a serving GPU is. DCGM’s SM activity is “the fraction of time at least one warp was active on a multiprocessor, averaged over all multiprocessors”, and NVIDIA writes that “A value of 0.8 or greater is necessary, but not sufficient, for effective use of the GPU.” Our guide to DCGM metrics and XID errors lists the fields and the exporter settings. On the serving side, vLLM’s gauge vllm:num_requests_running gives the “Number of requests currently running”, on the server’s /metrics endpoint.

  1. Export a month of vllm:num_requests_running and DCGM’s SM activity per GPU, from the cloud instances you run today or from a pilot server.
  2. Count, per model server and the GPUs it runs on, the hours with at least one running request or a batch job; their share of the month is the productive share.
  3. Record when those hours fall (working days, evenings, nights, month-end and seasonal peaks) and the SM activity within them.
  4. Decide in which hours the model must answer without delay, because that sets the pricing model on the rented side.
  5. Enter the rented side’s billed hours and the owned side’s life-cycle cost into the formula above.

Cost items on rented and owned GPU servers

COST ITEMRENTED GPUSOWN GPU SERVERS
GPU capacityinstance hours, or a commitment billed every hourpurchase, spread over the years in service
Capacity when neededdepends on the pricing model, see abovethe server is there; a second one for failover
Model weights and datapersistent storage; instance store is erased on stopNVMe in the server
Data transferdata transfer out billed as its own itemyour own uplink
Power, cooling, spacepart of the provider’s ratekWh, cooling, rack or colocation
Operationsyou run the OS, drivers and model serverthe same, plus hardware and firmware
Failed hardwarethe provider’s; a start typically moves to a new hostmanufacturer warranty, spare parts

AWS EC2 user guide (stop and start, Capacity Reservations) and EC2 On-Demand pricing page, read on 10 October 2026; the other rows are our reading.

When you compare H200 offers, note that NVIDIA lists the H200 as an SXM module for HGX boards “with 4 or 8 GPUs” and as the H200 NVL, a PCIe card at up to 600 W that “2- or 4-way NVIDIA NVLink bridge” joins at 900 GB/s per GPU. Both have 141 GB at 4.8 TB/s. AWS’s P5 page does not name the module but lists “900 GB/s NVSwitch” between the GPUs of each H200 instance. H200 NVL cards join in NVLink domains of two or four, so a model that needs eight GPUs in one NVLink domain needs an SXM system, while eight NVL cards form two such domains. A rented 8-GPU instance bills all eight, while an owned server can start with four cards and grow. For the owned side, our guide to comparing GPU server quotes lists the line items and the five-year cost categories.

Worked example: a company of 1,000 staff on eight GPUs

As an example, take a company of 1,000 staff whose chat and RAG assistant runs on eight GPUs from 07:00 to 19:00 on 21 working days, 252 hours. A nightly batch indexes new documents and summarises tickets for four hours on the same 21 days, another 84 hours. With the 730 hours of an average month (8,760 / 12), the GPUs are productive for 336 hours, a share of 46 per cent. Your own export replaces these example figures.

On an 8-GPU instance kept allocated all month, under On-Demand or a Capacity Reservation, the bill counts 730 instance hours, or 5,840 GPU-hours, for 2,688 productive GPU-hours. Stopped outside the 336 hours, the bill counts those hours plus each start and model load, two a day in this example, but every start depends on available On-Demand capacity, and the weights load from persistent storage because the instance store is erased on stop.

Over the example’s five years, eight owned GPUs serve 350,400 GPU-hours, 161,280 of them productive at the same share. If the model must answer from 07:00 every working day, the rented capacity is held for all 730 hours of the month, and the comparison is owned cost per GPU-hour against the rented rate under a reservation. If a late start on some mornings is acceptable, renting only for the busy hours and accepting the capacity risk is the case in which the rented side can come out lower, depending on the rates in your quotes.

Our Private AI/ML service includes a TCO calculation against cloud GPUs before the purchase. Send us a month of your GPU and request metrics through the form below, with the instance types you rent today.

Which workloads fit rented GPUs and which fit owned servers

WORKLOAD PATTERNBETTER FITWHY
Chat and RAG on working daysowned, or rented with a reservationcapacity is needed every morning, so idle hours are paid either way
Inference most of the dayownedhigh productive share over the service life
Planned peaks on known datesrented block on top of owned basereserved for fixed dates, released after
Overflow on busy dayson-demand or Flex-start, if the data may go thereshort and occasional, capacity not assured
Restartable batch jobsSpot, or owned GPUs at nightinterruptions are tolerated
Occasional fine-tuningrented blocklarge and short
Pilot with unknown volumeon-demand, or an APIno commitment until the volume is measured

Our reading of the pricing models in the table above and of the worked example; Google Cloud recommends Spot VMs for “Batch processing and data analytics”.

The hybrid pattern keeps the base load on owned servers and rents the peaks. It needs the same model and serving engine version on both sides, a gateway that sends overflow to the rented instances when the queue on the owned servers grows, which vLLM reports as vllm:num_requests_waiting, and a data class that may be processed at the provider. Planned peaks suit reserved blocks. AWS offers both H200 sizes as Capacity Blocks in selected Regions; a block can be booked to start up to eight weeks ahead and runs in 1-day steps up to 14 days and 7-day steps up to 182 days, at a price that “depends on available supply and demand” at purchase.

We build AI servers to order, sized by model and concurrent users. Tell us the base load you would move in-house and the peaks you would keep renting, and the configuration and quote follow within one business day.

Data transfer and leaving a cloud GPU provider

AWS lists data transfer out to the internet as its own item on the EC2 pricing page, with rate tiers across services. Streamed answers are small; document sets, vector indexes, logs, backups and model copies are the larger transfers, for the hybrid pattern and for the move itself.

The EU Data Act, Regulation (EU) 2023/2854, in its text as published in 2023, includes a move to on-premises ICT infrastructure in its definition of switching (Article 2(34)) and counts data egress charges among switching charges (Article 2(36)). Until 12 January 2027 providers may impose reduced switching charges, capped at their costs directly linked to the switch, and from that date none (Article 29). Standard service fees and early termination penalties are not switching charges (Article 2(36)), so check the remaining term of any Savings Plan or committed use discount before you fix a date. These rules concern the switch itself; for the hybrid pattern, which keeps using the provider in parallel, ask the provider in writing which charges apply. How these rules apply to a contract is a legal assessment for the company’s legal department, and our guide to Data Act cloud switching sets out the periods, the exceptions and the status of the Digital Omnibus proposal.

What we supply

We build AI servers to order for inference and RAG, with 2 to 8 GPUs per node, assembled and burn-in tested, with manufacturer warranty and delivery anywhere in the EU, on one EU contract and invoice. The GPUs for steady inference include the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges and the L40S, with NVIDIA AI Enterprise and vGPU licences on the same invoice. We check the rack, power and airflow before we quote. Our Private AI/ML service, with engineering by our partner Vixen.UNO, runs models on-premise or on dedicated hardware in a Tier-3 data centre in Lithuania.

FAQ

Is a cloud GPU cheaper than an on-premise GPU server?
It depends on the productive GPU-hours and on the hours you pay for. An owned server costs the same busy or idle, and rented capacity costs nothing only while it is released, so for inference that must answer every working morning the rented capacity is held all month. Compare the owned life-cycle cost per productive GPU-hour with the rented bill per productive GPU-hour under the pricing model you would need.
Should I rent or buy a GPU server for LLM inference?
Rent for pilots with unknown volume, short fine-tuning runs, planned peaks and restartable batch jobs. Own servers fit inference that runs most hours of the day or every working day, where rented capacity would be reserved for every hour anyway. The two can be combined, with the base load owned and peaks rented.
How do I calculate GPU utilisation for a break-even?
Export a month of vLLM’s running-request gauge and DCGM’s SM activity per GPU, and count the hours with at least one running request or a batch job. Their share of the month is the productive share, and the owned cost per productive GPU-hour is the life-cycle cost divided by GPUs, hours in service and that share. NVIDIA calls an SM activity of 0.8 or more necessary but not sufficient for effective use of the GPU.
Can I rent a single H200 GPU on AWS?
As AWS’s P5 page read on 10 October 2026, the H200 comes in two instance sizes, p5e.48xlarge and p5en.48xlarge, each with eight GPUs and 1,128 GB of HBM3e. The only one-GPU size on that page is p5.4xlarge with an H100. An owned server can start with fewer H200 NVL cards and add more later.
Do Savings Plans or committed use discounts reserve GPU capacity?
AWS states that Savings Plans do not provide a capacity reservation; capacity is held by an On-Demand Capacity Reservation, billed at the On-Demand rate whether instances run in it or not, or by a Capacity Block for ML on fixed dates. Each hour’s Savings Plan commitment can only be used within that hour. Google offers resource-based commitments for GPUs for one or three years and holds capacity through reservations.
Does the EU Data Act remove egress fees when moving GPU workloads on-premise?
The Data Act, as published in 2023, counts a move to on-premises infrastructure as switching and data egress charges as switching charges, which providers may impose in reduced form, capped at their directly linked costs, until 12 January 2027 and not at all from that date. Standard service fees and early termination penalties are not switching charges, so check the remaining term of any commitment before planning the move. How this applies to a given contract is a legal assessment for the company’s legal department.

Send us a month of GPU and request metrics from your cloud instances or pilot, the instance types and pricing models you use today and the hours your users need the model. We reply within one business day with a configuration and quote for that load, which you can compare with your cloud GPUs before buying.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna