Buying an AI server for LLMs: what to specify before you ask for a quote
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Specify the workload and the peak number of requests in flight first, then the model, its precision and the context length; these fix the GPU memory, and the CPU, RAM, NVMe, network, power and licences are sized around the cards
- Plan memory for the KV cache as 0.9 × the memory the driver reports, minus 3 GiB, minus the weights; with an FP8 cache, one 96 GB RTX PRO 6000 holds 12 concurrent 8K sessions of Llama 3.3 70B in FP8 and an H200 NVL holds 44
- NVIDIA’s configuration guide for NVIDIA-Certified Systems (30 September 2026) gives two CPU sockets, six physical cores per GPU and twice the GPU memory in system RAM as a starting point; for two 96 GB inference cards, our 70B worked example plans one processor and 192 to 256 GB
- Power is specified per rack position in amps and phases: two 600 W cards and one processor draw about 2 kW at the wall and fit a 16 A single-phase feed, while eight 600 W cards draw 4.8 kW on their own
- vLLM on bare metal needs no NVIDIA licence; NIM in production and compute vGPU profiles need NVIDIA AI Enterprise, licensed per GPU, and each H200 NVL includes a five-year subscription
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What to specify when buying an AI server
To buy an on-premise AI server that fits the job, specify the workload and the peak number of concurrent requests first, then the model, its precision and the context length. Together they fix the GPU memory, and the GPU class and card count follow from it. The CPU, system memory, NVMe storage, network, power, cooling and licences are sized around the cards, and warranty and support terms complete the request.
| DECISION | WHAT TO DECIDE | READ MORE |
|---|---|---|
| Workload and users | inference, RAG or fine-tuning; peak requests in flight | users per RTX PRO 6000 |
| Model, precision, context | parameters, weight precision, KV cache type, context length | how much VRAM an LLM needs |
| GPU class and count | memory and bandwidth per card, MIG, vGPU, NVLink, copies or a split model | which GPU for a first AI project |
| CPU, RAM and NVMe | cores per GPU, x16 links, RAM against GPU memory, storage tiers | 70B server worked example |
| Network, power, cooling | ports, feed in amps and phases, PDU outlets, inlet temperature | GPU rack power and cooling |
| Software and licences | bare metal or VMs, serving engine, AI Enterprise, vGPU | AI Enterprise and vGPU licences |
| Warranty and support | warranty, burn-in report, commissioning, support | AI servers built to order |
Decisions in the order this article takes them; the linked pages cover each step in depth.
Workload and concurrent users
Inference for an assistant holds the weights plus one KV cache per active conversation. RAG adds an embedding model, a vector index and often a reranker, and its prompts carry the retrieved passages, so its requests are longer than in plain chat. Fine-tuning keeps gradients and optimiser states in memory on top of the weights; we build training and fine-tuning nodes with H200 NVL cards and NVLink bridges or with the RTX PRO 6000 Server Edition. List every model the server keeps loaded, since each takes memory of its own.
The user count that sizes a server is the peak number of requests in flight. A hundred employees with an assistant open in a browser produce far fewer simultaneous generations, and we found no published ratio between the two. Take the busy-hour peak from a pilot’s gateway logs, or estimate it and say so. Give the context length you will configure as well. If none is set, vLLM takes it from the model configuration, 40,960 tokens for Qwen3-32B, and declaring 8K instead fits five times as many full-length sessions.
Model size, precision and GPU memory
Weights take two bytes per parameter in BF16, one in FP8 and about 0.6 in NVIDIA’s NVFP4 checkpoints. Our sizing guides give the KV cache 0.9 × the memory the driver reports, minus 3 GiB for CUDA graphs and activations, minus the weights. The factor 0.9 is the share of memory the serving engine may use, set with gpu_memory_utilization in vLLM, whose current default is 0.92. The driver reports 95.6 GiB on a 96 GB RTX PRO 6000 and 140.4 GiB on an H200 NVL. In an FP8 cache, one conversation at a full 8K context takes 0.625 GiB on Qwen3-14B, 1 GiB on Qwen3-32B and 1.25 GiB on Llama 3.3 70B. Each figure is 2 × layers × KV heads × head dimension × 8,192 tokens at one byte, from the model’s configuration file.
| MODEL, FP8 WEIGHTS | WEIGHTS | KV PER 8K SESSION | RTX PRO 6000, 96 GB | H200 NVL, 141 GB |
|---|---|---|---|---|
| Qwen3-14B | 15.2 GiB | 0.625 GiB | 108 | 173 |
| Qwen3-32B | 32.0 GiB | 1 GiB | 51 | 91 |
| Llama 3.3 70B | 67.7 GiB | 1.25 GiB | 12 | 44 |
Concurrent sessions at a full 8K context, FP8 KV cache. Weights: Qwen’s FP8 checkpoints (16.3 and 34.3 GB) and NVIDIA’s Llama 3.3 70B FP8 checkpoint (72.7 GB) on Hugging Face; layers and KV heads from each config.json.
These are sessions at full length at the same moment; shorter conversations leave room for more, and response time can set a lower limit. vLLM keeps the cache in the model’s 16-bit type unless --kv-cache-dtype fp8 is set, and a 16-bit cache halves every count. On a test card, vLLM prints the KV cache size it reserves at startup. Thirty users of the 70B model at 8K exceed what one RTX PRO 6000 holds in FP8, and the 70B worked example compares an H200 NVL, two cards and 4-bit weights for that case.
We size inference servers from the model and the number of concurrent users. Tell us the model, its precision, your context length and peak requests in the form below, and we reply with a configuration and quote within one business day.
Which GPU class and how many cards
The GPU class follows from memory per card, bandwidth and features. Bandwidth sets the speed each user sees. For a 70B model in FP8 the arithmetic ceiling for one user is about 23 tokens per second on the RTX PRO 6000 Server Edition and 68 on the H200 NVL.
| GPU | MEMORY, BANDWIDTH | BOARD POWER | TYPICAL USE |
|---|---|---|---|
| NVIDIA L4 | 24 GB, 300 GB/s | 72 W, low profile | embedding and reranking models for RAG; LLMs up to about 14B in FP8, few sessions at that size |
| NVIDIA L40S | 48 GB, 864 GB/s | 350 W | 14B in FP8 for tens of sessions, 32B for a handful; vGPU, no MIG, no NVLink |
| RTX PRO 6000 Server Edition | 96 GB, 1,597 GB/s | up to 600 W | 32B for tens of sessions, 70B in FP8 for a dozen; MIG, vGPU |
| NVIDIA H200 NVL | 141 GB, 4.8 TB/s | up to 600 W | 70B in FP8 for tens of sessions, long contexts, NVLink across 2 or 4 cards |
NVIDIA product pages, read on 6 October 2026; typical use is our reading of the sizing rule above, with FP8 weights and cache.
Server Edition and data-centre cards are passive and need a server chassis. A deskside workstation takes actively cooled cards, the RTX PRO 6000 Workstation or Max-Q, the RTX PRO 5000 or the RTX PRO 4500. While the choice of model is still open, one Spark on a developer’s desk can come first. Its 128 GB of unified memory holds large models, but its 273 GB/s of memory bandwidth, lower than any card in the table, makes it generate more slowly.
Where the model fits one card, add cards as copies, one instance per card behind a load balancer, which adds throughput and, in a second host, removes the single point of failure. Split a model across cards only when it, or its cache for long conversations, does not fit on one. The RTX PRO 6000 and the L40S have no NVLink, so a split model exchanges data over PCIe, while the H200 NVL bridges 2 or 4 cards at 900 GB/s per GPU. NVIDIA’s configuration guide for NVIDIA-Certified Systems calls 2, 4 or 8 GPUs per server balanced.
CPU cores, system RAM and NVMe storage
NVIDIA’s configuration guide for NVIDIA-Certified Systems, updated on 30 September 2026, starts an inference server from two CPU sockets and at least six physical cores per GPU. Each RTX PRO 6000 or H200 NVL should sit on a Gen5 x16 link, an L40S on Gen4 x16 or better. The guide’s minimum for system memory is twice the total GPU memory, spread evenly across all sockets and memory channels. NVIDIA calls these recommendations “a starting point for addressing workload-specific needs”. For two inference cards, the 70B worked example plans one processor and at least the total GPU memory, because the GPUs do the inference and the checkpoints load through host memory. The two-card starter configuration on our AI servers page has one processor and 256 GB. A RAG pipeline that parses documents and holds its vector index on the same server needs more cores and memory.
The guide recommends one NVMe drive per CPU socket, with 1 TB as the minimum. Above that, plan tiers: the operating system on a mirrored pair, a model store for every checkpoint you keep, the RAG index sized from the corpus and, for fine-tuning, a scratch tier for datasets and checkpoints. Llama 3.3 70B in FP8 alone is 72.7 GB, and the same example reserves 2 to 4 TB for its model store, logs and evaluation set.
Network, power and cooling
Users of one inference server need little bandwidth, since streamed text is small; checkpoint copies are the larger transfers, and the 70B worked example uses two 25 GbE ports. For several servers working on one model, NVIDIA’s guide asks for at least 200 Gbps for multi-node inference and up to 400 Gbps per GPU. It also lists a Redfish-compatible management controller and a TPM 2.0 module.
Specify power per rack position, in amps and phases. A 2U server with two 600 W cards and one processor draws about 2 kW at the wall. That fits a 16 A single-phase feed, which carries about 3.7 kW at 230 V. Eight 600 W cards draw 4.8 kW before processors and fans, more than a 16 A feed delivers, so plan three-phase feeds for an eight-card server before the order. Check the PDU outlets too, since a C13 outlet carries 10 A, about 2.3 kW at 230 V, and larger GPU-server power supplies need C19 outlets. On a GPU power cable set for 450 W, an RTX PRO 6000 Server Edition runs capped and an H200 NVL does not boot, so the 600 W cable belongs in the order.
Passive cards depend on the chassis fans and the room, and a server needs about 160 CFM of airflow per kW it draws at an 11 °C rise. Room air without containment reaches its limit at about 20 to 25 kW per rack. Inlet limits are set per server model and GPU, and one 2U model allows 30 °C with the H200 NVL but 25 °C with the RTX PRO 6000 Server Edition.
Software stack, licences and support
Decide whether models run on bare metal or in virtual machines, because the licences follow from it. An open-source engine such as vLLM, under the Apache 2.0 licence, needs no NVIDIA licence on bare metal. NVIDIA AI Enterprise is licensed per GPU, for every GPU in a server that hosts its software. NVIDIA’s NIM FAQ says “Using NIM in production requires an NVIDIA AI Enterprise license”, and compute vGPU profiles in virtual machines enforce the licence in software. Of the RTX PRO 6000 editions, only the Server Edition supports vGPU. As of October 2026, each H200 NVL includes a five-year subscription, activated with its serial number; for the L40S, L4 and RTX PRO cards the licence is bought per GPU. NVIDIA supports AI Enterprise on NVIDIA-Certified Systems, so if you need its support or the H200 NVL subscription, specify a server model that NVIDIA lists as certified with your cards.
For the hardware, ask for the warranty terms per component, who handles a replacement and a burn-in report showing the server ran under load before shipping. Decide who operates it afterwards and whether commissioning on site is part of the order.
What to send with a quote request
These ten facts are enough to size and quote a server, and an estimate marked as such will do where a figure is still open.
- Workload: assistant inference, RAG, fine-tuning, or which of them share the server.
- Model: the model or shortlist, its size and the precision it will run at.
- Context: the context length you will declare per request.
- Concurrency: peak requests in flight, now and as you expect them to grow.
- Response targets, if any: time to first token and tokens per second per user.
- Data: RAG corpus size, fine-tuning datasets and how many model versions you keep.
- Site: rack position and depth, feed in amps and phases, PDU outlets and inlet temperature, or an EU data centre instead.
- Network: uplink ports and speed, and whether several servers work on one model.
- Software: bare metal or virtual machines, the serving engine, and whether NIM, vGPU or NVIDIA support is needed.
- Support: warranty, commissioning on site and who runs the server after delivery.
We check the rack, power and airflow before we quote. Send us the ten facts through the form below, estimates included, and the configuration and quote follow within one business day.
What we supply
We build AI servers to order around the workload, assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU, on one EU contract and invoice. The cards include the RTX PRO 6000 Server and Workstation editions, the H200 NVL with NVLink bridges, the L40S and the L4, with NVIDIA AI Enterprise and vGPU licences on the same invoice. We also build on your own chassis or with parts you already own, after a compatibility check of the platform, power and cooling. Operating system, drivers, CUDA and a container runtime are installed on request; models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What should I specify when I buy an AI server?
What does an AI server for a business need besides GPUs?
How do I size a GPU server for an LLM?
Which GPU should a private AI server have?
Can an on-premise AI server run in an existing server room?
Do I need an NVIDIA AI Enterprise licence for an AI server?
Send us the workload, the model and its precision, the context length, the peak number of concurrent requests and the rack position’s power feed, outlets and inlet temperature. We reply within one business day with a configuration and a quote in writing, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day