GPU cluster without InfiniBand for LLM inference: Ethernet, RoCE and when multi-node matters
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Most models a company of 500 to 2,000 staff runs fit in one server of eight RTX PRO 6000 (768 GB) or eight H200 NVL (1,128 GB), so two to four servers run as independent replicas behind a load balancer and need no InfiniBand
- Replicas exchange nothing with each other; requests, model downloads and monitoring run over ordinary Ethernet, with two 25 GbE ports per server as in our 70B worked example and a 1 Gb management port as in NVIDIA’s reference architecture
- A fast east-west fabric matters when one model spans servers or when KV cache moves between them; NVIDIA’s configuration guide asks for a minimum of 200 Gbps for multi-node inference and InfiniBand or RoCE between the nodes
- vLLM’s documentation sets tensor parallelism to the GPUs per node and pipeline parallelism to the number of nodes; for DeepSeek-V3.2, vLLM passes the hidden state and the residual across a stage boundary, about 57 MB for a 2,000-token prompt, about 18 ms at 25 Gb/s by our arithmetic
- NVIDIA’s RTX PRO AI Factory reference architecture builds on scalable units of four servers with eight RTX PRO 6000 each, four 400 Gb/s SuperNICs for the compute network per server and a separate 1 Gb management network
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Two to four GPU servers without InfiniBand: when Ethernet is enough
For LLM inference on two to four servers, a GPU cluster without InfiniBand is usually enough. Most models that a company of 500 to 2,000 staff runs fit in one server with eight RTX PRO 6000 Server Edition cards (768 GB of GPU memory) or eight H200 NVL (1,128 GB). Each server then runs its own copies of the models behind a load balancer, and the servers exchange no data with each other. Such a cluster needs ordinary Ethernet for requests, model downloads and monitoring. Our 70B worked example plans two 25 GbE ports per server, and NVIDIA’s reference architecture gives the management controller of each node a 1 Gb port.
A fast east-west fabric, RDMA over InfiniBand or RoCE at 200 to 400 Gb/s, matters in two cases: one model spans servers through pipeline, tensor or expert parallelism, or KV cache moves between servers in disaggregated serving. NVIDIA’s configuration guide for NVIDIA-Certified Systems, last updated on 30 September 2026, draws the same line: “Single node deployments typically do not require high-speed networking to connect multiple nodes for your AI workload.” For clustered workloads it connects the nodes with “high-speed networking (either InfiniBand or RoCE)” and asks for a “Minimum 200 Gbps for multi-node inference”.
What crosses the network in each cluster pattern
| PATTERN | CROSSES THE NETWORK | HOW OFTEN | WHAT TO PLAN |
|---|---|---|---|
| Replicas behind a balancer | requests, streamed answers, metrics | every request, kilobytes of text | 2 × 25 GbE per server, 1 GbE management |
| Model loading | checkpoint files from the model store | at start, on update, after a failure | 25 to 100 GbE to the store, or a local NVMe copy |
| Pipeline parallel, 2 nodes | activations at the stage boundary | once per boundary and step | RDMA, at least 200 Gbps per NVIDIA’s guide; light traffic |
| Tensor parallel across nodes | partial results, all-reduced | twice per layer, every token | vLLM’s docs keep it inside a node; up to 400 Gbps per GPU |
| Expert parallel across nodes | tokens, all-to-all | every MoE layer | RDMA; vLLM’s DeepEP backends for multi-node |
| Disaggregated serving | KV cache of each request | once per request, prefill to decode | NIXL over RDMA; TCP for tests only |
vLLM documentation on parallelism and scaling (6 May 2026), expert parallel deployment (2 October 2026) and the NixlConnector (1 October 2026); NVIDIA’s configuration guide for NVIDIA-Certified Systems (30 September 2026); port counts from our 70B worked example and NVIDIA’s RTX PRO AI Factory reference architecture.
vLLM’s documentation on parallelism starts from the single card: “if the model fits on a single GPU, distributed inference is probably unnecessary.” For several servers it says: “Set tensor_parallel_size to the number of GPUs per node and pipeline_parallel_size to the number of nodes.” Only pipeline traffic then crosses the network. The same page also shows tensor parallelism over all GPUs of the cluster, and notes that “Efficient tensor parallelism requires fast internode communication”. One 400 Gb/s port carries 50 GB/s in each direction, less than the 64 GB/s of a PCIe 5.0 x16 slot, and every switch hop adds latency to the two all-reduces that tensor parallelism needs per layer. Our guide to one model on several GPUs over PCIe and NVLink covers that traffic.
Independent replicas behind a load balancer
Our guide to high availability for an on-premise LLM sizes a company of 2,000 staff with a peak of 80 requests in flight on gpt-oss-120b at a 32K context. Two servers with five RTX PRO 6000 or two H200 NVL each carry that peak alone if the other fails, and three servers need three RTX PRO 6000 or one H200 NVL each. In both layouts every server runs complete replicas, and nothing passes between the GPU servers.
The network of such a cluster carries requests and streamed answers, which are text, the calls of a RAG pipeline to the embedding service and the vector database, and metrics. Our 70B worked example plans two 25 GbE ports per server, and the second keeps backup traffic and model pulls off the user network. The management controller gets its own 1 GbE port on a separate management network.
Eight cards in one server or four in each of two is a question of failure domains, covered in our comparison of one 8-GPU server and two 4-GPU servers.
We build AI servers with 2 to 8 GPUs per node and fit the NICs, 25 to 400G, to the cluster pattern. Tell us how many servers you plan and which models each one serves through the form below.
When one model spans two servers
One server is no longer enough when the weights and the cache of the peak do not fit its cards. Our DeepSeek hardware guide gives an example: DeepSeek-V3.2 leaves 2.7 GiB per card on eight RTX PRO 6000, about one conversation at 32K, while eight H200 NVL in tensor parallelism leave 43 GiB per card, about 18 conversations. A model of that size either moves to the larger card or spans two servers, and the network between them then carries traffic of every request.
Pipeline parallelism sends the least. Only the activations at the boundary between two stages cross the network. DeepSeek-V3.2’s config.json gives a hidden size of 7,168, and vLLM’s DeepSeek code passes two tensors of that size per token to the next stage, the hidden state and the residual, so in BF16 one token takes 28 KiB. A 2,000-token prompt sends about 57 MB across the boundary, about 18 ms at a 25 Gb/s line rate and about 2.3 ms at 200 Gb/s by our arithmetic. Each generation step then sends 28 KiB per request.
Mixture-of-experts models add a third option, expert parallelism, which vLLM pairs with data parallelism. Across servers, expert parallelism sends tokens all-to-all at every MoE layer. vLLM’s expert parallel documentation lists the DeepEP backends for “Multi-node prefill” and “Multi-node decode” and has troubleshooting notes for InfiniBand and RoCE. This is the pattern that needs RDMA.
Disaggregated serving moves the KV cache instead. vLLM’s NixlConnector is “a high-performance KV cache transfer connector for vLLM’s disaggregated prefilling feature”, with UCX as its default transport; when to split prefill from decode is the subject of our guide to disaggregated serving with NVIDIA Dynamo.
RoCE or InfiniBand for the east-west fabric
InfiniBand is a separate fabric with its own switches, and the adapters run in InfiniBand mode. It needs a subnet manager, and NVIDIA’s DOCA documentation (v3.5.0) states: “One SM must be running for each InfiniBand subnet.” NCCL, the library through which vLLM exchanges data between GPUs, reaches the adapters through IB verbs, and NCCL_IB_HCA sets which adapters it uses.
RoCE runs RDMA over the Ethernet switches. NVIDIA’s Cumulus Linux 5.16 documentation says that “RoCE uses the Infiniband (IB) Protocol over converged Ethernet” and that “RoCEv2 requires flow control for lossless Ethernet”. Its default RoCE mode is lossless, with priority flow control and ECN on every switch the traffic crosses, and a lossy mode relies on ECN alone. NCCL uses the same verbs and picks the address with NCCL_IB_GID_INDEX, which NCCL 2.32.3 documents as “the Global ID index used in RoCE mode”, with AUTO as the default since 2.32u1.
NVIDIA’s own Ethernet option is Spectrum-X, which, in NVIDIA’s words, “combines purpose-built Ethernet switches and SuperNICs”; NVIDIA writes that its SuperNICs provide “RDMA over Converged Ethernet (RoCE) network connectivity between GPU servers”. The BlueField-3 SuperNIC in NVIDIA’s reference architecture is listed as “400GbE (default mode) /NDR IB”, so the same card serves either fabric.
What NVIDIA’s RTX PRO AI Factory reference architecture specifies
NVIDIA’s RTX PRO AI Factory enterprise reference architecture, last updated on 18 May 2026, “is based on a 2-8-5-200 infrastructure configuration (2 CPUs, 8 GPUs, 5 NICs at 200 Gbps each)”, with up to eight RTX PRO 6000 Blackwell Server Edition cards per server. Of the five NICs, four BlueField-3 SuperNICs serve the compute (east-west) network and one BlueField-3 DPU the converged (north-south) network. For the compute network the architecture gives four 200 Gb/s or two 400 Gb/s NICs as the minimum and four 400 Gb/s NICs as recommended, and its scalable units use the recommended four.
It builds clusters from “scalable units (SU) based on 4 compute nodes”. Per scalable unit, the compute network has “16x 400Gb/s connections”, four per server, and the converged network “8x 200Gb/s connections” used “for node communications with compute along with storage, in-band management, and end-user connections”. A 1 Gb/s out-of-band network reaches every node. The fabric is Ethernet throughout, with SN5610 switches of 128 ports at 400 GbE.
Four servers of eight cards are one scalable unit. By our reading, its compute network carries no inference traffic while every server runs its own replicas, and it can be added later if each chassis keeps PCIe slots for the NICs. NVIDIA’s configuration guide asks that “NICs and NVMe drives should be placed within the same PCIe switch or root complex as the GPUs”.
Model loading, storage and the management network
Model files are the largest transfers in a replica cluster. DeepSeek-V4-Flash-0731 is 167 GB, or 1,335 Gbit. At line rate that takes about 53 seconds at 25 Gb/s, 13 seconds at 100 Gb/s and under 7 seconds at 200 Gb/s by our arithmetic, before anything loads into GPU memory. Every replica restart reads them again, so our high-availability guide keeps a pinned copy on local NVMe in each server.
The management controller belongs on a separate network. NVIDIA’s configuration guide asks for a controller that is “Redfish 1.0 (or greater) compatible”, and the reference architecture connects every node at 1 Gb/s through 48-port management switches.
The east-west network needs isolation of its own. vLLM’s security documentation, updated on 9 October 2026, states that “All communications between nodes in a multi-node vLLM deployment are insecure by default” and “must be protected by placing the nodes on an isolated network”, and that “Inter-node communication is unencrypted by default.”
Checking a multi-node vLLM set-up before go-live
When a model spans servers, confirm that NCCL uses RDMA before any load test.
- Install the same image on every server; vLLM asks that “every node provides an identical execution environment, including the model path and Python packages”.
- Set
VLLM_HOST_IPto each server’s address on the isolated east-west network,NCCL_SOCKET_IFNAMEto that interface andNCCL_IB_HCAto the RDMA adapters. - For RoCE, check that every switch between the servers runs the same RoCE mode, lossless with priority flow control and ECN, or lossy with ECN.
- Run the all-reduce test from NVIDIA’s nccl-tests between the servers and compare the bus bandwidth it reports with the line rate of the NICs.
- Start vLLM with
NCCL_DEBUG=TRACEand look for “[send] via NET/IB/GDRDMA” in the log; “[send] via NET/Socket” means plain TCP, which vLLM calls “not efficient for cross-node tensor parallelism”.
Layouts for two, three and four servers
| SERVERS | LAYOUT | NETWORK PER SERVER | EXAMPLE |
|---|---|---|---|
| Two | replicas, each server holds the whole peak | 2 × 25 GbE, 1 GbE management | 5 RTX PRO 6000 or 2 H200 NVL each for 80 requests in flight |
| Three | replicas, each server holds half the peak | 2 × 25 GbE, 1 GbE management | 3 RTX PRO 6000 or 1 H200 NVL each for the same peak |
| Four | replicas, or two servers that share one large model | 2 × 25 GbE, plus RDMA at 200 to 400 Gb/s on the servers that split a model | one scalable unit in NVIDIA’s reference architecture |
| Two, one model across both | pipeline or expert parallelism over 16 cards | RDMA at 200 to 400 Gb/s on an isolated network | a large MoE model beyond eight RTX PRO 6000 |
Card counts from our high-availability guide (gpt-oss-120b at 32K, 16-bit cache, peak of 80 requests in flight); 25 GbE ports from our 70B worked example; network minimums from NVIDIA’s configuration guide (30 September 2026) and RTX PRO AI Factory reference architecture (18 May 2026).
Four servers leave room to give two of them to one large model while the other two run replicas, and only the two that share a model need the RDMA fabric.
We check the rack, power and airflow before we quote, and the configuration and quote follow within one business day. Send us the models, the peak requests in flight and the number of servers through the form below.
What we supply
We build AI servers to order with 2 to 8 GPUs per node, sized by model size and concurrent users, assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU. The cards include the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S and the L4, and each server comes with 25, 100, 200 or 400G NICs and out-of-band management sized to the cluster pattern. NVIDIA AI Enterprise and vGPU licences come on the same EU contract and invoice. The platform on top, deployment of the models on-premise with vLLM on a Kubernetes-based platform, is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
Can a GPU cluster for LLM inference run without InfiniBand?
Can vLLM run across multiple nodes over Ethernet?
RoCE or InfiniBand for a GPU inference cluster?
What network does a GPU inference cluster of four servers need?
How much network bandwidth does multi-node LLM inference need?
How do I check that vLLM uses RDMA between nodes?
Send us the models you plan to serve, the peak number of requests in flight, the number of servers you have in mind and the network ports at the rack. We reply within one business day with a configuration per server, cards and NICs included, and a written quote.
Talk to an expertWe reply within one business day