Serving multiple LLMs on one GPU server: vLLM instances, MIG, LoRA adapters and sleep mode
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- vLLM’s FAQ says one OpenAI-compatible server serving several models at once “is not currently supported”; each model runs as its own instance, and a router in front exposes all of them by name on one endpoint
- Instances are placed on whole cards with CUDA_VISIBLE_DEVICES, on MIG instances, or several on one card, where
--gpu-memory-utilizationis “a per-instance limit” and vLLM’s own example gives two instances 0.5 each - On one RTX PRO 6000, two Qwen3-14B instances at 0.45 each hold about 39 conversations of 8K each in an FP8 cache by our estimate; MIG adds hardware isolation in fixed slices of 24 GB, or 18 and 35 GB on the H200 NVL
- Fine-tuned variants of one base model share a single instance as LoRA adapters, chosen per request in the model field;
max_loras, the number of adapters in one batch, defaults to 1 in vLLM - vLLM’s sleep mode at level 1 moves the weights to CPU memory and drops the cache; a post on vLLM’s blog reports a 0.26 s wake for Qwen3-0.6B on an A100, against 37.6 s per switch without sleep mode
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
How to serve multiple LLMs on one GPU server
To serve multiple LLMs on one GPU server with vLLM, you run one instance per model, because one vLLM server process serves one base model. vLLM’s FAQ says that serving several models at once from its OpenAI-compatible server “is not currently supported” and points to running several instances, each with its own model, behind “another layer to route the incoming request”. On a GPU server that leaves five documented ways to place several models: one instance per card or set of cards, several instances sharing one card through memory fractions, one instance per MIG instance, many LoRA adapters on one base model, and models swapped in and out with vLLM’s sleep mode or a model manager such as Triton’s. A router in front exposes every model by name on one OpenAI-compatible endpoint.
| METHOD | ISOLATION | MEMORY | SWITCH TIME | DOCUMENTED BY |
|---|---|---|---|---|
| Whole cards per instance | own process and cards per model | the card’s memory beyond the weights goes to the cache | none, all models stay loaded | CUDA guide, vLLM FAQ |
| Shared card, fractions | own process; compute shared in time, no fault isolation | a limit each instance sets for itself | none, all models stay loaded | vLLM engine arguments, GPU Operator docs |
| MIG instances | hardware: dedicated compute and memory | fixed slices: 24 GB on RTX PRO 6000, 18 or 35 GB on H200 NVL | relayout only with the card idle | NVIDIA MIG guide |
| LoRA adapters | one process and one base model | adapter slots set by max_loras and max_ | per request | vLLM and NIM LoRA docs |
| Sleep mode | one process per model, one awake | level 1 keeps the weights in CPU RAM | 0.26 to 2.58 s to wake in vLLM’s test | vLLM docs and blog |
| Triton explicit mode | one Triton process for all models | only the models loaded on request | a full load | Triton model management |
vLLM FAQ, engine arguments, LoRA and sleep mode pages, and vLLM blog of 26 October 2025 (A100, vLLM 0.11.0); CUDA Programming Guide 13.4.2; NVIDIA MIG user guide of 11 September 2026; GPU Operator GPU sharing page of 23 September 2026; NIM for LLMs 2.0.13 docs; Triton model management docs; all read 10 October 2026.
Isolation sets how many models stop when one process or card fails, and every gigabyte that weights or a second instance take is missing from the KV cache.
We build AI servers with 4 or 8 cards for several models at once. Tell us which models you want to serve and how many users each one has, and we reply with a configuration and quote within one business day.
One vLLM instance per card with CUDA_VISIBLE_DEVICES
The CUDA Programming Guide says CUDA_VISIBLE_DEVICES “controls which GPU devices are visible to a CUDA application and in what order they are enumerated”. A vLLM instance started with CUDA_VISIBLE_DEVICES=0 sees only the first card; a second one with CUDA_VISIBLE_DEVICES=1 and its own --port sees only the second. Visible devices are renumbered from 0, so a model on cards 2 and 3 runs with --tensor-parallel-size 2 as if they were the only cards in the server.
--served-model-name sets “The model name(s) used in the API”, and without it the name is the --model argument, so give each instance a short name for the router and the clients.
Each model has its own process and cards, so a restart of one instance leaves the others serving and each cache gets all the memory beyond the weights, but a 14B model with few users still occupies a whole 96 GB card.
Two instances on one card: gpu-memory-utilization
vLLM claims a share of each card’s memory, 0.92 by default, for weights, activations and the KV cache. Its engine arguments describe the setting as “a per-instance limit, and only applies to the current vLLM instance”, and give the example that with two instances on the same GPU “you can set the GPU memory utilization to 0.5 for each instance”. The alternative is --kv-cache-memory-bytes, the “Size of KV Cache per GPU in bytes”, which when set “ignores gpu_memory_utilization”.
Our sizing guides give the KV cache 0.9 × the memory the driver reports, minus 3 GiB, minus the weights. Split into two instances of 0.45 on an RTX PRO 6000 Server Edition, which reports 95.6 GiB, each instance has 43.0 GiB. Qwen3-14B in FP8 takes 15.2 GiB of weights, and the 3 GiB overhead applies to each instance, which leaves 24.8 GiB of cache, room for about 39 conversations of 8K each at 0.625 GiB in an FP8 cache. Alone on the card, the model holds about 108. Two instances hold about 78 together, because the second set of weights and overhead comes out of the cache.
The fraction is a limit that each vLLM instance applies to itself, not a partition of the card. Both processes run on the same streaming multiprocessors and memory bandwidth, which the GPU shares between them in time, so a burst on one model slows the other. NVIDIA’s GPU Operator documentation says of this time-slicing that, unlike MIG, “there is no memory or fault-isolation between replicas”. A GPU reset stops both as well, since nvidia-smi resets a card only when “There can’t be any applications using these devices”. Time-slicing, MPS and MIG in Kubernetes are compared in our guide to GPU sharing in Kubernetes.
MIG instances for small models with dedicated memory
MIG partitions a GPU “into multiple isolated instances, each with dedicated compute and memory resources”, in the words of NVIDIA’s MIG user guide. Its profile tables give the RTX PRO 6000, in all three editions, up to four 1g.24gb instances or two 2g.48gb. For the H200 141 GB, which the guide’s GPU list also names as H200 NVL, they give seven 1g.18gb, four 1g.35gb, three 2g.35gb or two 3g.71gb. One vLLM instance runs in each MIG instance, selected through CUDA_VISIBLE_DEVICES with the MIG UUID; the CUDA guide adds that “Only single MIG instance enumeration is supported”, so a model must fit into one instance.
The slice sizes decide which models fit. An 8B embedding model has 15.1 GB of weights in BF16 and fits a 1g.24gb instance of an RTX PRO 6000 or a 1g.35gb instance of an H200 NVL, but not a 1g.18gb one once the CUDA context and activations are added, while a 0.6B embedder at 1.2 GB fits any of them. NVIDIA’s MIG guide lists reconfiguration as possible “When Idle”, so the layout changes only with no workloads on the card, and our MIG runbook for RTX PRO Blackwell gives the steps.
Many fine-tuned variants: multi-LoRA on one base model
Where several models are fine-tuned versions of one base model, one instance serves all of them as LoRA adapters. vLLM starts with --enable-lora and registers each adapter as name=path through --lora-modules. Adapters then appear in /v1/models beside the base model, and a client chooses one “as if it were any other model via the model request parameter”. max_loras, the “Max number of LoRAs in a single batch”, defaults to 1, so requests for different adapters are not batched together until it is raised to the number of adapters in use at once. max_lora_rank defaults to 16 and must be at least the highest rank among the adapters, and max_cpu_loras sets how many adapters wait in CPU memory.
Adapters can also be loaded while the server runs, with VLLM_ and the endpoint /v1/load_lora_adapter. vLLM warns that this “comes with security risks” and should not be used in production unless in “an isolated, fully trusted environment”. NVIDIA’s NIM for LLMs 2.0.13, built on vLLM, reads adapters from the directory in NIM_PEFT_SOURCE, rescans it when NIM_PEFT_REFRESH_INTERVAL is set, and states: “You can serve multiple adapters at the same time, subject to GPU memory limits.” NIM in production needs an NVIDIA AI Enterprise licence per GPU.
Swapping models: vLLM sleep mode, Triton and Ollama
Models used a few times a day need not hold GPU memory. vLLM’s sleep mode, enabled with --enable-sleep-mode on CUDA and ROCm, has two levels. At level 1 “The model weights are backed up in CPU memory” and the KV cache is discarded; at level 2 the weights are discarded too and must be reloaded. The server endpoints /sleep, /wake_up and /is_sleeping exist only with VLLM_SERVER_DEV_MODE=1, and vLLM says they “should not be exposed to users”.
A post on vLLM’s blog of 26 October 2025 reports switch times on an A100 with vLLM 0.11.0. Qwen3-0.6B woke from level 1 in 0.26 s, against an average of 37.6 s per switch without sleep mode, and Phi-3-vision in 0.82 s against 58.1 s. From level 2 the wake took 0.85 and 2.58 s. We found no published figures for the RTX PRO 6000 or the H200 NVL. Level 1 needs host memory for every sleeping model’s weights, 67.7 GiB for Llama 3.3 70B in FP8; the four-card inference node on our AI servers page has 512 GB of DDR5 ECC memory.
In the EXPLICIT model control mode of Triton Inference Server, “Triton loads only those models specified explicitly with the --load-model command-line option” at start-up and loads or unloads others on request through its API. By default, Ollama keeps a model loaded for 5 minutes after use and holds up to three times as many models as there are GPUs; when a new one does not fit, requests queue while idle models unload. Our comparison of vLLM, SGLang, TensorRT-LLM and Ollama covers the engines.
One endpoint for all models: routing by model name
Every OpenAI-style request names its model in the model field, and a router forwards it to the instance that serves that model. In Kubernetes, the Gateway API Inference Extension documents this as body-based routing, which “extracts the model name from the request body and adds it to the X-Gateway-Model-Name header”, on which the gateway routes. The guide warns that “LoRA names must be unique across the base AI models”, since adapters share their base model’s server. Outside Kubernetes, an OpenAI-compatible proxy does the same from a list of model names and backends.
The router is also where keys per team, rate limits and the query log belong, because vLLM’s own API key covers only part of its endpoints. Our guide to a private LLM platform on Kubernetes covers the gateway, the endpoint picker and their versions.
Example placements on a 4-card and an 8-card server
The examples use the sizing rule above, 3 GiB of overhead per instance, an FP8 cache and conversations of 8K at full length. A company plan, which department gets which cards, is in our allocation guide for eight RTX PRO 6000.
| SERVER AND CARDS | PLACEMENT | METHOD | CONVERSATIONS OF 8K |
|---|---|---|---|
| 4 × RTX PRO 6000, card 0 | Llama 3.3 70B, FP8 | whole card | about 12 |
| 4 × RTX PRO 6000, card 1 | Qwen3-32B, FP8, with LoRA adapters | whole card, multi-LoRA | about 51, less adapter slots |
| 4 × RTX PRO 6000, card 2 | two models of Qwen3-14B size, FP8 | two instances at 0.45 | about 39 each |
| 4 × RTX PRO 6000, card 3 | 8B embedder, 8B reranker, 0.6B embedder, test slot | four 1g.24gb MIG instances | not a chat model |
| 8 × H200 NVL, cards 0 to 3 | one model too large for one card | tensor parallel in one NVLink domain | sized on its model page |
| 8 × H200 NVL, cards 4 and 5 | Llama 3.3 70B, FP8 | one copy per card behind the router | about 44 per card |
| 8 × H200 NVL, card 6 | Qwen3-32B, FP8, with LoRA adapters | whole card, multi-LoRA | about 91, less adapter slots |
| 8 × H200 NVL, card 7 | embedders, reranker, test slot | four 1g.35gb MIG instances | not a chat model |
Our estimates: driver-visible memory 95.6 GiB on the RTX PRO 6000 Server Edition and 140.4 GiB on the H200 NVL; weights 67.7 GiB (Llama 3.3 70B), 32.0 GiB (Qwen3-32B) and 15.2 GiB (Qwen3-14B) from the FP8 checkpoints on Hugging Face; MIG profiles from NVIDIA’s MIG user guide.
The two Llama copies keep that model serving if one card fails. On the RTX PRO 6000 server every model sits on one card, since the cards have no NVLink and a split model would exchange its data over PCIe.
We check the rack, power and airflow before we quote. Send us your models, their precision and the users per model through the form below.
What we supply
We build AI servers to order for multi-model serving, with 4 or 8 cards: the RTX PRO 6000 Server Edition, which supports MIG and vGPU, the H200 NVL with NVLink bridges, and the L40S or L4 for small models, each with manufacturer warranty on one EU contract and invoice. The GPUs are also available on their own for servers you already run. On request the servers arrive with the operating system, drivers, CUDA and a container runtime installed, and NVIDIA AI Enterprise licences for NIM come on the same invoice. The serving platform on top, with private models on vLLM or NVIDIA AI Enterprise and a query log, is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
Can vLLM serve multiple models at the same time?
How do I run two vLLM instances on one GPU?
--gpu-memory-utilization, which vLLM describes as a per-instance limit; its documentation gives two instances 0.5 each as the example. Both instances share the card’s compute and memory bandwidth in time, without memory or fault isolation, so one model’s load slows the other, and a GPU reset stops both.What is vLLM sleep mode?
VLLM_SERVER_DEV_MODE=1 and should not be exposed to users.How does multi-LoRA serving work in vLLM?
--enable-lora, and each adapter is registered by name with --lora-modules. Adapters appear in /v1/models beside the base model, and a request chooses one through its model field. max_loras sets how many adapters can be in one batch and defaults to 1, while max_lora_rank, 16 by default, must be at least the highest rank among your adapters.Should I use MIG or shared memory fractions for multiple models on one GPU?
--gpu-memory-utilization allow any split but no memory or fault isolation. MIG suits embedders, rerankers and other small services that need dedicated resources; fractions suit two small models with light, predictable traffic.How do I serve several models behind one OpenAI-compatible endpoint?
--served-model-name in vLLM or NIM_SERVED_MODEL_NAME in NIM. In Kubernetes, the Gateway API Inference Extension does this with body-based routing, which copies the model name into a header the gateway routes on.Send us the models you want to serve, their precision and context length, and the peak number of concurrent users per model. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day