Small language models on L4 and RTX PRO 2000: GPU sizing for branch offices and edge servers
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- By our memory estimate, a 24 GB L4, RTX PRO 4000 SFF or RTX PRO 4000 holds each of ten current 2B to 14B models with at least two conversations of 8,192 tokens, and eight of them with at least two of 32,768 tokens
- The 16 GB RTX PRO 2000 holds Gemma 4 E2B, Qwen3.5-4B, Ministral 3 3B, Nemotron 3 Nano 4B, Gemma 4 12B in Google’s 4-bit QAT version and the 8B models in FP8, the 8B models for one or two short conversations only
- Hybrid models keep far less KV cache than dense 8B models: 0.30 GiB per 8K conversation on Qwen3.5-4B and 0.14 GiB on Nemotron 3 Nano 4B, against 1 GiB on Llama 3.1 8B and 1.06 GiB on Ministral 3 8B
- The L4 is a passive single-slot low-profile card at 72 W for servers; the RTX PRO 4000 SFF is an actively cooled low-profile dual-slot card at 70 W, and the RTX PRO 2000 is a 70 W dual-slot card of the same 2.7 by 6.6 inch size; the RTX PRO 4000 is a 145 W single-slot card with 672 GB/s
- vLLM’s quantisation table marks llm-compressor FP8 and the AWQ and GPTQ 4-bit formats as supported on Ada, the L4’s architecture; the three RTX PRO Blackwell cards also compute FP4, which the L4 does not
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Small language models on one low-power GPU
Small language models of about 2 to 14 billion parameters run on one GPU of 70 to 145 W. By our memory estimate, a 24 GB card (the NVIDIA L4, the RTX PRO 4000 SFF or the RTX PRO 4000) holds each of the ten models in this article with at least two conversations of 8,192 tokens, and eight of them with at least two of 32,768 tokens. The 16 GB RTX PRO 2000 holds the 2B to 4B models, Gemma 4 12B in Google’s 4-bit QAT version and an 8B model in FP8, the 8B model for one or two short conversations only.
These are memory ceilings from the publishers’ Hugging Face files, not measurements. We use the rule of our guide to how much VRAM an LLM needs: 90 per cent of the card’s memory, less 3 GiB, for weights and KV cache. That leaves 18.6 GiB on a 24 GB card and 11.4 GiB on a 16 GB card. The context length you allow per conversation changes the result more than the parameter count does.
L4, RTX PRO 2000, RTX PRO 4000 SFF and RTX PRO 4000
| CARD | MEMORY | BANDWIDTH | POWER | FORMAT, COOLING |
|---|---|---|---|---|
| NVIDIA L4 | 24 GB GDDR6 | 300 GB/s | 72 W | low profile, single slot, passive |
| RTX PRO 2000 Blackwell | 16 GB GDDR7 | 288 GB/s | 70 W | 2.7 by 6.6 in, dual slot, active |
| RTX PRO 4000 SFF Blackwell | 24 GB GDDR7 | 432 GB/s | 70 W | low profile, dual slot, active |
| RTX PRO 4000 Blackwell | 24 GB GDDR7 | 672 GB/s | 145 W | full height, single slot, active |
NVIDIA product pages for the four cards and NVIDIA’s RTX PRO 2000 datasheet (June 2026), read on 10 October 2026; the L4’s passive cooling from Lenovo’s ThinkEdge SE455 V3 product guide. All four have ECC memory; none of the three 70 W cards supports MIG.
NVIDIA describes the L4 as “1-slot low-profile” and says it operates “in a 72W low-power envelope”. It is a passive data-centre card that takes its airflow from the server’s fans. The RTX PRO 2000 and RTX PRO 4000 SFF measure 2.7 by 6.6 inches, take two slots and have their own fans, so they suit compact desktops as well as servers that support them. NVIDIA calls the SFF card low profile. For the RTX PRO 2000 its page and datasheet give only the size, and the datasheet lists PCIe 5.0 x8 on a full-length connector, so check that the maker lists the card for the chassis and slot. The RTX PRO 4000 draws twice the L4’s power and has 2.2 times its bandwidth, in a full-height single-slot card. Our comparison of the RTX PRO 4000 SFF, RTX PRO 2000 and L4 covers slot power, display outputs, video engines and vGPU.
NVIDIA rates the three Blackwell cards in “AI TOPS”, which its pages and datasheets give as FP4 TOPS with sparsity: 545 for the RTX PRO 2000, 770 for the RTX PRO 4000 SFF and 1,290 for the RTX PRO 4000. The L4 has FP8 Tensor Cores, rated at 485 TFLOPS with sparsity and half that without, and no FP4.
Which small models fit in 16 and 24 GB
The KV cache per conversation comes from the model’s config.json, as 2 × layers × KV heads × head dimension × bytes per value for every layer that keeps the whole context, plus a fixed amount for layers that keep a window or a state.
| MODEL, FORMAT | WEIGHTS | CACHE PER 8K | 16 GB: 8K / 32K | 24 GB: 8K / 32K |
|---|---|---|---|---|
| Gemma 4 E2B, BF16 | 10.2 GB | 0.05 GiB | 36 / 9 | 172 / 47 |
| Gemma 4 E4B, BF16 | 16.0 GB | 0.14 GiB | does not fit | 25 / 7 |
| Gemma 4 12B, QAT W4A16 | 10.3 GB | 0.44 GiB | 4 / 2 | 20 / 11 |
| Qwen3.5-4B, BF16 | 9.3 GB | 0.30 GiB | 9 / 2 | 33 / 9 |
| Qwen3.5-9B, BF16 | 19.3 GB | 0.30 GiB | does not fit | 2 / none |
| Ministral 3 3B, FP8 | 4.7 GB | 0.81 GiB | 8 / 2 | 17 / 4 |
| Ministral 3 8B, FP8 | 10.4 GB | 1.06 GiB | 1 / none | 8 / 2 |
| Ministral 3 14B, FP8 | 15.7 GB | 1.25 GiB | does not fit | 3 / none |
| Llama 3.1 8B, NVIDIA FP8 | 9.1 GB | 1.00 GiB | 2 / none | 10 / 2 |
| Nemotron 3 Nano 4B, FP8 | 5.3 GB | 0.14 GiB | 45 / 19 | 97 / 41 |
Our estimates, not measurements: concurrent conversations at a full 8,192 or 32,768 tokens, from 90 per cent of nominal memory less 3 GiB. 16-bit KV cache, except Nemotron 3 Nano 4B in FP8 as NVIDIA’s vLLM command sets it. Weights from the Hugging Face file lists and cache from each config.json, read on 10 October 2026. 16 GB is the RTX PRO 2000; 24 GB is the L4, RTX PRO 4000 SFF or RTX PRO 4000.
The Gemma figures at 32K are those of our Gemma 4 hardware guide, which explains the sliding-window layers and why its 12B figure is an upper bound.
The other hybrid models keep little cache as well. Qwen3.5-4B and Qwen3.5-9B each have 32 layers, of which 8 use full attention with four KV heads of dimension 256, 32 KiB per token in 16-bit; the 24 Gated DeltaNet layers hold a fixed state of about 48 MiB in float32. Nemotron 3 Nano 4B keeps a cache in 4 of its 42 layers, and our Nemotron hardware guide gives its state. The dense models keep a cache in every layer: 128 KiB per token on Llama 3.1 8B (32 layers, eight KV heads) and 136 KiB on Ministral 3 8B (34 layers), four times the Qwen3.5 figure.
Qwen3.5-9B is a case where the weights leave too little room. Its BF16 checkpoint of 19.3 GB leaves 0.6 GiB of the 18.6 GiB budget on a 24 GB card, and its card links to 558 quantised versions, which we did not evaluate. Mistral publishes the Ministral 3 instruct models in FP8, and its cards say the 8B is “capable of fitting in 12GB of VRAM in FP8” and the 3B “in 8GB”. Our arithmetic agrees: the 8B with one conversation of 8,192 tokens comes to about 10.8 GiB.
Gemma 4, Qwen3.5-4B and 9B and Ministral 3 are published under Apache 2.0. Llama 3.1 8B comes under the Llama 3.1 Community License and Nemotron 3 Nano 4B under the NVIDIA Nemotron Open Model License. Whether a clause applies to your company is a legal assessment for your legal department.
We supply the L4, RTX PRO 2000, RTX PRO 4000 and RTX PRO 4000 SFF as cards or in servers built to order. Tell us the model, its context length and the users per site at peak, and we name the card that holds them.
Context length and thinking mode decide the user count
A serving engine reserves cache for the context you declare. vLLM takes it from the model configuration unless --max-model-len sets it, and Qwen3.5 and Ministral 3 declare 262,144 tokens. At that length one Ministral 3 8B conversation needs 34 GiB of 16-bit cache by our arithmetic, more than any card in this article holds. For classification or extraction of short texts, declare 8K or less; for a branch assistant that answers from a few retrieved passages, 8K to 32K.
Qwen3.5 models “operate in thinking mode by default”, and Qwen’s card advises “maintaining a context length of at least 128K tokens” to keep that mode working well. At 128K, Qwen3.5-4B in BF16 holds two conversations on a 24 GB card by our estimate. For extraction and routing, thinking can be switched off per request with enable_thinking set to false, which keeps outputs short and the context small. The same card documents --language-model-only, which skips the vision encoder so that its memory goes to the KV cache.
An FP8 KV cache, set in vLLM with --kv-cache-dtype fp8 as in NVIDIA’s command for Nemotron 3 Nano 4B, halves the attention cache per token and nearly doubles the counts of the dense models in the table. The fixed state of the hybrid layers does not shrink. Check engine support on your card first.
FP8 and 4-bit formats on Ada and Blackwell
vLLM’s quantisation documentation, dated 14 September 2026 in its footer, marks llm-compressor FP8 (W8A8) as supported on Ada and Hopper GPUs, and AWQ, GPTQ and the Marlin kernels for GPTQ, AWQ, FP8 and FP4 weights on Ada as well. The L4 is an Ada Lovelace card with no FP4 arithmetic of its own. The FP8 checkpoints of Ministral 3 and Llama 3.1 8B and Google’s QAT W4A16 come in their publishers’ own formats, which the table does not name, and it has no Blackwell column. Confirm each checkpoint with the vLLM release you run on your card.
None of the checkpoints in our table is in NVFP4, NVIDIA’s 4-bit format for Blackwell; our Gemma and Nemotron guides list NVFP4 versions of the larger models in those families.
Generation speed on 288 to 672 GB/s
A dense model reads all its weights from memory for every token it generates, so bandwidth divided by the size of the weights sets a ceiling for the speed one user sees. On that basis the RTX PRO 4000 has a ceiling about 2.2 times that of the L4 for the same model, the RTX PRO 4000 SFF about 1.4 times, and the RTX PRO 2000 about the same as the L4. These are ratios from the specifications, not measured speeds. Several users share each read of the weights, so total throughput rises with concurrent requests until the cache or the compute runs out.
Where a small model is enough, and where it is not
Small models fit work with short inputs and constrained outputs: classifying tickets or e-mails, extracting fields from forms into JSON, routing a request to the right queue or to a larger model, and answering routine questions at a branch from a limited set of documents through RAG. Mistral’s card for Ministral 3 3B calls it “Ideal for lightweight, real-time applications on edge or low-resource devices”. NVIDIA calls Nemotron 3 Nano 4B “an edge-ready small language model”. A position paper on arXiv (2506.02153, revised on 22 September 2026), whose abstract links to an NVIDIA Research page, argues that small models are “sufficiently powerful” for many agentic invocations and that systems with general conversational needs combine several models.
The limits follow from the memory arithmetic. Long contracts, reports read in full or many users with long conversations need more cache than 16 or 24 GB holds beside the weights, and broad knowledge and multi-step reasoning across documents point to a larger model. One layout keeps the small model at the branch for routing and extraction and sends the rest to a larger model in the data centre, on an RTX PRO 6000 or an H200 NVL.
Branch and edge servers: low-profile, SFF and short-depth
For a server at a branch or on a factory floor, the L4 is the one of the four designed for server airflow. Lenovo’s product guide for the ThinkEdge SE455 V3, updated on 25 August 2026, describes a “2U rack server, short depth (438mm depth, from EIA front rack flange)” and lists up to six L4 cards in it. Our comparison of L4 and L40S for inference lists data-centre servers that take up to eight L4 in 2U. Of the three 70 W cards, only the L4 supports vGPU, which lets one branch host give GPU slices to several virtual machines.
Where there is no server room, an RTX PRO 2000 or RTX PRO 4000 SFF in a compact desktop runs a small model from the PCIe slot, provided the maker lists that card, at 70 W and 2.7 inches high, for the chassis. Four L4 cards draw 288 W at their rated power, less than one RTX PRO 6000 Server Edition at 600 W, which matters where a branch rack has a single small feed.
We build GPU workstations and short-depth servers for branch sites to order, and we check the rack, power and airflow before we quote. Describe the branch site, its rack depth or desk space and the power available in the form below.
What we supply
We supply the NVIDIA L4, the RTX PRO 2000, the RTX PRO 4000 and the RTX PRO 4000 SFF as cards, in configured workstations or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We size the card from the model, its precision, the context length and the users per site, and we confirm that the chassis takes the card in that slot and can cool it. vGPU licences for L4 hosts come on the same invoice as the hardware. Our professional GPU range lists every card, and the models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What GPU do I need for an 8B model?
Can an LLM run on an NVIDIA L4?
Can the RTX PRO 2000 run an LLM?
What is a small language model?
How much VRAM does a small language model need?
Which GPU is suitable for an edge LLM server?
Send us the model, its precision, the context length, the number of sites and the users per site at peak, and whether the card goes into a server or a compact desktop. We reply within one business day with the card that fits the slot, the cooling and the memory you need, and a written quote.
Talk to an expertWe reply within one business day