BLOG · GUIDE ·

Small language models on L4 and RTX PRO 2000: GPU sizing for branch offices and edge servers

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • By our memory estimate, a 24 GB L4, RTX PRO 4000 SFF or RTX PRO 4000 holds each of ten current 2B to 14B models with at least two conversations of 8,192 tokens, and eight of them with at least two of 32,768 tokens
  • The 16 GB RTX PRO 2000 holds Gemma 4 E2B, Qwen3.5-4B, Ministral 3 3B, Nemotron 3 Nano 4B, Gemma 4 12B in Google’s 4-bit QAT version and the 8B models in FP8, the 8B models for one or two short conversations only
  • Hybrid models keep far less KV cache than dense 8B models: 0.30 GiB per 8K conversation on Qwen3.5-4B and 0.14 GiB on Nemotron 3 Nano 4B, against 1 GiB on Llama 3.1 8B and 1.06 GiB on Ministral 3 8B
  • The L4 is a passive single-slot low-profile card at 72 W for servers; the RTX PRO 4000 SFF is an actively cooled low-profile dual-slot card at 70 W, and the RTX PRO 2000 is a 70 W dual-slot card of the same 2.7 by 6.6 inch size; the RTX PRO 4000 is a 145 W single-slot card with 672 GB/s
  • vLLM’s quantisation table marks llm-compressor FP8 and the AWQ and GPTQ 4-bit formats as supported on Ada, the L4’s architecture; the three RTX PRO Blackwell cards also compute FP4, which the L4 does not

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Small language models on one low-power GPU

Small language models of about 2 to 14 billion parameters run on one GPU of 70 to 145 W. By our memory estimate, a 24 GB card (the NVIDIA L4, the RTX PRO 4000 SFF or the RTX PRO 4000) holds each of the ten models in this article with at least two conversations of 8,192 tokens, and eight of them with at least two of 32,768 tokens. The 16 GB RTX PRO 2000 holds the 2B to 4B models, Gemma 4 12B in Google’s 4-bit QAT version and an 8B model in FP8, the 8B model for one or two short conversations only.

These are memory ceilings from the publishers’ Hugging Face files, not measurements. We use the rule of our guide to how much VRAM an LLM needs: 90 per cent of the card’s memory, less 3 GiB, for weights and KV cache. That leaves 18.6 GiB on a 24 GB card and 11.4 GiB on a 16 GB card. The context length you allow per conversation changes the result more than the parameter count does.

L4, RTX PRO 2000, RTX PRO 4000 SFF and RTX PRO 4000

CARDMEMORYBANDWIDTHPOWERFORMAT, COOLING
NVIDIA L424 GB GDDR6300 GB/s72 Wlow profile, single slot, passive
RTX PRO 2000 Blackwell16 GB GDDR7288 GB/s70 W2.7 by 6.6 in, dual slot, active
RTX PRO 4000 SFF Blackwell24 GB GDDR7432 GB/s70 Wlow profile, dual slot, active
RTX PRO 4000 Blackwell24 GB GDDR7672 GB/s145 Wfull height, single slot, active

NVIDIA product pages for the four cards and NVIDIA’s RTX PRO 2000 datasheet (June 2026), read on 10 October 2026; the L4’s passive cooling from Lenovo’s ThinkEdge SE455 V3 product guide. All four have ECC memory; none of the three 70 W cards supports MIG.

NVIDIA describes the L4 as “1-slot low-profile” and says it operates “in a 72W low-power envelope”. It is a passive data-centre card that takes its airflow from the server’s fans. The RTX PRO 2000 and RTX PRO 4000 SFF measure 2.7 by 6.6 inches, take two slots and have their own fans, so they suit compact desktops as well as servers that support them. NVIDIA calls the SFF card low profile. For the RTX PRO 2000 its page and datasheet give only the size, and the datasheet lists PCIe 5.0 x8 on a full-length connector, so check that the maker lists the card for the chassis and slot. The RTX PRO 4000 draws twice the L4’s power and has 2.2 times its bandwidth, in a full-height single-slot card. Our comparison of the RTX PRO 4000 SFF, RTX PRO 2000 and L4 covers slot power, display outputs, video engines and vGPU.

NVIDIA rates the three Blackwell cards in “AI TOPS”, which its pages and datasheets give as FP4 TOPS with sparsity: 545 for the RTX PRO 2000, 770 for the RTX PRO 4000 SFF and 1,290 for the RTX PRO 4000. The L4 has FP8 Tensor Cores, rated at 485 TFLOPS with sparsity and half that without, and no FP4.

Which small models fit in 16 and 24 GB

The KV cache per conversation comes from the model’s config.json, as 2 × layers × KV heads × head dimension × bytes per value for every layer that keeps the whole context, plus a fixed amount for layers that keep a window or a state.

MODEL, FORMATWEIGHTSCACHE PER 8K16 GB: 8K / 32K24 GB: 8K / 32K
Gemma 4 E2B, BF1610.2 GB0.05 GiB36 / 9172 / 47
Gemma 4 E4B, BF1616.0 GB0.14 GiBdoes not fit25 / 7
Gemma 4 12B, QAT W4A1610.3 GB0.44 GiB4 / 220 / 11
Qwen3.5-4B, BF169.3 GB0.30 GiB9 / 233 / 9
Qwen3.5-9B, BF1619.3 GB0.30 GiBdoes not fit2 / none
Ministral 3 3B, FP84.7 GB0.81 GiB8 / 217 / 4
Ministral 3 8B, FP810.4 GB1.06 GiB1 / none8 / 2
Ministral 3 14B, FP815.7 GB1.25 GiBdoes not fit3 / none
Llama 3.1 8B, NVIDIA FP89.1 GB1.00 GiB2 / none10 / 2
Nemotron 3 Nano 4B, FP85.3 GB0.14 GiB45 / 1997 / 41

Our estimates, not measurements: concurrent conversations at a full 8,192 or 32,768 tokens, from 90 per cent of nominal memory less 3 GiB. 16-bit KV cache, except Nemotron 3 Nano 4B in FP8 as NVIDIA’s vLLM command sets it. Weights from the Hugging Face file lists and cache from each config.json, read on 10 October 2026. 16 GB is the RTX PRO 2000; 24 GB is the L4, RTX PRO 4000 SFF or RTX PRO 4000.

The Gemma figures at 32K are those of our Gemma 4 hardware guide, which explains the sliding-window layers and why its 12B figure is an upper bound.

The other hybrid models keep little cache as well. Qwen3.5-4B and Qwen3.5-9B each have 32 layers, of which 8 use full attention with four KV heads of dimension 256, 32 KiB per token in 16-bit; the 24 Gated DeltaNet layers hold a fixed state of about 48 MiB in float32. Nemotron 3 Nano 4B keeps a cache in 4 of its 42 layers, and our Nemotron hardware guide gives its state. The dense models keep a cache in every layer: 128 KiB per token on Llama 3.1 8B (32 layers, eight KV heads) and 136 KiB on Ministral 3 8B (34 layers), four times the Qwen3.5 figure.

Qwen3.5-9B is a case where the weights leave too little room. Its BF16 checkpoint of 19.3 GB leaves 0.6 GiB of the 18.6 GiB budget on a 24 GB card, and its card links to 558 quantised versions, which we did not evaluate. Mistral publishes the Ministral 3 instruct models in FP8, and its cards say the 8B is “capable of fitting in 12GB of VRAM in FP8” and the 3B “in 8GB”. Our arithmetic agrees: the 8B with one conversation of 8,192 tokens comes to about 10.8 GiB.

Gemma 4, Qwen3.5-4B and 9B and Ministral 3 are published under Apache 2.0. Llama 3.1 8B comes under the Llama 3.1 Community License and Nemotron 3 Nano 4B under the NVIDIA Nemotron Open Model License. Whether a clause applies to your company is a legal assessment for your legal department.

We supply the L4, RTX PRO 2000, RTX PRO 4000 and RTX PRO 4000 SFF as cards or in servers built to order. Tell us the model, its context length and the users per site at peak, and we name the card that holds them.

Context length and thinking mode decide the user count

A serving engine reserves cache for the context you declare. vLLM takes it from the model configuration unless --max-model-len sets it, and Qwen3.5 and Ministral 3 declare 262,144 tokens. At that length one Ministral 3 8B conversation needs 34 GiB of 16-bit cache by our arithmetic, more than any card in this article holds. For classification or extraction of short texts, declare 8K or less; for a branch assistant that answers from a few retrieved passages, 8K to 32K.

Qwen3.5 models “operate in thinking mode by default”, and Qwen’s card advises “maintaining a context length of at least 128K tokens” to keep that mode working well. At 128K, Qwen3.5-4B in BF16 holds two conversations on a 24 GB card by our estimate. For extraction and routing, thinking can be switched off per request with enable_thinking set to false, which keeps outputs short and the context small. The same card documents --language-model-only, which skips the vision encoder so that its memory goes to the KV cache.

An FP8 KV cache, set in vLLM with --kv-cache-dtype fp8 as in NVIDIA’s command for Nemotron 3 Nano 4B, halves the attention cache per token and nearly doubles the counts of the dense models in the table. The fixed state of the hybrid layers does not shrink. Check engine support on your card first.

FP8 and 4-bit formats on Ada and Blackwell

vLLM’s quantisation documentation, dated 14 September 2026 in its footer, marks llm-compressor FP8 (W8A8) as supported on Ada and Hopper GPUs, and AWQ, GPTQ and the Marlin kernels for GPTQ, AWQ, FP8 and FP4 weights on Ada as well. The L4 is an Ada Lovelace card with no FP4 arithmetic of its own. The FP8 checkpoints of Ministral 3 and Llama 3.1 8B and Google’s QAT W4A16 come in their publishers’ own formats, which the table does not name, and it has no Blackwell column. Confirm each checkpoint with the vLLM release you run on your card.

None of the checkpoints in our table is in NVFP4, NVIDIA’s 4-bit format for Blackwell; our Gemma and Nemotron guides list NVFP4 versions of the larger models in those families.

Generation speed on 288 to 672 GB/s

A dense model reads all its weights from memory for every token it generates, so bandwidth divided by the size of the weights sets a ceiling for the speed one user sees. On that basis the RTX PRO 4000 has a ceiling about 2.2 times that of the L4 for the same model, the RTX PRO 4000 SFF about 1.4 times, and the RTX PRO 2000 about the same as the L4. These are ratios from the specifications, not measured speeds. Several users share each read of the weights, so total throughput rises with concurrent requests until the cache or the compute runs out.

Where a small model is enough, and where it is not

Small models fit work with short inputs and constrained outputs: classifying tickets or e-mails, extracting fields from forms into JSON, routing a request to the right queue or to a larger model, and answering routine questions at a branch from a limited set of documents through RAG. Mistral’s card for Ministral 3 3B calls it “Ideal for lightweight, real-time applications on edge or low-resource devices”. NVIDIA calls Nemotron 3 Nano 4B “an edge-ready small language model”. A position paper on arXiv (2506.02153, revised on 22 September 2026), whose abstract links to an NVIDIA Research page, argues that small models are “sufficiently powerful” for many agentic invocations and that systems with general conversational needs combine several models.

The limits follow from the memory arithmetic. Long contracts, reports read in full or many users with long conversations need more cache than 16 or 24 GB holds beside the weights, and broad knowledge and multi-step reasoning across documents point to a larger model. One layout keeps the small model at the branch for routing and extraction and sends the rest to a larger model in the data centre, on an RTX PRO 6000 or an H200 NVL.

Branch and edge servers: low-profile, SFF and short-depth

For a server at a branch or on a factory floor, the L4 is the one of the four designed for server airflow. Lenovo’s product guide for the ThinkEdge SE455 V3, updated on 25 August 2026, describes a “2U rack server, short depth (438mm depth, from EIA front rack flange)” and lists up to six L4 cards in it. Our comparison of L4 and L40S for inference lists data-centre servers that take up to eight L4 in 2U. Of the three 70 W cards, only the L4 supports vGPU, which lets one branch host give GPU slices to several virtual machines.

Where there is no server room, an RTX PRO 2000 or RTX PRO 4000 SFF in a compact desktop runs a small model from the PCIe slot, provided the maker lists that card, at 70 W and 2.7 inches high, for the chassis. Four L4 cards draw 288 W at their rated power, less than one RTX PRO 6000 Server Edition at 600 W, which matters where a branch rack has a single small feed.

We build GPU workstations and short-depth servers for branch sites to order, and we check the rack, power and airflow before we quote. Describe the branch site, its rack depth or desk space and the power available in the form below.

What we supply

We supply the NVIDIA L4, the RTX PRO 2000, the RTX PRO 4000 and the RTX PRO 4000 SFF as cards, in configured workstations or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. We size the card from the model, its precision, the context length and the users per site, and we confirm that the chassis takes the card in that slot and can cool it. vGPU licences for L4 hosts come on the same invoice as the hardware. Our professional GPU range lists every card, and the models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What GPU do I need for an 8B model?
An 8B model in FP8 is about 9 to 10.4 GB, so a 24 GB card such as the L4, RTX PRO 4000 SFF or RTX PRO 4000 holds Llama 3.1 8B with ten conversations of 8,192 tokens and Ministral 3 8B with eight, by our estimate with a 16-bit KV cache. A 16 GB RTX PRO 2000 holds either model for one or two such conversations. An FP8 cache roughly doubles these counts where the engine supports it on your card.
Can an LLM run on an NVIDIA L4?
Yes. The L4 has 24 GB at 300 GB/s and runs models up to about 14B parameters in FP8, which vLLM’s quantisation table marks as supported on its Ada Lovelace architecture. By our estimate it holds Qwen3.5-4B in BF16 with 33 conversations of 8,192 tokens and Ministral 3 14B in FP8 with three, while the card has no FP4 arithmetic.
Can the RTX PRO 2000 run an LLM?
It can run small models in its 16 GB. By our estimate it holds Gemma 4 E2B, Qwen3.5-4B, Ministral 3 3B and Nemotron 3 Nano 4B with several conversations, and Gemma 4 12B in Google’s 4-bit QAT version with four conversations of 8,192 tokens. Gemma 4 E4B in BF16, Qwen3.5-9B in BF16 and Ministral 3 14B do not fit.
What is a small language model?
The term is used for models of roughly 1 to 14 billion parameters that run on one low-power GPU or an edge device. Current examples are Gemma 4 E2B, E4B and 12B, Qwen3.5-4B and 9B, Ministral 3 in 3B, 8B and 14B, Llama 3.1 8B and Nemotron 3 Nano 4B, which NVIDIA describes as “an edge-ready small language model”. They suit classification, extraction, routing and branch assistants over a limited set of documents.
How much VRAM does a small language model need?
The weights take 4.7 to 19.3 GB for the models in this article, depending on size and format, and each conversation adds KV cache on top. Dense 8B models need about 1 GiB of 16-bit cache per 8,192 tokens, while hybrid models such as Qwen3.5-4B and Nemotron 3 Nano 4B need 0.30 and 0.14 GiB. Plan on 90 per cent of the card’s memory, less 3 GiB, for weights plus cache.
Which GPU is suitable for an edge LLM server?
In a server, the L4 is a passive 72 W single-slot low-profile card that fits short-depth edge servers; Lenovo lists up to six in its 2U ThinkEdge SE455 V3, which is 438 mm deep. Without server airflow, the actively cooled RTX PRO 2000 or RTX PRO 4000 SFF at 70 W fits a compact desktop whose maker lists that card for the slot. Of these cards, only the L4 supports vGPU.

Send us the model, its precision, the context length, the number of sites and the users per site at peak, and whether the card goes into a server or a compact desktop. We reply within one business day with the card that fits the slot, the cooling and the memory you need, and a written quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna