Gemma hardware requirements: VRAM and GPUs for Gemma 4 from E2B to 31B on-premise
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Every Gemma 4 size runs for one user on one GPU we supply: E2B and E4B on a 24 GB RTX PRO 4000 or L4 in BF16, the 12B on 24 GB in Google’s 4-bit QAT version, and the 26B A4B and 31B on one RTX PRO 6000, one H200 NVL or one DGX Spark in BF16
- The 31B keeps a 16-bit KV cache of up to about 3.3 GiB per conversation of 32,768 tokens by our count, four times that of the 12B or 26B A4B, because its 10 full-attention layers have four KV heads of dimension 512
- Fifty of the 31B’s 60 layers attend over a 1,024-token sliding window, so they hold a fixed 0.78 GiB per conversation, and a 256K conversation takes about 20.8 GiB in all
- For 20 conversations at 32K one RTX PRO 6000 serves every size except the 31B, which needs two; 100 conversations of the 31B need four H200 NVL, or six to eight RTX PRO 6000 depending on the format, by our estimate
- Gemma 4 is published under Apache 2.0, while Gemma 3, Gemma 3n and the other models in the appendix of the Gemma Terms of Use stay under those terms and their Prohibited Use Policy
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Gemma hardware requirements in brief
Gemma 4, Google’s current open model family as of October 2026, comes in five sizes, and each of them runs for one user on one GPU we supply. Gemma 4 E2B and E4B fit a 24 GB RTX PRO 4000 or L4 in BF16. The 12B needs a 32 GB card in BF16, or 24 GB in Google’s 4-bit QAT version. The 26B A4B and the 31B fit one RTX PRO 6000, one H200 NVL or one DGX Spark in BF16.
For 20 concurrent conversations of 32,768 tokens, one RTX PRO 6000 serves every size except the 31B, which needs a second card. For 100 conversations the 31B needs four H200 NVL, or six to eight RTX PRO 6000 depending on the format. Our estimates use Google’s Hugging Face files as read on 9 October 2026 and the rule of our guide to how much VRAM an LLM needs: 90 per cent of the memory the driver reports, less 3 GiB per card. Our LLM hardware requirements by model compares other model families on the same basis.
Gemma 4 sizes, checkpoints and context in October 2026
| MODEL | PARAMETERS | GOOGLE CHECKPOINTS | CONTEXT | INPUTS |
|---|---|---|---|---|
| Gemma 4 E2B | 5.1B with embeddings, 2.3B effective | BF16, 10.2 GB | 128K | text, image, audio, video |
| Gemma 4 E4B | 8B with embeddings, 4.5B effective | BF16, 16 GB | 128K | text, image, audio, video |
| Gemma 4 12B | 11.95B, dense | BF16 23.9 GB; QAT W4A16 10.3 GB | 256K | text, image, audio, video |
| Gemma 4 26B A4B | 25.2B MoE, 3.8B active | BF16, 51.6 GB | 256K | text, image, video |
| Gemma 4 31B | 30.7B, dense | BF16 62.6 GB; QAT W4A16 23.3 GB | 256K | text, image, video |
Model card of Gemma 4 31B, file lists and config.json files on huggingface.co/google, read on 9 October 2026. Video is read as frames.
Google’s release notes list Gemma 4 in the E2B, E4B, 31B and 26B A4B sizes on 31 March 2026, and the 12B Unified on 3 June 2026. The “E” stands for effective parameters; the checkpoints also hold the Per-Layer Embeddings. The 26B A4B has 128 experts, eight of them active per token, plus one shared expert.
Google publishes quantisation-aware trained (QAT) versions as W4A16 checkpoints, 4-bit weights with 16-bit activations, which the cards describe as made “for native, optimized inference with vLLM”, and as Q4_0 GGUF files. FP8 versions come from Red Hat, and NVFP4 versions from NVIDIA and Red Hat. Red Hat’s FP8 checkpoints are 33.3 GB for the 31B and 28.6 GB for the 26B A4B, and NVIDIA’s NVFP4 checkpoints 32.6 GB and 18.8 GB. The RTX PRO cards and the DGX Spark compute FP4 directly; the H200 NVL, L4 and L40S do not.
Google’s own memory table, last updated 8 July 2026, puts the 31B at 69.9 GB in BF16 for loading, with 20 per cent overhead, and notes that the figures “don’t include the additional VRAM needed for supporting software or the context window.”
KV cache per conversation with sliding-window layers
The cache per token comes from config.json: 2 × layers × KV heads × head dimension × bytes per value, counted only for layers that keep the whole context. The model card says Gemma 4 “interleaves local sliding window attention with full global attention, ensuring the final layer is always global.” Sliding-window layers keep only the last 1,024 tokens (512 on E2B and E4B), a fixed amount per conversation.
The 31B lists 60 layers, of which 10 are full attention with four KV heads of dimension 512. That gives 80 KiB per token in 16-bit. Its 50 sliding layers have 16 KV heads of dimension 256 and hold 0.78 GiB per conversation. A conversation of 32K therefore takes about 3.3 GiB, and one of 256K about 20.8 GiB.
The 12B has 8 full layers with one KV head, and the 26B A4B has 5 with two. Both come to about 0.8 GiB per 32K conversation, a quarter of the 31B’s. On E2B and E4B the configuration marks 20 of 35 and 18 of 42 layers as KV-shared, and LMCache’s Gemma 4 recipe states that it “stores the cache-owning layers only”. We count the 15 and 24 layers that own a cache, three and four of them with full attention.
The 12B, 26B A4B and 31B also set attention_k_eq_v to true; the Transformers documentation says this setting means “the key projection output is reused as the value projection”. We found no vLLM document stating that the engine then stores one tensor instead of two, so we count both, and our cache figures for these three models are upper bounds. With one tensor, a 32K conversation of the 31B takes about 2.0 GiB.
Which card holds which Gemma size
| MODEL, FORMAT | CACHE PER 32K | 24 GB CARD | 32 GB CARD | 48 GB CARD |
|---|---|---|---|---|
| E2B, BF16 | 0.19 GiB | 47 | 84 | 158 |
| E4B, BF16 | 0.52 GiB | 7 | 20 | 48 |
| 12B, BF16 | 0.81 GiB | does not fit | 4 | 22 |
| 12B, QAT W4A16 | 0.81 GiB | 11 | 19 | 37 |
| 26B A4B, FP8 | 0.82 GiB | does not fit | does not fit | 16 |
| 26B A4B, NVFP4 | 0.82 GiB | 1 | 10 | 27 |
| 31B, QAT W4A16 | 3.28 GiB | does not fit | 1 | 5 |
| 31B, FP8 | 3.28 GiB | does not fit | does not fit | 2 |
Our estimates, not measurements: conversations of 32,768 tokens that fit on one card with a 16-bit cache, keys and values counted in every layer, from the nominal memory of these cards, so the counts are upper bounds. 24 GB is the RTX PRO 4000 or L4, 32 GB the RTX PRO 4500, 48 GB the RTX PRO 5000 48 GB or L40S; the NVFP4 row applies to the RTX PRO cards only. Weights from the Hugging Face file lists, 9 October 2026.
A 24 GB card leaves 18.6 GiB for weights and cache by our rule. That holds E2B and E4B in BF16, and the 12B in its QAT version with 11 conversations at 32K. The 72 GB RTX PRO 5000 holds the 31B in FP8 with nine conversations and the 26B A4B in FP8 with 42. Our comparison of RTX PRO 6000, 5000 and 4500 covers these cards beyond memory.
The L4 and L40S are Ada Lovelace cards with FP8 but no FP4. vLLM’s Gemma 4 recipes name Hopper and Blackwell for the FP8 checkpoints, so check FP8 on an Ada card with your engine version first; our comparison of L4 and L40S for inference covers both cards.
We supply every card in this table, on its own or in servers built to order. Tell us which Gemma size you plan to run, its format and the users at peak, and we reply within one business day with a configuration and quote.
GPUs for 1, 20 and 100 users of each Gemma size
| MODEL, FORMAT | ONE DGX SPARK | RTX PRO 6000 | H200 NVL |
|---|---|---|---|
| E4B, BF16 | yes / yes / yes | 1 / 1 / 1 | 1 / 1 / 1 |
| 12B, BF16 | yes / yes / no | 1 / 1 / 2 | 1 / 1 / 1 |
| 26B A4B, BF16 | yes / yes / no | 1 / 1 / 2 | 1 / 1 / 2 |
| 26B A4B, FP8 | yes / yes / no | 1 / 1 / 2 | 1 / 1 / 1 |
| 31B, BF16 | yes / no / no | 1 / 2 / 8 | 1 / 2 / 4 |
| 31B, FP8 | yes / no / no | 1 / 2 / 6 | 1 / 1 / 4 |
| 31B, QAT W4A16 | yes / yes / no | 1 / 2 / 6 | 1 / 1 / 4 |
Our estimates, not measurements: cards needed for 1, 20 and 100 concurrent conversations of 32,768 tokens with a 16-bit cache, as identical copies of the model on one, two or four cards each; DGX Spark shows a memory fit only, from a 102 GB working set. E2B fits one card or one Spark for all three loads.
The figures in a cell differ where the cache sets the limit. The 31B in BF16 leaves 24.7 GiB on one RTX PRO 6000, room for seven conversations at 32K. In FP8 it leaves 52 GiB, room for 15, and one H200 NVL holds 28. One DGX Spark holds the 31B in FP8 with 19 conversations, one short of the 20 in the table, and its QAT version with 22. The model card says the 26B A4B “runs almost as fast as a 4B-parameter model”, but it needs memory for all 25.2B parameters.
A full 256K conversation of the 31B, about 20.8 GiB of cache, fits one RTX PRO 6000 or one DGX Spark in BF16, and one H200 NVL holds three. vLLM’s 31B recipe says an FP8 cache, --kv-cache-dtype fp8, “saves ~50% KV memory”.
Gemma 4 31B and 26B A4B on several GPUs
vLLM’s recipe for the 31B, updated 11 May 2026, runs it on two A100 or H100 with --tensor-parallel-size 2 and a context of 32,768 tokens, and the 26B A4B on one. The 26B A4B recipe, updated 12 September 2026, sets a memory utilisation of 0.8 on the DGX Spark, with its “128 GB shared with the CPU”, and 0.92 on the RTX PRO 6000. That matches our 102 GB for one Spark.
Tensor parallelism splits the KV heads across the cards. The 31B has four KV heads in its full layers, the 26B A4B two and the 12B one, so beyond that many cards the full-layer cache is duplicated instead of split further, and on the 12B every card keeps it whole. For 100 users of the 31B, identical copies on one, two or four cards each therefore make better use of memory than one copy across eight. The RTX PRO 6000 has no NVLink, so a two-card split runs over PCIe, as our guide to one model on several GPUs explains. For the 26B A4B the recipe states that “TEP (tensor-expert parallelism) and DEP (data-expert parallelism) strategies scale better than pure TP at large node counts.”
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the Gemma size, your context length and peak requests in flight through the form below, and we reply with a configuration and quote.
vLLM settings and versions for Gemma 4
vLLM’s recipe index lists vLLM 0.19.1 as the minimum for E2B, E4B and the 31B, 0.23.0 for the 12B and 0.25.0 for the 26B A4B. The 12B recipe, updated 4 June 2026, says the context is pinned to 131,072 tokens from config.json, while the config.json we read lists 262,144, so check which limit your vLLM version applies.
Google released draft models for multi-token prediction (MTP) on 16 April 2026, as “assistant” repositories of 78M to 0.5B parameters. vLLM’s 31B recipe of 11 May 2026 states that “MTP speculative decoding for Gemma 4 is only available on the vLLM nightly build”, and LMCache notes that “The draft layers carry their own KV cache”.
The 31B recipe marks NVIDIA’s NVFP4 version as requiring “Blackwell (B200/B300)”, while the 26B A4B recipe lists NVIDIA’s NVFP4 version as validated on single-GPU Blackwell systems, among them the DGX Spark and the RTX PRO 6000. For the 31B, Google’s QAT W4A16 checkpoint is smaller.
Gemma 3 27B and earlier Gemma models
Gemma 3 27B, the previous generation, is 54.8 GB in BF16 and fits one RTX PRO 6000, one H200 NVL, one DGX Spark or one 72 GB RTX PRO 5000, but not a 48 GB card. Google’s QAT version in Q4_0 GGUF is 17.2 GB plus an 858 MB vision projector. Its repository is gated, so we did not read its config.json and give no cache figure here.
Gemma licence: Apache 2.0 and the Gemma Terms of Use
The Gemma 4 model cards list Apache 2.0, and Google’s developer site links its Apache 2.0 page as the “Gemma 4 license”. The Gemma Terms of Use, last modified on 1 April 2026, state: “For Gemma 4 terms, see the Gemma 4 license.” They cover the models in their appendix, among them Gemma 3, Gemma 3n and TranslateGemma. Their Section 3.2 bars use of the Gemma Services for the restricted uses in the Gemma Prohibited Use Policy, and states that “Google reserves the right to restrict (remotely or otherwise) usage” it reasonably believes violates the agreement. Whether a clause applies to your company is a legal assessment for your legal department.
What we supply
We supply the RTX PRO 4000, 4500, 5000 and 6000, the L4 and L40S, the H200 NVL and the DGX Spark Founders Edition, as cards or in AI servers built to order. We size the card or server from the Gemma size, its format, your context and peak requests, and check the rack, power and airflow before we quote. Everything comes on one EU contract and invoice with manufacturer warranty, and our professional GPU range lists every card. The models, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What are the hardware requirements for Gemma 4?
How much VRAM does Gemma 4 31B need?
How much VRAM does Gemma 27B need?
Can I run Gemma locally on a single GPU?
Can Gemma 4 run on a DGX Spark?
Can Gemma be used commercially?
Send us the Gemma size and format, the context length you will configure, your peak number of concurrent requests and the rack position or workstation it goes into. We reply within one business day with a configuration and quote for that card or server, after checking the rack, power and airflow.
Talk to an expertWe reply within one business day