GLM hardware requirements: GLM-5.3, GLM-5.3-Flash and GLM-4.x on H200 NVL and RTX PRO 6000
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- GLM-5.3 (753B parameters) is a 756 GB FP8 checkpoint that needs eight H200 NVL in one server; its FP8 weights take 704 GiB, more than the 664 GiB that eight RTX PRO 6000 provide by our sizing rule, and the 1.51 TB BF16 version needs several servers
- By our estimate for vLLM, eight H200 NVL in two pipeline stages of four leave room for about 23 conversations of 32K tokens with a 16-bit cache or 37 with an FP8 cache, against 11 or 18 with tensor parallelism over all eight; 100 users at 32K need about three two-stage servers
- GLM-5.3-Flash (320B, 18B active, 328 GB in FP8) runs on four H200 NVL for 1 to 20 users or four RTX PRO 6000 for one user; GLM-4.7 and GLM-4.5 (FP8, 362 and 361 GB) need four H200 NVL or eight RTX PRO 6000
- GLM-5.3, GLM-5.3-Flash and GLM-4.7-Flash keep a compressed latent cache that vLLM stores in full on every card of a tensor-parallel group, so the memory left on one card, not on all cards, sets the number of conversations
- GLM-4.7-Flash (30B, 3B active, 62.5 GB in BF16) runs on one RTX PRO 6000, one H200 NVL or one DGX Spark; GLM-5.3 has its own GLM-5.3 License with a security review clause for large Model-as-a-Service operators, the others are MIT
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
GLM hardware requirements: the short answer
GLM-5.3, Z.ai’s 753B-parameter open-weight model, needs eight H200 NVL in one server for its 756 GB FP8 checkpoint. Run as two pipeline stages of four cards, that server holds about 23 concurrent conversations of 32,768 tokens with a 16-bit cache by our estimate, or 37 with the FP8 cache of vLLM’s recipe. With the recipe’s tensor parallelism over all eight cards, the figures are about 11 and 18. The FP8 weights do not fit eight RTX PRO 6000 or a DGX Spark. The smaller GLM-5.3-Flash runs on four H200 NVL, or on four RTX PRO 6000 for a single user, and GLM-4.7-Flash, a 30B model, runs on one RTX PRO 6000, one H200 NVL or one DGX Spark.
The configurations are our estimates for 1, 20 and 100 concurrent conversations at 32,768 tokens each on vLLM 0.28.0, the minimum version of vLLM’s GLM-5.3 recipe, made with the same memory rule as our LLM hardware requirements by model hub.
GLM-5.3, GLM-5.3-Flash and GLM-4.x variants in October 2026
| MODEL | PARAMETERS | CHECKPOINT | CONTEXT | CACHE PER 32K |
|---|---|---|---|---|
| GLM-5.3 | 753B, active not stated | FP8 756 GB; BF16 1.51 TB | 1,048,576 | 3.06 GiB; FP8 1.88 GiB |
| GLM-5. | 320B, 18B active | FP8 328 GB; BF16 643 GB | 1,048,576 | about 0.69 GiB |
| GLM-4.7 | 358B, active not stated | FP8 362 GB | 202,752 | 11.5 GiB |
| GLM-4.5 | 355B, 32B active | FP8 361 GB | 131,072 | 11.5 GiB |
| GLM-4.5-Air | 106B, 12B active | FP8 113 GB | 131,072 | 5.75 GiB |
| GLM-4. | 30B, 3B active | BF16 62.5 GB | 202,752 | 1.65 GiB |
zai-org model cards, file lists and config.json files and docs.z.ai, read on 9 October 2026. Cache per 32K is our estimate for one conversation, 16-bit unless FP8 is stated.
GLM-5.3 is a mixture-of-experts model whose config.json names the architecture GlmMoeDsaForCausalLM: 78 layers, 256 routed experts of which 8 are active per token, one shared expert and one multi-token prediction (MTP) layer. The FP8 checkpoint uses 128 × 128 block scales. Z.ai’s documentation gives text-only input, a 1M-token context window and a maximum output of 128K tokens, but neither source states the active parameters.
For GLM-5.3-Flash, Z.ai’s documentation states “320B total parameters with 18B activated” and native input of images, video and files, and its config.json has 45 layers, 34 with linear attention and 11 with sparse attention. GLM-4.7 and GLM-4.5 have 92 layers with grouped-query attention and 8 KV heads, GLM-4.5-Air the same attention in 46 layers. GLM-4.7-Flash, which Z.ai’s card calls “a 30B-A3B MoE model”, has BF16 weights only on its card.
KV cache per conversation from config.json
Our guide to how much VRAM an LLM needs has the general method: weights and cache share 90 per cent of the memory the driver reports, minus 3 GiB per card, which is 83.04 GiB per RTX PRO 6000 and 123.36 GiB per H200 NVL. For one DGX Spark (128 GB) we take 102 GB.
GLM-4.7, GLM-4.5 and GLM-4.5-Air store keys and values per KV head: 2 × 92 layers × 8 heads × 128 values × 2 bytes is 368 KiB per token, or 11.5 GiB for a 32K conversation in 16-bit. GLM-5.3 and GLM-4.7-Flash store one compressed latent vector per layer instead, 512 values plus a 64-value positional part, the kv_lora_rank and qk_rope_head_dim of their configuration files. For GLM-5.3 that is 1,152 bytes per layer in 16-bit, plus 132 bytes for the key of the sparse-attention indexer, which we count on all 78 layers as an upper bound. The result is 97.8 KiB per token and 3.06 GiB per 32K conversation.
With an FP8 cache, vLLM stores each token’s latent in 656 bytes per layer, as its post of 29 September 2025 describes for DeepSeek-V3.2, whose latent has the same 512 plus 64 values. That gives 60 KiB per token and 1.88 GiB per 32K conversation. For GLM-5.3-Flash we apply Z.ai’s own figure, a KV cache reduced “by 4.44× versus GLM-5.3”, to our GLM-5.3 number, which gives about 0.69 GiB.
vLLM’s blog on decode context parallelism of 7 August 2026 states for MLA models that under tensor parallelism “the latent KV cache is replicated in full across every TP rank”. GLM-5.3 keeps the same 512 plus 64 latent, so eight cards in one tensor-parallel group hold eight copies of each conversation’s cache. Pipeline parallelism halves the cache per card, because each stage keeps only the cache of its own layers, and vLLM lists pipeline parallelism as supported for the GLM-5 architecture. vLLM’s data parallel guide names a third layout for MoE models with MLA, data parallel attention, in which “Each DP engine has an independent KV cache” but every card repeats the non-expert weights; we have not sized it for GLM-5.3. The blog presents decode context parallelism (DCP) as the remedy, which splits each cache across the cards by token, and vLLM 0.30.0 lists “PCP+DCP on sparse-MLA models”, the attention type of GLM-5.3. By our estimate -dcp 8 would give about 92 conversations of 32K in 16-bit, but no recipe we found runs GLM-5.3 that way, so the table does not count on it.
GPUs for 1, 20 and 100 users
| MODEL, FORMAT | ONE DGX SPARK | RTX PRO 6000 | H200 NVL |
|---|---|---|---|
| GLM-5.3, FP8 | no / no / no | over 8 | 8 / 8 / over 8 |
| GLM-5.3, BF16 | no / no / no | over 8 | over 8 |
| GLM-5. | no / no / no | 4 / 8 / 8 | 4 / 4 / 8 |
| GLM-4.7 or GLM-4.5, FP8 | no / no / no | 8 / 8 / over 8 | 4 / 8 / over 8 |
| GLM-4.5-Air, FP8 | no / no / no | 2 / 4 / over 8 | 1 / 2 / 8 |
| GLM-4. | yes / yes / no | 1 / 2 / 8 | 1 / 1 / 4 |
Our estimates for vLLM 0.28.0, not measurements: cards for 1, 20 and 100 concurrent conversations of 32,768 tokens, in configurations of 1, 2, 4 or 8 cards, with a 16-bit cache. Latent caches (GLM-5.3, GLM-5.3-Flash, GLM-4.7-Flash) are counted in full on every card of a tensor-parallel group; eight-card rows of GLM-5.3 and GLM-5.3-Flash assume two pipeline stages of four. DGX Spark shows a memory fit only.
GLM-5.3 on eight H200 NVL leaves 35.35 GiB per card after the weights. Split into two pipeline stages of four cards, each card holds the cache of 39 layers, room for about 23 conversations of 32K in 16-bit or 37 with the FP8 cache that vLLM’s recipe sets. With the recipe’s tensor parallelism over all eight cards, every card holds all 78 layers of each cache, and the figures fall to about 11 and 18. A hundred users at 32K therefore need about three servers in the two-stage layout, while one such server holds 100 conversations of 8K with an FP8 cache.
GLM-5.3-Flash fits four cards of either type for its 305 GiB of weights. On four RTX PRO 6000, 6.7 GiB per card remains, about nine conversations of 32K, so 20 users need eight cards in two stages. Four H200 NVL leave 47 GiB per card, about 68 conversations, and 100 users need eight.
GLM-4.7 and GLM-4.5 in FP8 take about 337 GiB, more than the 332 GiB of four RTX PRO 6000, so they need eight. Four H200 NVL hold them with 13 conversations of 32K, and 20 users need eight. This matches Z.ai’s card for GLM-4.5, which gives four H200 at FP8 as the minimum and eight for the full 128K context. An FP8 cache halves the cache figures and lets eight H200 NVL serve 100 users of GLM-4.7.
GLM-4.7-Flash leaves room for about 15 conversations of 32K on one RTX PRO 6000 and 39 on one H200 NVL. One DGX Spark holds it with about 22. For 100 users, one copy per card on eight RTX PRO 6000 or four H200 NVL needs no traffic between cards.
On the RTX PRO 6000, TensorRT-LLM’s support matrix (version 1.3.0rc29) lists only per-tensor FP8, while these checkpoints use block or per-channel scales, so check that your engine runs them on that card.
We build AI servers with four or eight H200 NVL and their NVLink bridges, or with RTX PRO 6000 Server Edition cards. Send us the GLM variant, your context length and peak conversations, and we reply within one business day.
Why GLM-5.3 needs eight H200 NVL and not eight RTX PRO 6000
GLM-5.3’s FP8 weights take 704 GiB. Eight RTX PRO 6000 give 664 GiB by our rule, so the weights alone do not fit. Eight H200 NVL give 987 GiB, and vLLM’s GLM-5.3 recipe, updated on 3 October 2026, states that “The FP8 checkpoint fits on a single 8xH200 / 8xH20 node”. The BF16 repository, at 1.51 TB, exceeds the 1,128 GB of eight H200 NVL, and the recipe says these weights “need multi-node deployment”.
As of 9 October 2026, neither zai-org nor NVIDIA publishes a 4-bit checkpoint of GLM-5.3 on Hugging Face. vLLM’s recipe points to an NVFP4 checkpoint for Blackwell published by Inferact, and other NVFP4 conversions come from third parties; we do not size against them here. NVFP4 runs natively on Blackwell cards such as the RTX PRO 6000, while the H200 NVL has no FP4 arithmetic, as our guide to H200 NVL card counts for large models explains.
vLLM settings for GLM-5.3 on eight H200 NVL
vLLM’s GLM-5.3 recipe serves the FP8 checkpoint on eight H200 with --tensor-parallel-size 8 and --kv-cache-dtype fp8. It adds MTP speculative decoding with --speculative-config.method mtp and five speculative tokens. The recipe states that DeepGEMM is required.
In an eight-card H200 NVL server the bridges join at most four cards, so tensor parallelism over eight sends every layer’s all-reduce across PCIe between the two domains. Our article on splitting one model over PCIe and NVLink describes tensor parallelism inside each domain and pipeline parallelism between them, in vLLM --tensor-parallel-size 4 with --pipeline-parallel-size 2. This layout also halves the latent cache per card. The recipe does not test it, with or without MTP, so compare both layouts on the delivered server.
Set --max-model-len to the context you serve; the recipe’s AMD command uses 524,288 tokens. At the full 1M tokens, one conversation needs about 30 GiB of FP8 cache per card in the two-stage layout, which eight H200 NVL hold for one user by our estimate. Under plain tensor parallelism it needs 60 GiB per card, more than the 35 GiB left. vLLM’s blog of 8 September 2026 reports GLM-5.3 at its full 1M context on one node of eight H200 with Hybrid HiSparse, which moves the cache pages the indexer does not select to pinned CPU memory when GPU memory runs short. HiSparse came with release 0.30.0, 0.31.0 lists “HiSparse hardening”, and the post says it is “currently implemented only for NVIDIA GPUs”. For the full 1M context the recipe names eight B200 GPUs with 180 GB each, a system class we do not supply.
Rack, power and host memory for an eight-card H200 NVL server
Eight H200 NVL at up to 600 W each draw 4.8 kW for the cards alone, before processors, memory and fans. That is more than a single-phase 16 A circuit of about 3.7 kW carries, so plan three-phase feeds at the rack position and C19 outlets on the PDU for the larger power supplies. Each card needs the 600 W GPU power cable, since an H200 NVL on a 450 W cable does not boot. Z.ai’s GLM-4.5 card asks for more than 1 TB of server memory, and the eight-card H200 NVL configuration on our AI servers page has 1 to 2 TB. The model store needs 756 GB for the FP8 checkpoint, and another 1.51 TB if you keep the BF16 weights. Our GPU rack power and cooling article covers containment and floor loading.
We check the rack, power and airflow before we quote. Describe the rack position, its feed in amps and phases and the PDU outlets in the form below.
GLM licence terms for company use
GLM-5.3 comes under the GLM-5.3 License, which permits use of the software “without restriction” under its conditions. Its clause 2 defines Model as a Service as “giving a third party access to language model inference or fine-tuning (e.g., via API)” with meaningful control over inputs, parameters or training data. The definition excludes “end-user products with model capabilities solely embedded within specific features or harnesses”. If a licensee operating such a business exceeds the revenue threshold the clause sets, “the Licensee must pass Z.AI’s security review before using the Software or its derivative works for any commercial purpose.” We found no territorial restriction in the text. GLM-5.3-Flash, GLM-4.7, GLM-4.7-Flash and GLM-4.5 carry the MIT licence on their model cards. Whether clause 2 applies to a deployment is a legal assessment for your legal department.
What we supply
We supply the H200 NVL with its two-way and four-way NVLink bridges, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the DGX Spark Founders Edition, as cards or in AI servers built to order. All of it comes on one EU contract and invoice with manufacturer warranty; our professional GPU range lists every card. Models, RAG and MLOps on top of the hardware are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What are the hardware requirements for GLM-5.3?
How much VRAM does GLM-5 need per user?
Can I run GLM locally on one GPU?
How much VRAM does GLM-4.5 need?
Does GLM-5.3 run on RTX PRO 6000?
Can GLM-5.3 be used commercially?
Send us the GLM variant and precision, the context length, the peak number of concurrent conversations and the rack position’s power feed and PDU outlets. We reply within one business day with a configuration and a quote in writing, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day