gpt-oss-120b hardware requirements: VRAM and GPUs for 1 to 100 users, with gpt-oss-20b
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- gpt-oss-120b is a 65.3 GB checkpoint (60.8 GiB) with MXFP4 expert weights and BF16 for the rest; OpenAI says it runs on a single 80 GB GPU, and gpt-oss-20b, at 13.8 GB, within 16 GB of memory
- Only half the layers keep the full context, so a 16-bit KV cache costs 36 KiB per token on gpt-oss-120b and 24 KiB on gpt-oss-20b, or 1.125 GiB and 0.75 GiB per conversation of 32,768 tokens
- By our estimate, gpt-oss-120b serves 20 conversations at 32K on one H200 NVL or two RTX PRO 6000, and 100 on two H200 NVL or four RTX PRO 6000; one DGX Spark holds about 30 by memory
- At the full 131,072 tokens a gpt-oss-120b conversation takes 4.5 GiB of cache, and 100 of them need eight H200 NVL or eight RTX PRO 6000 as two copies of four, or four H200 NVL with an FP8 cache
- A higher reasoning level lengthens the chain of thought and the time each request stays in flight; the licence is Apache 2.0, with a usage policy that asks users to comply with all applicable law
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
gpt-oss-120b hardware requirements
The hardware requirements of gpt-oss-120b follow from about 61 GiB of GPU memory for its weights plus a KV cache for every conversation in flight, so one 96 GB RTX PRO 6000, one 141 GB H200 NVL or one DGX Spark with 128 GB of unified memory runs it for one user. For 20 concurrent conversations of 32,768 tokens it needs one H200 NVL or two RTX PRO 6000, and for 100 it needs two H200 NVL or four RTX PRO 6000, by our estimate. The smaller gpt-oss-20b takes 12.8 GiB and serves 100 such conversations on one H200 NVL or two RTX PRO 6000.
OpenAI’s model card says that MXFP4 quantisation of the MoE weights makes gpt-oss-120b “run on a single 80GB GPU” and gpt-oss-20b “run within 16GB of memory”. Our hardware requirements by model compare it with other open models.
gpt-oss variants and checkpoint sizes in October 2026
OpenAI has published two gpt-oss models on Hugging Face, and two safety reasoning models built on them. On 9 October 2026, OpenAI’s organisation page showed no newer gpt-oss text model.
| MODEL | PARAMETERS | WEIGHT FORMAT | CHECKPOINT | LICENCE |
|---|---|---|---|---|
| gpt-oss-120b | 117B, 5.1B active | MXFP4 experts, BF16 rest | 65.3 GB | Apache 2.0 |
| gpt-oss-20b | 21B, 3.6B active | MXFP4 experts, BF16 rest | 13.8 GB | Apache 2.0 |
| gpt-oss-safeguard-120b | 117B, 5.1B active | as gpt-oss-120b | 65.3 GB | Apache 2.0 |
| gpt-oss-safeguard-20b | 21B, 3.6B active | as gpt-oss-20b | 13.8 GB | Apache 2.0 |
| gpt-oss-puzzle-88B (NVIDIA) | about 88B, active not stated | MXFP4 experts, BF16 rest | 50.0 GB | NVIDIA Open Model License |
Model cards and file lists on Hugging Face, read on 9 October 2026; the checkpoint is the sum of the top-level safetensors files. OpenAI’s model card lists 60.8 GiB and 12.8 GiB for the two base checkpoints. Puzzle’s rest is BF16 by its listed tensor types.
Both base models are mixture-of-experts transformers. gpt-oss-120b has 36 layers with 128 experts, gpt-oss-20b has 24 layers with 32, and both route each token to 4 experts. Only the expert weights are in MXFP4, which the model card puts at 4.25 bits per parameter and at more than 90 per cent of all parameters. Attention, router, embeddings and output head stay in BF16, as the config.json files list. Both models accept 131,072 tokens of context, extended with YaRN from an original 4,096.
With its original folder of 65.2 GB and a metal folder, the gpt-oss-120b repository totals 196 GB and gpt-oss-20b 41.3 GB, which a full download needs on disk.
The safeguard models classify text against a safety policy you supply and size like their base models. NVIDIA’s gpt-oss-puzzle-88B, published on 26 March 2026, is derived from gpt-oss-120b, replaces a subset of global attention layers with an 8K window and keeps its cache in FP8. Its card lists only the B200 and the H100-80GB as supported hardware, so test it on your card first.
KV cache per conversation from config.json
The general method is in our guide to how much VRAM an LLM needs. Usable memory is 90 per cent of what the driver reports, less 3 GiB per card, and what remains after the weights holds the cache. That gives 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For one DGX Spark (128 GB) we take 102 GB, about 95 GiB, as in our hub.
gpt-oss keeps the full context in only half of its layers. The model card says the attention blocks “alternate between banded window and fully dense patterns, where the bandwidth is 128 tokens”, and vLLM reserves cache only for the window in such layers. Each full layer stores keys and values for 8 KV heads of dimension 64. One token therefore costs 2 × 18 × 8 × 64 × 2 bytes, or 36 KiB, on gpt-oss-120b in 16-bit, and 24 KiB over the 12 full layers of gpt-oss-20b. The windowed layers add about 4.5 MiB per conversation on the larger model at any length.
| CONTEXT | GPT-OSS-20B | GPT-OSS-120B | 120B, FP8 CACHE |
|---|---|---|---|
| 8,192 tokens | 0.19 GiB | 0.28 GiB | 0.14 GiB |
| 32,768 tokens | 0.75 GiB | 1.125 GiB | 0.56 GiB |
| 131,072 tokens | 3 GiB | 4.5 GiB | 2.25 GiB |
Our arithmetic from the config.json files: full-attention layers × 2 × 8 KV heads × 64 × bytes per value × tokens, per conversation at that length; the 128-token window layers are left out.
A conversation’s length counts the system prompt, the history and everything the model writes, its reasoning included, up to the context you declare.
Reasoning effort, harmony and the load per request
gpt-oss reasons before it answers, at three levels set in the system prompt with a line such as Reasoning: high. OpenAI’s harmony guide says the model uses medium by default. The model card states that “Increasing the reasoning level will cause the model’s average CoT length to increase.”
For sizing, a longer chain of thought means that each request writes more tokens, fills more cache and stays in flight longer, so the same number of employees produces more concurrent requests. OpenAI’s model card plots the average reasoning length per level for two benchmarks only, and reports that gpt-oss-20b used over 20,000 reasoning tokens per AIME problem on average. Measure the length on your own prompts during a pilot and add it to the context you declare.
Harmony is OpenAI’s response format for gpt-oss, with the reasoning in an analysis channel and the answer in a final channel. OpenAI’s harmony repository states that gpt-oss “should not be used without using the harmony format as it will not work correctly”, and vLLM, Ollama and Hugging Face apply it for you. OpenAI’s guide adds that you “should drop any previous CoT content on subsequent sampling” once a reply has ended in the final channel, except around tool and function calls, where the earlier reasoning goes back in. A chat history therefore grows by questions and answers, while an agent that calls tools carries its reasoning forward and reaches a long context sooner.
GPUs for gpt-oss: 1, 20 and 100 concurrent users
The table counts conversations in flight at the full stated length at the same moment, with a 16-bit cache. Two or four cards mean one copy split with tensor parallelism unless the text names copies. vLLM’s context parallel guide says the split shards the cache by KV head and duplicates it only when the tensor parallel size exceeds the number of KV heads; gpt-oss has eight, so the cache of up to eight cards adds up.
| MODEL, CONTEXT | ONE DGX SPARK | RTX PRO 6000 | H200 NVL |
|---|---|---|---|
| gpt-oss-120b, 32K | yes / yes / no | 1 / 2 / 4 | 1 / 1 / 2 |
| gpt-oss-120b, 131K | yes / no / no | 1 / 2 / 8 | 1 / 2 / 8 |
| gpt-oss-20b, 32K | yes / yes / yes | 1 / 1 / 2 | 1 / 1 / 1 |
| gpt-oss-20b, 131K | yes / yes / no | 1 / 1 / 4 | 1 / 1 / 4 |
Our estimates, not measurements: cards needed, in configurations of 1, 2, 4 or 8, for 1 / 20 / 100 concurrent conversations of 32,768 or 131,072 tokens with a 16-bit KV cache; eight cards run as two copies of four. DGX Spark shows a memory fit for one Spark, not its speed.
On one RTX PRO 6000, gpt-oss-120b leaves 22.2 GiB for the cache, room for 19 conversations at 32K, so 20 users need a second card, and two cards as replicas hold 38. For 100 users at 32K, four cards hold 241 conversations as one split copy, 186 as two copies of two cards and 76 as four single-card copies. One H200 NVL leaves 62.5 GiB, enough for 55 conversations at 32K, and two as replicas hold 110 without a bridge.
At the full 131,072 tokens the counts fall by a factor of four. One RTX PRO 6000 then holds four conversations, one H200 NVL 13 and one Spark 7. Four H200 NVL on a four-way bridge hold 96, so the table shows eight, run as two copies of four, one per NVLink domain, which hold 192. Eight RTX PRO 6000 in that layout hold 120, and four H200 NVL with an FP8 cache hold 192.
gpt-oss-20b leaves 70.2 GiB on one RTX PRO 6000, room for 93 conversations at 32K, and one H200 NVL holds 147. One Spark holds 109 by memory, although its 273 GB/s of memory bandwidth limits how many of them stream at a usable speed; the measured figures are in our DGX Spark benchmarks.
An FP8 cache roughly doubles every count, to 111 conversations of gpt-oss-120b at 32K on one H200 NVL. TensorRT-LLM’s support matrix lists an FP8 KV cache for Hopper and for the RTX PRO Blackwell generation (sm120). Among NVIDIA GPUs, vLLM’s gpt-oss recipe sets it only for the B200, so test answer quality first.
Size for the peak number of requests in flight, which differs from headcount; our article on how many users one RTX PRO 6000 serves shows how to take that peak from gateway logs.
We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the gpt-oss variant, your context length and peak requests in flight, and we reply within one business day with a configuration and quote.
H200 NVL, RTX PRO 6000 or DGX Spark for gpt-oss
TensorRT-LLM’s support matrix, version 1.3.0rc29 of 26 September 2026, lists MXFP4 for Blackwell, including the RTX PRO generation, and not for Hopper. The RTX PRO 6000 and the DGX Spark have FP4 arithmetic, which an engine uses only if it has an MXFP4 kernel for them; vLLM’s B200 path pairs MXFP4 weights with MXFP8 activations. The H200 NVL holds the weights in MXFP4 too, but without FP4 arithmetic it computes those layers at higher precision, so memory per card is the same and the compute path differs.
vLLM’s gpt-oss recipe, updated 22 September 2026, names the H100, H200 and B200 among NVIDIA GPUs and mentions “ongoing work for Ampere/Ada/RTX 5090”. It does not name the RTX PRO 6000 or the DGX Spark. StorageReview served an NVFP4 version of gpt-oss-120b with vLLM on four RTX PRO 6000 Server Edition cards, as our article on one model on several GPUs over PCIe and NVLink reports. Check the engine version on the card before production.
Because gpt-oss fits one card, replicas behind a load balancer are the default; they exchange nothing over PCIe and fail independently. A split pays where the cache runs short, as for 100 conversations at 32K on RTX PRO 6000 cards. That card has no NVLink, so the split runs over PCIe, where the peer-to-peer settings decide the speed, while four H200 NVL on a four-way bridge keep it on NVLink at 900 GB/s per GPU.
vLLM settings for gpt-oss-120b
vLLM’s recipe serves gpt-oss-120b on a B200 with --tensor-parallel-size 1 and a configuration file, and describes Hopper as the same “without kv-cache-dtype and without the FlashInfer MoE flags”. vLLM’s earlier gpt-oss guide, kept in its documentation as a historical reference, lists both files: the Hopper file sets no-enable-prefix-caching: true and max-num-batched-tokens: 8192, and the Blackwell file adds kv-cache-dtype: fp8. The guide turns off prefix caching “if running with synthetic dataset”, a benchmark setting, so test prefix caching for a chat service.
The guide names two defaults that matter for memory. The context length defaults to “the maximum sequence length supported by the model”, 131,072 tokens. The number of sequences defaults to “a large number like 1024 on GPUs with large memory sizes”. Set --max-model-len to the context your users need, and vLLM reports a higher maximum concurrency for the same cache.
Operating system, drivers, CUDA and a container runtime are installed on request, and deploying the model with vLLM is part of our Private AI/ML service. Write to us with the variant, the reasoning level you plan and your rack position.
Licence and usage policy for gpt-oss
gpt-oss-120b, gpt-oss-20b and both safeguard models are published under the Apache 2.0 licence, as their model cards state. Each OpenAI repository also carries a short USAGE_POLICY file, which for gpt-oss-120b reads: “By using OpenAI gpt-oss-120b, you agree to comply with all applicable law.” NVIDIA’s gpt-oss-puzzle-88B comes under the NVIDIA Open Model License instead. Whether these terms fit your intended use is a legal assessment for your legal department.
What we supply
We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its two-way and four-way NVLink bridges, as cards or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. For gpt-oss we size the server from the variant, the context and the peak requests in flight, and we check the rack, power and airflow before we quote. The configuration and quote follow within one business day, and our professional GPU range lists every card. Running the model, RAG and MLOps on top is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How much VRAM does gpt-oss-120b need?
How much VRAM does gpt-oss-20b need?
Does gpt-oss-120b run on one H200 NVL?
What GPU server does gpt-oss-120b need for 100 users?
Can I run gpt-oss-120b locally on a DGX Spark?
Can gpt-oss be used commercially on premise?
Send us the gpt-oss variant, the reasoning level you plan, the context you will declare and your peak requests in flight. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day