BLOG · GUIDE ·

gpt-oss-120b hardware requirements: VRAM and GPUs for 1 to 100 users, with gpt-oss-20b

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • gpt-oss-120b is a 65.3 GB checkpoint (60.8 GiB) with MXFP4 expert weights and BF16 for the rest; OpenAI says it runs on a single 80 GB GPU, and gpt-oss-20b, at 13.8 GB, within 16 GB of memory
  • Only half the layers keep the full context, so a 16-bit KV cache costs 36 KiB per token on gpt-oss-120b and 24 KiB on gpt-oss-20b, or 1.125 GiB and 0.75 GiB per conversation of 32,768 tokens
  • By our estimate, gpt-oss-120b serves 20 conversations at 32K on one H200 NVL or two RTX PRO 6000, and 100 on two H200 NVL or four RTX PRO 6000; one DGX Spark holds about 30 by memory
  • At the full 131,072 tokens a gpt-oss-120b conversation takes 4.5 GiB of cache, and 100 of them need eight H200 NVL or eight RTX PRO 6000 as two copies of four, or four H200 NVL with an FP8 cache
  • A higher reasoning level lengthens the chain of thought and the time each request stays in flight; the licence is Apache 2.0, with a usage policy that asks users to comply with all applicable law

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

gpt-oss-120b hardware requirements

The hardware requirements of gpt-oss-120b follow from about 61 GiB of GPU memory for its weights plus a KV cache for every conversation in flight, so one 96 GB RTX PRO 6000, one 141 GB H200 NVL or one DGX Spark with 128 GB of unified memory runs it for one user. For 20 concurrent conversations of 32,768 tokens it needs one H200 NVL or two RTX PRO 6000, and for 100 it needs two H200 NVL or four RTX PRO 6000, by our estimate. The smaller gpt-oss-20b takes 12.8 GiB and serves 100 such conversations on one H200 NVL or two RTX PRO 6000.

OpenAI’s model card says that MXFP4 quantisation of the MoE weights makes gpt-oss-120b “run on a single 80GB GPU” and gpt-oss-20b “run within 16GB of memory”. Our hardware requirements by model compare it with other open models.

gpt-oss variants and checkpoint sizes in October 2026

OpenAI has published two gpt-oss models on Hugging Face, and two safety reasoning models built on them. On 9 October 2026, OpenAI’s organisation page showed no newer gpt-oss text model.

MODELPARAMETERSWEIGHT FORMATCHECKPOINTLICENCE
gpt-oss-120b117B, 5.1B activeMXFP4 experts, BF16 rest65.3 GBApache 2.0
gpt-oss-20b21B, 3.6B activeMXFP4 experts, BF16 rest13.8 GBApache 2.0
gpt-oss-safeguard-120b117B, 5.1B activeas gpt-oss-120b65.3 GBApache 2.0
gpt-oss-safeguard-20b21B, 3.6B activeas gpt-oss-20b13.8 GBApache 2.0
gpt-oss-puzzle-88B (NVIDIA)about 88B, active not statedMXFP4 experts, BF16 rest50.0 GBNVIDIA Open Model License

Model cards and file lists on Hugging Face, read on 9 October 2026; the checkpoint is the sum of the top-level safetensors files. OpenAI’s model card lists 60.8 GiB and 12.8 GiB for the two base checkpoints. Puzzle’s rest is BF16 by its listed tensor types.

Both base models are mixture-of-experts transformers. gpt-oss-120b has 36 layers with 128 experts, gpt-oss-20b has 24 layers with 32, and both route each token to 4 experts. Only the expert weights are in MXFP4, which the model card puts at 4.25 bits per parameter and at more than 90 per cent of all parameters. Attention, router, embeddings and output head stay in BF16, as the config.json files list. Both models accept 131,072 tokens of context, extended with YaRN from an original 4,096.

With its original folder of 65.2 GB and a metal folder, the gpt-oss-120b repository totals 196 GB and gpt-oss-20b 41.3 GB, which a full download needs on disk.

The safeguard models classify text against a safety policy you supply and size like their base models. NVIDIA’s gpt-oss-puzzle-88B, published on 26 March 2026, is derived from gpt-oss-120b, replaces a subset of global attention layers with an 8K window and keeps its cache in FP8. Its card lists only the B200 and the H100-80GB as supported hardware, so test it on your card first.

KV cache per conversation from config.json

The general method is in our guide to how much VRAM an LLM needs. Usable memory is 90 per cent of what the driver reports, less 3 GiB per card, and what remains after the weights holds the cache. That gives 83.0 GiB per RTX PRO 6000 and 123.4 GiB per H200 NVL. For one DGX Spark (128 GB) we take 102 GB, about 95 GiB, as in our hub.

gpt-oss keeps the full context in only half of its layers. The model card says the attention blocks “alternate between banded window and fully dense patterns, where the bandwidth is 128 tokens”, and vLLM reserves cache only for the window in such layers. Each full layer stores keys and values for 8 KV heads of dimension 64. One token therefore costs 2 × 18 × 8 × 64 × 2 bytes, or 36 KiB, on gpt-oss-120b in 16-bit, and 24 KiB over the 12 full layers of gpt-oss-20b. The windowed layers add about 4.5 MiB per conversation on the larger model at any length.

CONTEXTGPT-OSS-20BGPT-OSS-120B120B, FP8 CACHE
8,192 tokens0.19 GiB0.28 GiB0.14 GiB
32,768 tokens0.75 GiB1.125 GiB0.56 GiB
131,072 tokens3 GiB4.5 GiB2.25 GiB

Our arithmetic from the config.json files: full-attention layers × 2 × 8 KV heads × 64 × bytes per value × tokens, per conversation at that length; the 128-token window layers are left out.

A conversation’s length counts the system prompt, the history and everything the model writes, its reasoning included, up to the context you declare.

Reasoning effort, harmony and the load per request

gpt-oss reasons before it answers, at three levels set in the system prompt with a line such as Reasoning: high. OpenAI’s harmony guide says the model uses medium by default. The model card states that “Increasing the reasoning level will cause the model’s average CoT length to increase.”

For sizing, a longer chain of thought means that each request writes more tokens, fills more cache and stays in flight longer, so the same number of employees produces more concurrent requests. OpenAI’s model card plots the average reasoning length per level for two benchmarks only, and reports that gpt-oss-20b used over 20,000 reasoning tokens per AIME problem on average. Measure the length on your own prompts during a pilot and add it to the context you declare.

Harmony is OpenAI’s response format for gpt-oss, with the reasoning in an analysis channel and the answer in a final channel. OpenAI’s harmony repository states that gpt-oss “should not be used without using the harmony format as it will not work correctly”, and vLLM, Ollama and Hugging Face apply it for you. OpenAI’s guide adds that you “should drop any previous CoT content on subsequent sampling” once a reply has ended in the final channel, except around tool and function calls, where the earlier reasoning goes back in. A chat history therefore grows by questions and answers, while an agent that calls tools carries its reasoning forward and reaches a long context sooner.

GPUs for gpt-oss: 1, 20 and 100 concurrent users

The table counts conversations in flight at the full stated length at the same moment, with a 16-bit cache. Two or four cards mean one copy split with tensor parallelism unless the text names copies. vLLM’s context parallel guide says the split shards the cache by KV head and duplicates it only when the tensor parallel size exceeds the number of KV heads; gpt-oss has eight, so the cache of up to eight cards adds up.

MODEL, CONTEXTONE DGX SPARKRTX PRO 6000H200 NVL
gpt-oss-120b, 32Kyes / yes / no1 / 2 / 41 / 1 / 2
gpt-oss-120b, 131Kyes / no / no1 / 2 / 81 / 2 / 8
gpt-oss-20b, 32Kyes / yes / yes1 / 1 / 21 / 1 / 1
gpt-oss-20b, 131Kyes / yes / no1 / 1 / 41 / 1 / 4

Our estimates, not measurements: cards needed, in configurations of 1, 2, 4 or 8, for 1 / 20 / 100 concurrent conversations of 32,768 or 131,072 tokens with a 16-bit KV cache; eight cards run as two copies of four. DGX Spark shows a memory fit for one Spark, not its speed.

On one RTX PRO 6000, gpt-oss-120b leaves 22.2 GiB for the cache, room for 19 conversations at 32K, so 20 users need a second card, and two cards as replicas hold 38. For 100 users at 32K, four cards hold 241 conversations as one split copy, 186 as two copies of two cards and 76 as four single-card copies. One H200 NVL leaves 62.5 GiB, enough for 55 conversations at 32K, and two as replicas hold 110 without a bridge.

At the full 131,072 tokens the counts fall by a factor of four. One RTX PRO 6000 then holds four conversations, one H200 NVL 13 and one Spark 7. Four H200 NVL on a four-way bridge hold 96, so the table shows eight, run as two copies of four, one per NVLink domain, which hold 192. Eight RTX PRO 6000 in that layout hold 120, and four H200 NVL with an FP8 cache hold 192.

gpt-oss-20b leaves 70.2 GiB on one RTX PRO 6000, room for 93 conversations at 32K, and one H200 NVL holds 147. One Spark holds 109 by memory, although its 273 GB/s of memory bandwidth limits how many of them stream at a usable speed; the measured figures are in our DGX Spark benchmarks.

An FP8 cache roughly doubles every count, to 111 conversations of gpt-oss-120b at 32K on one H200 NVL. TensorRT-LLM’s support matrix lists an FP8 KV cache for Hopper and for the RTX PRO Blackwell generation (sm120). Among NVIDIA GPUs, vLLM’s gpt-oss recipe sets it only for the B200, so test answer quality first.

Size for the peak number of requests in flight, which differs from headcount; our article on how many users one RTX PRO 6000 serves shows how to take that peak from gateway logs.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the gpt-oss variant, your context length and peak requests in flight, and we reply within one business day with a configuration and quote.

H200 NVL, RTX PRO 6000 or DGX Spark for gpt-oss

TensorRT-LLM’s support matrix, version 1.3.0rc29 of 26 September 2026, lists MXFP4 for Blackwell, including the RTX PRO generation, and not for Hopper. The RTX PRO 6000 and the DGX Spark have FP4 arithmetic, which an engine uses only if it has an MXFP4 kernel for them; vLLM’s B200 path pairs MXFP4 weights with MXFP8 activations. The H200 NVL holds the weights in MXFP4 too, but without FP4 arithmetic it computes those layers at higher precision, so memory per card is the same and the compute path differs.

vLLM’s gpt-oss recipe, updated 22 September 2026, names the H100, H200 and B200 among NVIDIA GPUs and mentions “ongoing work for Ampere/Ada/RTX 5090”. It does not name the RTX PRO 6000 or the DGX Spark. StorageReview served an NVFP4 version of gpt-oss-120b with vLLM on four RTX PRO 6000 Server Edition cards, as our article on one model on several GPUs over PCIe and NVLink reports. Check the engine version on the card before production.

Because gpt-oss fits one card, replicas behind a load balancer are the default; they exchange nothing over PCIe and fail independently. A split pays where the cache runs short, as for 100 conversations at 32K on RTX PRO 6000 cards. That card has no NVLink, so the split runs over PCIe, where the peer-to-peer settings decide the speed, while four H200 NVL on a four-way bridge keep it on NVLink at 900 GB/s per GPU.

vLLM settings for gpt-oss-120b

vLLM’s recipe serves gpt-oss-120b on a B200 with --tensor-parallel-size 1 and a configuration file, and describes Hopper as the same “without kv-cache-dtype and without the FlashInfer MoE flags”. vLLM’s earlier gpt-oss guide, kept in its documentation as a historical reference, lists both files: the Hopper file sets no-enable-prefix-caching: true and max-num-batched-tokens: 8192, and the Blackwell file adds kv-cache-dtype: fp8. The guide turns off prefix caching “if running with synthetic dataset”, a benchmark setting, so test prefix caching for a chat service.

The guide names two defaults that matter for memory. The context length defaults to “the maximum sequence length supported by the model”, 131,072 tokens. The number of sequences defaults to “a large number like 1024 on GPUs with large memory sizes”. Set --max-model-len to the context your users need, and vLLM reports a higher maximum concurrency for the same cache.

Operating system, drivers, CUDA and a container runtime are installed on request, and deploying the model with vLLM is part of our Private AI/ML service. Write to us with the variant, the reasoning level you plan and your rack position.

Licence and usage policy for gpt-oss

gpt-oss-120b, gpt-oss-20b and both safeguard models are published under the Apache 2.0 licence, as their model cards state. Each OpenAI repository also carries a short USAGE_POLICY file, which for gpt-oss-120b reads: “By using OpenAI gpt-oss-120b, you agree to comply with all applicable law.” NVIDIA’s gpt-oss-puzzle-88B comes under the NVIDIA Open Model License instead. Whether these terms fit your intended use is a legal assessment for your legal department.

What we supply

We supply the DGX Spark Founders Edition, the RTX PRO 6000 in its Workstation, Max-Q and Server editions, and the H200 NVL with its two-way and four-way NVLink bridges, as cards or in AI servers built to order, on one EU contract and invoice with manufacturer warranty. For gpt-oss we size the server from the variant, the context and the peak requests in flight, and we check the rack, power and airflow before we quote. The configuration and quote follow within one business day, and our professional GPU range lists every card. Running the model, RAG and MLOps on top is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

How much VRAM does gpt-oss-120b need?
Its checkpoint, with the expert weights in MXFP4, is 65.3 GB or 60.8 GiB, and OpenAI’s model card says it runs on a single 80 GB GPU. Each conversation adds 1.125 GiB of 16-bit KV cache at 32,768 tokens and 4.5 GiB at the full 131,072. One 96 GB RTX PRO 6000 holds about 19 conversations at 32K and one 141 GB H200 NVL about 55, by our estimate.
How much VRAM does gpt-oss-20b need?
The checkpoint is 13.8 GB or 12.8 GiB, and OpenAI states that the model runs within 16 GB of memory. Its KV cache costs 0.75 GiB per conversation at 32,768 tokens in 16-bit. By our estimate one RTX PRO 6000 holds about 93 such conversations and one H200 NVL about 147.
Does gpt-oss-120b run on one H200 NVL?
Yes. The 60.8 GiB of weights leave about 62.5 GiB for the cache on one H200 NVL, room for about 55 conversations at 32K or 13 at the full 131,072 tokens. Hopper has no FP4 arithmetic, so the MXFP4 expert weights are stored in 4-bit and computed at higher precision.
What GPU server does gpt-oss-120b need for 100 users?
For 100 concurrent conversations of 32,768 tokens with a 16-bit cache, two H200 NVL as separate copies, or four RTX PRO 6000 as one split copy or two copies of two cards, by our estimate. At the full 131,072 tokens the same load needs eight of either card, run as two copies of four, or four H200 NVL with an FP8 cache. The figure counts requests in flight at peak, which differ from headcount.
Can I run gpt-oss-120b locally on a DGX Spark?
Yes. One DGX Spark with 128 GB of unified memory holds the 60.8 GiB checkpoint with cache for about 30 conversations at 32K by our memory estimate. Its 273 GB/s of memory bandwidth limits speed before memory, so check measured figures for this model before you plan a team on one Spark.
Can gpt-oss be used commercially on premise?
gpt-oss-120b and gpt-oss-20b are published under the Apache 2.0 licence, and their repositories add a usage policy under which users agree to comply with all applicable law. NVIDIA’s derived gpt-oss-puzzle-88B has the NVIDIA Open Model License instead. Whether the terms fit a specific use is a legal assessment for the company’s legal department.

Send us the gpt-oss variant, the reasoning level you plan, the context you will declare and your peak requests in flight. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna