BLOG · GUIDE ·

Reasoning model hardware: why thinking tokens change GPU sizing for gpt-oss, Qwen and DeepSeek

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A reasoning model writes hidden thinking tokens before the answer, so the weights stay the same size while output tokens per request, KV cache per conversation and time in flight grow
  • Stated lengths as read on 10 October 2026: gpt-oss-20b used over 20,000 reasoning tokens per AIME problem on average, DeepSeek-R1-0528 averages 23,000 tokens per AIME question, and DeepSeek recommends 384K tokens of output for V4-Flash at high and max effort
  • Requests in flight equal arrival rate times time in flight, so at 18 tokens per second per stream a request that writes 8,000 tokens stays in flight about 444 seconds, against the 30 seconds we assume for a 500-token chat answer
  • For 500 employees on gpt-oss-120b at a declared 32K context, our estimate rises from 2 RTX PRO 6000 or 1 H200 NVL for chat to 5 to 16 RTX PRO 6000 or 2 to 6 H200 NVL at 8,000 output tokens per request, depending on how much of the context each request fills
  • Reasoning effort levels, enable_thinking and vLLM’s per-request thinking_token_budget control the length, and gpt-oss drops earlier reasoning from the history while Qwen3.8 keeps it by default

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Reasoning model hardware: what thinking tokens change

A reasoning model writes hidden thinking tokens before its answer, from a few hundred to tens of thousands per request, and the GPUs generate and store every one of them. The weights keep their size, so the memory for one copy of the model is unchanged. What grows is the number of output tokens per request, the KV cache each conversation holds while it reasons and the time each request stays in flight. With the same staff, more requests are in flight at once, each holding a longer cache.

In the worked example below, 500 employees on gpt-oss-120b need two RTX PRO 6000 or one H200 NVL for 500-token chat answers at a declared 32K context, by our estimate. When every request writes 8,000 tokens, they need 16 RTX PRO 6000 or six H200 NVL if each request keeps the full 32K in reserve, and five or two if it holds only the 10,000 tokens it uses. The reasoning level is therefore a sizing input, set per use case, like the context length.

NVIDIA described the trend in its Dynamo announcement of 18 March 2025: as models add reasoning abilities and run in agentic workflows, “they generate a far greater number of tokens during inference.”

Reasoning controls and stated lengths by model

Each model family documents its own control. gpt-oss takes a reasoning level in the system prompt, and OpenAI’s harmony guide of 5 August 2025 states that “By default, the model will do medium level reasoning.” Qwen3.8 thinks by default, accepts a reasoning_effort of xhigh, medium or low, and turns thinking off with enable_thinking set to false.

MODELREASONING CONTROLDEFAULTSTATED LENGTH
gpt-oss-120b and gpt-oss-20blow, medium, high in the system promptmediumgpt-oss-20b: over 20,000 tokens per AIME problem on average
Qwen3.8-27Breasoning_effort xhigh, medium, low; enable_thinking offthinking on, xhighoutput limit of 262,144 tokens recommended for reasoning content
DeepSeek-V4-Flash-0731reasoning_effort low, high, maxnot stated on the card384K tokens of output recommended for high and max
DeepSeek-R1-0528no effort setting on the cardthinking23,000 tokens per AIME question on average, 12,000 before

OpenAI’s gpt-oss model card (arXiv 2508.10925) and harmony guide, Qwen’s Qwen3.8-27B card, DeepSeek’s V4-Flash-0731 and R1-0528 cards on Hugging Face, all read on 10 October 2026. AIME figures are from competition mathematics, not office tasks.

The figures in the table are output ceilings and averages on competition mathematics, and the lengths of your own use cases are measured in a pilot, as described below. The gpt-oss model card says that “Increasing the reasoning level will cause the model’s average CoT length to increase”, and that longer reasoning brings “higher accuracy at a relatively large increase in final response latency and cost.” DeepSeek’s R1-0528 card reports that its update raised the average from 12,000 to 23,000 tokens per AIME question, and its evaluations allowed up to 64K tokens of generation. Qwen’s card adds a caution for agents: “In multi-turn agentic tasks, lower reasoning effort does not always reduce overall task completion time”, because weak analysis leads to retries.

Decode load and time in flight per request

Prefill reads the prompt once and is compute-bound, while decode writes one token after another and, in NVIDIA’s words, “is memory-bound”. Thinking tokens are decode tokens. The decode load of a service is its arrival rate times the output tokens per request, and by Little’s law the requests in flight are the arrival rate times the time each request takes.

We take the example values of our article on a private ChatGPT server by company size for 500 employees: 200 busy-hour users, 6 requests an hour each and a peak factor of 2. That is one request every 1.5 seconds at the peak. For the time in flight we assume 18 tokens per second per stream, the median of the gpt-oss-120b server scenario in MLPerf Inference v6.0 on an eight-card RTX PRO 6000 server, as our RTX PRO 6000 benchmarks report. A 500-token answer then takes about 28 seconds, close to the 30 seconds of the company-size article.

PROFILETIME IN FLIGHTPEAK IN FLIGHTPEAK DECODE LOADCARDS AT 32K
Chat, 500 tokens out30 s20333 tokens/s2 RTX PRO 6000 or 1 H200 NVL
Reasoning, 2,000 out111 s741,333 tokens/s4 RTX PRO 6000 or 2 H200 NVL
Reasoning, 8,000 out444 s2965,333 tokens/s16 RTX PRO 6000 or 6 H200 NVL

Our estimates, not measurements, for gpt-oss-120b with a 16-bit KV cache: example values from our company-size article, including its 30 seconds per chat request; reasoning rows at 18 tokens per second per stream from MLPerf Inference v6.0 (result 6.0-0047); cards by the company-size method of 19 conversations per RTX PRO 6000 and 55 per H200 NVL at a full 32K, one copy per card.

Output length enters twice. At 18 tokens per second an 8,000-token request stays in flight about 444 seconds, against the 30 seconds of a chat answer, so about 15 times as many requests are in flight at the peak, each holding a longer cache.

The card counts in the table follow the company-size method and reserve the full declared 32,768 tokens for every request in flight, as if all of them filled their context at the same moment. Counting only the tokens each request holds when it ends, 2,000 of prompt plus 8,000 of output at 36 KiB per token, the 296 requests of the last row take about 103 GiB of cache, which five RTX PRO 6000 or two H200 NVL hold. The gap between 5 and 16 cards is the context each request reserves but does not use, 10,000 tokens held against 32,768 reserved. The lower figure holds while prompts and history stay near 2,000 tokens. When users paste long documents or keep long chats, requests approach the declared context and the count moves towards 16. If the cache runs out, vLLM preempts requests and recomputes them later, as its tuning guide describes, so measure prompt lengths in a pilot before you choose a point between the two.

Memory is not the only limit. At 5,333 tokens per second at the peak, the 8,000-token profile needs at least three RTX PRO 6000 by throughput, since MLPerf’s eight-card system delivered 1,782 tokens per second per card in the same scenario, by our division of the system result. Those cards ran fully loaded, so 18 tokens per second is a loaded-card value; with fewer streams per card each request finishes sooner, and the table errs on the high side. If the service must survive a server failure, the two-server rule of the company-size article applies on top.

For a pilot, one DGX Spark (128 GB) holds gpt-oss-120b with about 30 conversations at 32K by memory, counting 102 GB of its memory as our hub does. Its 273 GB/s of memory bandwidth limits it before memory does. The llama.cpp results in our article on sharing one DGX Spark in a team work out to about 12 tokens per second per request at 16 at once, so an 8,000-token reasoning request would run for about 11 minutes.

We build inference servers with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the model, the reasoning level per use case and your peak requests through the form below, and we reply with a configuration and quote.

KV cache per conversation when the model reasons

The cache per token is fixed by the model’s attention design, and reasoning only multiplies the tokens. On gpt-oss-120b each token costs 36 KiB of 16-bit cache in the full-attention layers, as our gpt-oss-120b hardware guide derives from its config.json, so 20,000 reasoning tokens add about 0.7 GiB to one conversation. Qwen3.8-27B at 64 KiB per token reaches about 16 GiB at the recommended 262,144 tokens of reasoning output, which is also its native context length. One RTX PRO 6000 with its FP8 weights keeps 54.3 GiB for the cache, as our Qwen hardware guide shows, which is room for three such conversations.

The chat history adds to the cache as well. OpenAI’s harmony guide says to “drop any previous CoT content on subsequent sampling” once a reply has ended in the final channel, with tool and function calls as the exception. A gpt-oss chat therefore grows by questions and answers only. Qwen’s card states that “By default, Qwen3.8 retains thinking blocks from all historical messages”, so every turn carries the reasoning of earlier turns unless preserve_thinking is set to false. For a multi-turn assistant on Qwen3.8 that setting changes the context you declare.

DeepSeek R1 hardware and large reasoning models

DeepSeek-R1-0528 is listed at 685B parameters, and its config.json gives the DeepSeek-V3 architecture (DeepseekV3ForCausalLM), 61 layers, FP8 weights in 128 × 128 blocks and 163,840 tokens of context. By our sum of the parameter counts per data type on its Hugging Face page, the weights take about 689 GB, close to the 690 GB FP8 checkpoint of DeepSeek-V3.2, which our DeepSeek hardware guide puts on eight H200 NVL. By our reading R1-0528 needs the same eight cards by weights.

Its multi-head latent attention stores 512 + 64 values per layer and token, 68.6 KiB in 16-bit by our arithmetic, without V3.2’s sparse-attention indexer. An average AIME answer of 23,000 tokens then holds about 1.5 GiB of cache, and the 64K tokens of DeepSeek’s evaluations about 4.3 GiB. Under tensor parallelism the latent cache is copied to every card, so with the about 43 GiB per card our DeepSeek guide finds on eight H200 NVL for V3.2, about 28 such answers fit at once, by our estimate. The R1-0528 card also describes a distilled DeepSeek-R1-0528-Qwen3-8B with the architecture of Qwen3-8B, about 16 GB of BF16 weights for its 8.2B parameters, which fits a 24 GB L4 or RTX PRO 4000 by weights. Qwen3-8B’s config.json gives 144 KiB of 16-bit cache per token, about 3.2 GiB for a 23,000-token answer, so a 24 GB card has room beside the weights for two such answers at most.

DeepSeek-V4-Flash allows far longer reasoning with a much smaller cache per token. At the recommended 384K tokens of output, our DeepSeek guide estimates about 1.5 GiB of FP8 cache per conversation on every card.

Serving reasoning models in vLLM

vLLM separates reasoning from the answer when the server starts with --reasoning-parser; its documentation, a developer preview dated 1 October 2026, lists deepseek_r1 for the DeepSeek R1 series and qwen3 for the Qwen3 series. The response then carries a reasoning field next to the content, which earlier versions called reasoning_content.

Two controls act per request. For models that use enable_thinking, a reasoning_effort of low, medium or high sets it to true and none sets it to false, unless the request sets enable_thinking itself. The sampling parameter thinking_token_budget sets a per-request limit, and once it is reached “vLLM forces the model to produce reasoning_end_str”, the end marker set with --reasoning-config or taken from the reasoning parser. Without a budget, the documentation states, “no explicit reasoning limit is applied beyond normal generation constraints such as max_tokens”.

Set --max-model-len to the prompt, the history and the longest reasoning you allow, not to the model’s maximum. A budget that cuts reasoning short trades answer quality against load, so test it on your evaluation set.

Measuring reasoning load in a pilot

Model cards give benchmark averages, so a pilot measures the output length of your own use cases.

  1. Log every request at the gateway with its start and end time and its reasoning and answer token counts.
  2. Run each use case at the reasoning level you plan, and once at the next lower level.
  3. Take the busiest hour’s arrival rate and the median and 95th percentile of output tokens per request.
  4. Compute requests in flight and decode load as in the table above, and load-test the candidate GPU with those lengths.

Our Private AI/ML service starts with a pilot on one process with clear metrics. Tell us which process the pilot should cover and the reasoning level you expect it to need.

What we supply

We supply the RTX PRO 6000 in its Workstation, Max-Q and Server editions, the H200 NVL with two-way and four-way NVLink bridges, the L40S, the L4 and the DGX Spark Founders Edition, as cards from our professional GPU range or in AI servers built to order with 2 to 8 GPUs per node. For a reasoning model we size the server from the model, the reasoning level per use case, the declared context and the peak requests in flight. We check the rack, power and airflow before we quote, and the configuration and quote follow within one business day, on one EU contract and invoice with manufacturer warranty. Deploying the model with vLLM, RAG and MLOps is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What hardware does a reasoning model need?
The same GPU memory for its weights as the base model, plus more KV cache and decode capacity, because it writes thinking tokens before each answer. For 500 employees on gpt-oss-120b at a declared 32K context, our estimate rises from two RTX PRO 6000 or one H200 NVL for 500-token chat answers to between 5 and 16 RTX PRO 6000, or between 2 and 6 H200 NVL, when each request writes 8,000 tokens, depending on how much of the context each request fills. Measure prompt and output lengths in a pilot before you size the server.
Do reasoning models need more VRAM?
They need the same memory for the weights and more for the KV cache, since every thinking token is stored like an answer token. On gpt-oss-120b each token costs 36 KiB of 16-bit cache, so 20,000 reasoning tokens add about 0.7 GiB per conversation. Because each request stays in flight longer, more conversations hold cache at the same time.
Does gpt-oss reasoning effort change the GPU requirements?
Yes, through output length and time in flight rather than weights. OpenAI’s model card states that a higher reasoning level increases the average chain-of-thought length, and gpt-oss-20b used over 20,000 reasoning tokens per AIME problem on average. The level is set in the system prompt, and OpenAI’s harmony guide gives medium as the default.
What hardware does DeepSeek R1 need?
DeepSeek-R1-0528 has 685B parameters in FP8 with the DeepSeek-V3 architecture, about 689 GB of weights, which by our reading needs eight H200 NVL, as DeepSeek-V3.2 does. Its average AIME answer of 23,000 tokens holds about 1.5 GiB of 16-bit cache, copied to every card under tensor parallelism. The distilled DeepSeek-R1-0528-Qwen3-8B, about 16 GB in BF16, fits a 24 GB card such as the L4 by weights, but at about 3.2 GiB of cache per 23,000-token answer a 48 GB card such as the L40S serves several at once.
How do I limit thinking tokens in vLLM?
Start the server with a reasoning parser and pass thinking_token_budget per request; once the budget is reached, vLLM forces the model to emit the end-of-reasoning marker. Setting reasoning_effort to none turns thinking off for models that use enable_thinking. Without a budget only max_tokens limits the reasoning.
How much more throughput does a reasoning model need than chat?
Decode load is the arrival rate times the output tokens per request, so 16 times the output means 16 times the tokens per second. For 500 employees with our example values, that is about 333 tokens per second at the peak for 500-token answers and 5,333 for 8,000-token requests. These are our estimates, and a load test with your own prompts confirms them.

Send us the model, the reasoning level or thinking budget per use case, the context length and your peak requests in flight. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna