GPU energy per token for LLM inference: H200 NVL, RTX PRO 6000 and L40S power, joules and kWh
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Energy per token is the power drawn divided by the tokens produced per second, since one watt is one joule per second; 1 J per token equals about 278 Wh per million tokens
- NVIDIA rates the H200 NVL and the RTX PRO 6000 Server Edition at up to 600 W, with minimums of 200 W and 300 W in its product briefs, and the L40S at 350 W; nvidia-smi reports the range each card accepts as Min and Max Power Limit
- MLPerf Inference v6.0 publishes Llama 2 70B throughput for both 600 W cards but no power results; at 8 × 600 W, the offline results give, by our calculation, a ceiling of 0.150 J per token for eight H200 NVL in FP8 and 0.160 J for eight RTX PRO 6000 in FP4
- Batch size and task move energy per token strongly: ML.ENERGY measured 0.209 J per token for Qwen3 32B on one B200 at batch size 128 and 0.151 J at 512, and 0.312 J for problem-solving answers at 128
- Measure your own figure with DCGM’s energy counter, field 156 in millijoules, and vLLM’s token counters over the same window, at your concurrency and with the power limits you plan to run
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Energy per token for LLM inference: the formula
Energy per token in LLM inference is the power a GPU draws divided by the tokens it produces per second. A watt is one joule per second, so watts divided by tokens per second gives joules per token, and since one kWh is 3.6 million joules, 1 J per token equals about 278 Wh per million tokens. A card under full load spreads its power over many tokens, while a lightly loaded card spreads it over few, so the batch size and the utilisation of the server decide much of the result, alongside the card model.
As of October 2026 we found no published per-token energy measurement for the H200 NVL, the RTX PRO 6000 Server Edition or the L40S. MLPerf Inference v6.0 publishes throughput for the first two, and with their rated board power that throughput gives an upper bound for GPU energy, by our calculation. The method section below shows how to measure your own.
Board power and power limits: H200 NVL, RTX PRO 6000 and L40S
| CARD | MEMORY, BANDWIDTH | MAX BOARD POWER | POWER LIMIT SETTING |
|---|---|---|---|
| NVIDIA H200 NVL | 141 GB HBM3e, 4.8 TB/s | up to 600 W | 200 W minimum, 350 W power compliance limit; Lenovo lists 600 W |
| RTX PRO 6000 Server Edition | 96 GB GDDR7, 1,597 GB/s | up to 600 W | 300 W minimum in the 600 W and 450 W modes; Lenovo offers a 450 W capped option |
| NVIDIA L40S | 48 GB GDDR6, 864 GB/s | 350 W | NVIDIA’s page lists no configurable range |
NVIDIA product pages for the H200, the RTX PRO 6000 Blackwell Server Edition and the L40S, the Server Edition datasheet 4682150 (December 2025), read on 10 October 2026; NVIDIA product briefs PB-12128-001_v01 (H200 NVL, 11 April 2025) and SP-12355-001_v02 (Server Edition, 27 June 2025); Lenovo Press product guides LP2263 (updated 28 July 2026) and LP1944 (updated 7 November 2025).
NVIDIA gives the H200 NVL a maximum thermal design power of “Up to 600W (configurable)”, and uses the same words for the maximum power consumption of the RTX PRO 6000 Blackwell Server Edition. NVIDIA’s H200 NVL product brief PB-12128-001_v01 of 11 April 2025 lists a 600 W default maximum, a 350 W power compliance limit and a 200 W minimum. The Server Edition brief SP-12355-001_v02 of 27 June 2025 lists a 300 W minimum in its 600 W and 450 W cable modes. Lenovo’s product guide for the Server Edition lists “600 W (can also be power capped to 450 W to support increased density)”, an option that lets four cards fit in its SR650a V4. NVIDIA lists the L40S at 350 W. How the two passive cards compare beyond power is in our comparison of the L40S and the RTX PRO 6000 Server Edition.
The range a delivered card accepts comes from the driver. nvidia-smi -q -d POWER shows its Min Power Limit and its Max Power Limit, which NVIDIA describes as “The maximum value in watts that power limit can be set to.” On a Server Edition card it can also read 450 W because of the power cable, which our rack power guide explains. The rated figure caps the board, and the card draws less when the work does not need it. The ML.ENERGY researchers call energy estimates from TDP “nearly always an overestimation”, because a GPU rarely draws its maximum power at every moment.
We supply all three cards and build GPU servers around them. Tell us the power available at the rack position and the model you plan to run, and we reply with a configuration and quote within one business day.
Published results: MLPerf Inference v6.0 and ML.ENERGY
Where MLPerf® Inference measures power, it measures the whole system. MLCommons computes its power metrics from “the measured average AC power (energy) consumed by the entire system” during the benchmark run, and states that “MLPerf Power is only capable of measuring and validating the full system power.” In the summary results file of the v6.0 round, published on 1 April 2026, none of the 520 result rows carries a power measurement. The L40S appears in that round only in vision and speech recognition results, not in an LLM benchmark.
The throughput results still bound the energy. Eight cards limited to 600 W draw at most 4,800 W on their boards, and 4,800 W divided by the tokens per second of a result gives the most GPU energy each token can have taken. These ceilings are our calculation, not measurements. The table sets them beside published measurements from the ML.ENERGY team, taken on a B200, a GPU outside this comparison.
| RESULT | SETUP | TOKENS/S | J PER TOKEN | BASIS |
|---|---|---|---|---|
| 6.0-0021, Offline | 8 × H200 NVL, FP8, Llama 2 70B | 32,004.0 | 0.150 | our ceiling, 8 × 600 W |
| 6.0-0021, Interactive | 8 × H200 NVL, FP8, Llama 2 70B | 16,343.8 | 0.294 | our ceiling, 8 × 600 W |
| 6.0-0047, Offline | 8 × RTX PRO 6000 SE, FP4, Llama 2 70B | 29,908.4 | 0.160 | our ceiling, 8 × 600 W |
| 6.0-0004, Interactive | 8 × RTX PRO 6000 SE, FP4, Llama 2 70B | 6,238.1 | 0.769 | our ceiling, 8 × 600 W |
| ML.ENERGY, text chat | 1 × B200, Qwen3 32B, batch 128 | not stated | 0.209 | measured GPU energy |
| ML.ENERGY, text chat | 1 × B200, Qwen3 32B, batch 512 | not stated | 0.151 | measured GPU energy |
| ML.ENERGY, problem solving | 1 × B200, Qwen3 32B, batch 128 | not stated | 0.312 | measured GPU energy |
MLPerf Inference v6.0, published by MLCommons on 1 April 2026, data centre suite, closed division, available systems, benchmark llama2-70b-99: entry 6.0-0021 (Dell PowerEdge XE7740, listed with “H200 TGP 600W”), 6.0-0047 (HPE ProLiant Compute DL380a Gen12) and 6.0-0004; from MLCommons’ v6.0 summary results file; results verified by MLCommons Association. The J per token ceilings are our calculation (4,800 W ÷ tokens per second) for the GPU boards only, not measurements and not an MLPerf metric. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information. ML.ENERGY team, University of Michigan, “Where Do the Joules Go? Diagnosing Inference Energy Consumption”, arXiv 2601.22076v2 (30 January 2026), Table 1.
In the offline scenario, where latency does not count, the ceilings of the two 600 W cards end up close, at 0.150 J per token for the H200 NVL in FP8 and 0.160 J for the RTX PRO 6000 in FP4. Under the latency limits of the interactive scenario, another eight-card RTX PRO 6000 system (6.0-0004) produced far fewer tokens, so the same 4,800 W ceiling works out at 0.769 J per token against 0.294 J for the H200 NVL. Neither pair compares measured efficiency, because the cards may have drawn well below 600 W, above all under the latency limits. Our comparison of the RTX PRO 6000 and the H200 NVL for LLM inference explains the memory bandwidth behind that gap.
The ML.ENERGY team measured GPU energy across 46 models and 1,858 configurations on H100 and B200 GPUs. The authors also report energy differences of three to five times caused by differences in GPU utilisation.
Why batch size and utilisation decide energy per token
During generation, every step reads the weights of a dense model such as Llama 2 70B once for the whole batch, and each request in the batch adds its own KV cache traffic. A larger batch therefore spreads the same weight reads, and much of the same board power, over more tokens. The ML.ENERGY study puts it this way: as batch size increases, “energy per token drops at first, then plateaus as GPU utilization approaches saturation.” In the team’s earlier benchmark paper, the energy per answer of a distilled 8B reasoning model on one H100 fell to about 29 per cent of its value at a maximum batch size of 4 when the maximum was raised to 64.
Utilisation over the day works the same way. A server that waits for requests still draws idle power, and that energy belongs to the tokens it produces in the busy hours. A pilot with a handful of users therefore reports a much higher figure per token than the same server under steady load, so compare figures only at the same load.
Answer length changes the energy per request more than the per-token figure does. For Qwen3 32B on one B200, each task at its largest batch size, ML.ENERGY measured 95 J per response for text conversation and 2,192 J for problem solving, whose answers averaged 7,035 tokens against 627. Plan energy per answer for reasoning models, not only per token.
Measuring energy per token on your own server
A figure for a budget has to come from your model, prompts and concurrency; DCGM and vLLM provide the two counters.
- With the model loaded and no requests, note each GPU’s idle board power from
nvidia-smi -q -d POWER, whose Average Power Draw NVIDIA defines as “The average power draw for the entire board for the last second, in watts.” - Before the run, read DCGM field 156, DCGM_
FI_ DEV_ TOTAL_ ENERGY_ CONSUMPTION, defined as “Total energy consumption for the GPU in mJ since the driver was last reloaded”, for every GPU; dcgmi dmon -e 155,156shows it beside board power in watts. - At the same moment, read vLLM’s counters
vllm:generation_tokens_totalandvllm:prompt_tokens_totalfrom its/metricsendpoint. - Run a load with the concurrency, prompt length and answer length of your workload for a fixed period, then read both counters again.
- Divide the energy difference, summed over all GPUs and converted from millijoules to joules, by the generated tokens. Report prompt tokens beside it, since long RAG prompts add prefill energy that a figure per output token hides.
- Repeat at two or three concurrency levels and at each power limit you consider, and keep tokens per second per user beside every result.
Board energy leaves out the processors, memory, fans and power-supply losses. For the whole server, read the input power from the BMC or a metered PDU over the same window and divide by the same token count. The other DCGM fields worth exporting, and the field names DCGM 4.6 changed, are in our guide to GPU server monitoring with DCGM.
We build AI servers to order and check the rack, power and airflow before we quote. Send us the model, your daily token volume and the power budget of the rack position through the form below.
Power limits as an efficiency setting
On a 600 W card, nvidia-smi -i 0 -pl 450 sets a lower limit on the first GPU. NVIDIA’s documentation says the value “needs to be between Min and Max Power Limit as reported by nvidia-smi” and that the command “Requires root.” A lower limit reduces peak draw and heat per card, and throughput can fall by an amount that depends on the model, the precision and the batch size. Whether energy per token falls with it is a question for step 6 above, at the default limit and one or two lower ones.
After a reboot or driver reload, check that the limit still applies. DCGM reports it in DCGM_
From joules per token to kWh per day
To convert, multiply joules per token by one million and divide by 3,600 for Wh per million tokens. The offline ceiling of 0.150 J per token for eight H200 NVL comes to 41.7 Wh per million tokens of GPU board energy. The RTX PRO 6000 result comes to 44.6 Wh.
A daily figure is simpler to bound. Eight cards at 600 W use at most 115.2 kWh of board energy in 24 hours, whatever they produce, and idle hours count towards that total. The wall power of the server adds processors, fans and power-supply losses, and a data centre adds its cooling and distribution on top. All of that energy ends up as heat in the room, which our article on a GPU server in an office or small server room converts to BTU/h.
At EU level, Commission Delegated Regulation (EU) 2024/1364 requires “operators of data centres with an installed information technology power demand of at least 500 kW” to report to a European database every year, directly or through a national reporting scheme, including the total energy consumption of the data centre and of its IT equipment in kWh. Whether a given site falls under it is a legal assessment for the company’s legal department.
What we supply
We supply the H200 NVL, the RTX PRO 6000 Server Edition and the L40S on their own for servers you already run, and in AI servers built to order, with manufacturer warranty on one EU contract and invoice. Each server is assembled and burn-in tested, and we check the rack, power and airflow before we quote. Operating system, drivers, CUDA and a container runtime are installed on request, and the driver includes nvidia-smi for reading power and setting limits. The full card line-up with each card’s board power is on our professional NVIDIA GPU page.
FAQ
How much energy does a GPU use per token in LLM inference?
How do I calculate energy per token for LLM inference?
How much electricity does an AI server use for LLM inference?
Is the H200 NVL more energy efficient than the RTX PRO 6000?
How many tokens per watt does a GPU deliver?
Does lowering the GPU power limit save energy?
Send us the model and its precision, your expected daily token volume, the peak number of concurrent requests and the power available at the rack position. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day