BLOG · COMPARISON ·

GPU energy per token for LLM inference: H200 NVL, RTX PRO 6000 and L40S power, joules and kWh

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Energy per token is the power drawn divided by the tokens produced per second, since one watt is one joule per second; 1 J per token equals about 278 Wh per million tokens
  • NVIDIA rates the H200 NVL and the RTX PRO 6000 Server Edition at up to 600 W, with minimums of 200 W and 300 W in its product briefs, and the L40S at 350 W; nvidia-smi reports the range each card accepts as Min and Max Power Limit
  • MLPerf Inference v6.0 publishes Llama 2 70B throughput for both 600 W cards but no power results; at 8 × 600 W, the offline results give, by our calculation, a ceiling of 0.150 J per token for eight H200 NVL in FP8 and 0.160 J for eight RTX PRO 6000 in FP4
  • Batch size and task move energy per token strongly: ML.ENERGY measured 0.209 J per token for Qwen3 32B on one B200 at batch size 128 and 0.151 J at 512, and 0.312 J for problem-solving answers at 128
  • Measure your own figure with DCGM’s energy counter, field 156 in millijoules, and vLLM’s token counters over the same window, at your concurrency and with the power limits you plan to run

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Energy per token for LLM inference: the formula

Energy per token in LLM inference is the power a GPU draws divided by the tokens it produces per second. A watt is one joule per second, so watts divided by tokens per second gives joules per token, and since one kWh is 3.6 million joules, 1 J per token equals about 278 Wh per million tokens. A card under full load spreads its power over many tokens, while a lightly loaded card spreads it over few, so the batch size and the utilisation of the server decide much of the result, alongside the card model.

As of October 2026 we found no published per-token energy measurement for the H200 NVL, the RTX PRO 6000 Server Edition or the L40S. MLPerf Inference v6.0 publishes throughput for the first two, and with their rated board power that throughput gives an upper bound for GPU energy, by our calculation. The method section below shows how to measure your own.

Board power and power limits: H200 NVL, RTX PRO 6000 and L40S

CARDMEMORY, BANDWIDTHMAX BOARD POWERPOWER LIMIT SETTING
NVIDIA H200 NVL141 GB HBM3e, 4.8 TB/sup to 600 W200 W minimum, 350 W power compliance limit; Lenovo lists 600 W
RTX PRO 6000 Server Edition96 GB GDDR7, 1,597 GB/sup to 600 W300 W minimum in the 600 W and 450 W modes; Lenovo offers a 450 W capped option
NVIDIA L40S48 GB GDDR6, 864 GB/s350 WNVIDIA’s page lists no configurable range

NVIDIA product pages for the H200, the RTX PRO 6000 Blackwell Server Edition and the L40S, the Server Edition datasheet 4682150 (December 2025), read on 10 October 2026; NVIDIA product briefs PB-12128-001_v01 (H200 NVL, 11 April 2025) and SP-12355-001_v02 (Server Edition, 27 June 2025); Lenovo Press product guides LP2263 (updated 28 July 2026) and LP1944 (updated 7 November 2025).

NVIDIA gives the H200 NVL a maximum thermal design power of “Up to 600W (configurable)”, and uses the same words for the maximum power consumption of the RTX PRO 6000 Blackwell Server Edition. NVIDIA’s H200 NVL product brief PB-12128-001_v01 of 11 April 2025 lists a 600 W default maximum, a 350 W power compliance limit and a 200 W minimum. The Server Edition brief SP-12355-001_v02 of 27 June 2025 lists a 300 W minimum in its 600 W and 450 W cable modes. Lenovo’s product guide for the Server Edition lists “600 W (can also be power capped to 450 W to support increased density)”, an option that lets four cards fit in its SR650a V4. NVIDIA lists the L40S at 350 W. How the two passive cards compare beyond power is in our comparison of the L40S and the RTX PRO 6000 Server Edition.

The range a delivered card accepts comes from the driver. nvidia-smi -q -d POWER shows its Min Power Limit and its Max Power Limit, which NVIDIA describes as “The maximum value in watts that power limit can be set to.” On a Server Edition card it can also read 450 W because of the power cable, which our rack power guide explains. The rated figure caps the board, and the card draws less when the work does not need it. The ML.ENERGY researchers call energy estimates from TDP “nearly always an overestimation”, because a GPU rarely draws its maximum power at every moment.

We supply all three cards and build GPU servers around them. Tell us the power available at the rack position and the model you plan to run, and we reply with a configuration and quote within one business day.

Published results: MLPerf Inference v6.0 and ML.ENERGY

Where MLPerf® Inference measures power, it measures the whole system. MLCommons computes its power metrics from “the measured average AC power (energy) consumed by the entire system” during the benchmark run, and states that “MLPerf Power is only capable of measuring and validating the full system power.” In the summary results file of the v6.0 round, published on 1 April 2026, none of the 520 result rows carries a power measurement. The L40S appears in that round only in vision and speech recognition results, not in an LLM benchmark.

The throughput results still bound the energy. Eight cards limited to 600 W draw at most 4,800 W on their boards, and 4,800 W divided by the tokens per second of a result gives the most GPU energy each token can have taken. These ceilings are our calculation, not measurements. The table sets them beside published measurements from the ML.ENERGY team, taken on a B200, a GPU outside this comparison.

RESULTSETUPTOKENS/SJ PER TOKENBASIS
6.0-0021, Offline8 × H200 NVL, FP8, Llama 2 70B32,004.00.150our ceiling, 8 × 600 W
6.0-0021, Interactive8 × H200 NVL, FP8, Llama 2 70B16,343.80.294our ceiling, 8 × 600 W
6.0-0047, Offline8 × RTX PRO 6000 SE, FP4, Llama 2 70B29,908.40.160our ceiling, 8 × 600 W
6.0-0004, Interactive8 × RTX PRO 6000 SE, FP4, Llama 2 70B6,238.10.769our ceiling, 8 × 600 W
ML.ENERGY, text chat1 × B200, Qwen3 32B, batch 128not stated0.209measured GPU energy
ML.ENERGY, text chat1 × B200, Qwen3 32B, batch 512not stated0.151measured GPU energy
ML.ENERGY, problem solving1 × B200, Qwen3 32B, batch 128not stated0.312measured GPU energy

MLPerf Inference v6.0, published by MLCommons on 1 April 2026, data centre suite, closed division, available systems, benchmark llama2-70b-99: entry 6.0-0021 (Dell PowerEdge XE7740, listed with “H200 TGP 600W”), 6.0-0047 (HPE ProLiant Compute DL380a Gen12) and 6.0-0004; from MLCommons’ v6.0 summary results file; results verified by MLCommons Association. The J per token ceilings are our calculation (4,800 W ÷ tokens per second) for the GPU boards only, not measurements and not an MLPerf metric. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information. ML.ENERGY team, University of Michigan, “Where Do the Joules Go? Diagnosing Inference Energy Consumption”, arXiv 2601.22076v2 (30 January 2026), Table 1.

In the offline scenario, where latency does not count, the ceilings of the two 600 W cards end up close, at 0.150 J per token for the H200 NVL in FP8 and 0.160 J for the RTX PRO 6000 in FP4. Under the latency limits of the interactive scenario, another eight-card RTX PRO 6000 system (6.0-0004) produced far fewer tokens, so the same 4,800 W ceiling works out at 0.769 J per token against 0.294 J for the H200 NVL. Neither pair compares measured efficiency, because the cards may have drawn well below 600 W, above all under the latency limits. Our comparison of the RTX PRO 6000 and the H200 NVL for LLM inference explains the memory bandwidth behind that gap.

The ML.ENERGY team measured GPU energy across 46 models and 1,858 configurations on H100 and B200 GPUs. The authors also report energy differences of three to five times caused by differences in GPU utilisation.

Why batch size and utilisation decide energy per token

During generation, every step reads the weights of a dense model such as Llama 2 70B once for the whole batch, and each request in the batch adds its own KV cache traffic. A larger batch therefore spreads the same weight reads, and much of the same board power, over more tokens. The ML.ENERGY study puts it this way: as batch size increases, “energy per token drops at first, then plateaus as GPU utilization approaches saturation.” In the team’s earlier benchmark paper, the energy per answer of a distilled 8B reasoning model on one H100 fell to about 29 per cent of its value at a maximum batch size of 4 when the maximum was raised to 64.

Utilisation over the day works the same way. A server that waits for requests still draws idle power, and that energy belongs to the tokens it produces in the busy hours. A pilot with a handful of users therefore reports a much higher figure per token than the same server under steady load, so compare figures only at the same load.

Answer length changes the energy per request more than the per-token figure does. For Qwen3 32B on one B200, each task at its largest batch size, ML.ENERGY measured 95 J per response for text conversation and 2,192 J for problem solving, whose answers averaged 7,035 tokens against 627. Plan energy per answer for reasoning models, not only per token.

Measuring energy per token on your own server

A figure for a budget has to come from your model, prompts and concurrency; DCGM and vLLM provide the two counters.

  1. With the model loaded and no requests, note each GPU’s idle board power from nvidia-smi -q -d POWER, whose Average Power Draw NVIDIA defines as “The average power draw for the entire board for the last second, in watts.”
  2. Before the run, read DCGM field 156, DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION, defined as “Total energy consumption for the GPU in mJ since the driver was last reloaded”, for every GPU; dcgmi dmon -e 155,156 shows it beside board power in watts.
  3. At the same moment, read vLLM’s counters vllm:generation_tokens_total and vllm:prompt_tokens_total from its /metrics endpoint.
  4. Run a load with the concurrency, prompt length and answer length of your workload for a fixed period, then read both counters again.
  5. Divide the energy difference, summed over all GPUs and converted from millijoules to joules, by the generated tokens. Report prompt tokens beside it, since long RAG prompts add prefill energy that a figure per output token hides.
  6. Repeat at two or three concurrency levels and at each power limit you consider, and keep tokens per second per user beside every result.

Board energy leaves out the processors, memory, fans and power-supply losses. For the whole server, read the input power from the BMC or a metered PDU over the same window and divide by the same token count. The other DCGM fields worth exporting, and the field names DCGM 4.6 changed, are in our guide to GPU server monitoring with DCGM.

We build AI servers to order and check the rack, power and airflow before we quote. Send us the model, your daily token volume and the power budget of the rack position through the form below.

Power limits as an efficiency setting

On a 600 W card, nvidia-smi -i 0 -pl 450 sets a lower limit on the first GPU. NVIDIA’s documentation says the value “needs to be between Min and Max Power Limit as reported by nvidia-smi” and that the command “Requires root.” A lower limit reduces peak draw and heat per card, and throughput can fall by an amount that depends on the model, the precision and the batch size. Whether energy per token falls with it is a question for step 6 above, at the default limit and one or two lower ones.

After a reboot or driver reload, check that the limit still applies. DCGM reports it in DCGM_FI_DEV_BOARD_POWER_LIMIT_ENFORCED_WATTS, the “Effective power limit that the driver enforces after taking into account all limiters.” A capped card also lowers the power budget of the rack position, which our guide to GPU rack power and cooling sizes in amps and phases.

From joules per token to kWh per day

To convert, multiply joules per token by one million and divide by 3,600 for Wh per million tokens. The offline ceiling of 0.150 J per token for eight H200 NVL comes to 41.7 Wh per million tokens of GPU board energy. The RTX PRO 6000 result comes to 44.6 Wh.

A daily figure is simpler to bound. Eight cards at 600 W use at most 115.2 kWh of board energy in 24 hours, whatever they produce, and idle hours count towards that total. The wall power of the server adds processors, fans and power-supply losses, and a data centre adds its cooling and distribution on top. All of that energy ends up as heat in the room, which our article on a GPU server in an office or small server room converts to BTU/h.

At EU level, Commission Delegated Regulation (EU) 2024/1364 requires “operators of data centres with an installed information technology power demand of at least 500 kW” to report to a European database every year, directly or through a national reporting scheme, including the total energy consumption of the data centre and of its IT equipment in kWh. Whether a given site falls under it is a legal assessment for the company’s legal department.

What we supply

We supply the H200 NVL, the RTX PRO 6000 Server Edition and the L40S on their own for servers you already run, and in AI servers built to order, with manufacturer warranty on one EU contract and invoice. Each server is assembled and burn-in tested, and we check the rack, power and airflow before we quote. Operating system, drivers, CUDA and a container runtime are installed on request, and the driver includes nvidia-smi for reading power and setting limits. The full card line-up with each card’s board power is on our professional NVIDIA GPU page.

FAQ

How much energy does a GPU use per token in LLM inference?
It depends on batch size and utilisation as well as on the card. MLPerf Inference v6.0 publishes no power for these cards, but at their 600 W rating the offline Llama 2 70B results give, by our calculation, a ceiling of 0.150 J per token for eight H200 NVL and 0.160 J for eight RTX PRO 6000. ML.ENERGY measured 0.151 to 0.312 J per token for Qwen3 32B on one B200, depending on batch size and task.
How do I calculate energy per token for LLM inference?
Divide the average power drawn in watts by the tokens produced per second, which gives joules per token. On a running server, read DCGM’s energy counter in millijoules and vLLM’s generated-token counter before and after a load run, and divide the energy difference by the token difference. Repeat at the concurrency and power limit you plan to use.
How much electricity does an AI server use for LLM inference?
The cards set most of the load: eight 600 W cards use at most 4.8 kW, or 115.2 kWh of board energy in 24 hours, before processors, fans and power-supply losses. The server draws less when it is idle or lightly loaded, but idle power still counts towards the daily total. For an exact figure, read the input power from the BMC or a metered PDU.
Is the H200 NVL more energy efficient than the RTX PRO 6000?
In the MLPerf Inference v6.0 offline Llama 2 70B results, the two give similar ceilings by our calculation at 600 W per card, 0.150 J per token for the H200 NVL in FP8 and 0.160 J for the RTX PRO 6000 in FP4. Under the interactive latency limits, the H200 NVL system produced more tokens, giving a ceiling of 0.294 J against 0.769 J. These are bounds from rated power, not measured energy.
How many tokens per watt does a GPU deliver?
Tokens per second per watt equals tokens per joule, the inverse of joules per token. The offline MLPerf ceiling of 0.150 J per token for eight H200 NVL means at least 6.6 tokens per joule of GPU board energy for Llama 2 70B in FP8. Your own figure depends on batch size, utilisation and precision.
Does lowering the GPU power limit save energy?
It lowers peak draw and heat per card, and throughput can fall by an amount that depends on the model, precision and batch size. Whether energy per token falls as well has to be measured on your workload at the default and at lower limits. Set the limit with nvidia-smi -pl, within the Min and Max Power Limit the card reports.

Send us the model and its precision, your expected daily token volume, the peak number of concurrent requests and the power available at the rack position. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna