GPU for CFD simulation: H200 NVL or RTX PRO 6000 for double- and single-precision solvers
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- The solver’s precision decides the card; NVIDIA rates the H200 NVL at 30 TFLOPS of FP64, while the RTX PRO 6000 runs FP64 at 1/64 of its FP32 rate, about 1.9 TFLOPS on the Server Edition by our arithmetic
- For single-precision and mixed-precision solvers the RTX PRO 6000 Server Edition offers 120 TFLOPS of FP32 and 96 GB per card; the H200 NVL offers 60 TFLOPS of FP32, 141 GB and 4.8 TB/s of memory bandwidth
- Ansys’s Fluent GPU hardware guide plans 1.2 GB per million hexahedral cells in single precision with the segregated solver and 3.6 GB in double precision with the coupled solver, so one 96 GB card holds about 80 or 26 million cells
- STAR-CCM+ 2602 runs VOF and mixture multiphase on its GPU-native solver; HP’s paper of January 2025 used the mixed-precision version of release 2310, and Siemens wrote in 2024 that one Power Session Plus licence covers unlimited CPUs or GPUs
- Fluent counts GPU licences by streaming multiprocessors, with 40 SMs included in CFD Enterprise, and Ansys asks for system RAM of at least the total GPU memory
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Which GPU for CFD: H200 NVL or RTX PRO 6000
Choose a GPU for CFD by the precision your solver runs in. A case that has to run in double precision belongs on the H200 NVL, which NVIDIA rates at 30 TFLOPS of FP64. The RTX PRO 6000 runs FP64 at 1/64 of its FP32 rate, about 1.9 TFLOPS on the Server Edition by our arithmetic. It suits solvers that run in single or mixed precision, where its 120 TFLOPS of FP32 and 96 GB per card count. After precision, card memory sets the mesh size per GPU and memory bandwidth sets much of the speed.
| GPU | FP64 | FP32 | MEMORY | BANDWIDTH |
|---|---|---|---|---|
| NVIDIA H200 NVL | 30 TFLOPS, 60 on Tensor Cores | 60 TFLOPS | 141 GB HBM3e | 4.8 TB/s |
| RTX PRO 6000 Server Edition | about 1.9 TFLOPS | 120 TFLOPS | 96 GB GDDR7, ECC | 1,597 GB/s |
| RTX PRO 6000 Workstation | about 2.0 TFLOPS | 125 TFLOPS | 96 GB GDDR7, ECC | 1,792 GB/s |
NVIDIA H200, RTX PRO 6000 Server Edition and RTX PRO 6000 Workstation Edition product pages, read on 10 October 2026; the 1/64 FP64 rate from NVIDIA’s RTX Blackwell PRO architecture whitepaper v1.1. RTX PRO 6000 FP64 values are our arithmetic, since NVIDIA’s product pages list none.
Both are PCIe Gen5 cards rated up to 600 W, so the same server class takes either. The H200 NVL also joins 2 or 4 cards over an NVLink bridge at 900 GB/s per GPU, and the RTX PRO 6000 has no NVLink. Our comparison of the two as A100 replacements covers the power cable, driver and MIG differences, which apply to a simulation server in the same way.
Double precision on the RTX PRO 6000 and the H200 NVL
NVIDIA’s RTX Blackwell PRO architecture whitepaper (version 1.1) says the GB202 chip has two FP64 cores per SM and that “The FP64 TFLOP rate is 1/64th the TFLOP rate of FP32 operations.” NVIDIA gives the reason too, that the few FP64 cores “are included to ensure any programs with FP64 code operate correctly”. A double-precision code therefore runs on the RTX PRO 6000, but at about 1.9 TFLOPS on the Server Edition. The H200 NVL is a Hopper card with native FP64 cores, rated at 30 TFLOPS and at 60 TFLOPS on its Tensor Cores.
Whether a case needs double precision is a solver and validation question for the simulation team. Ansys’s Fluent GPU solver hardware buying guide, last modified on 9 April 2026, calls it “The most important question to consider”. Where double precision is needed, the guide finds the high-end server products “more appealing”. In single precision, it says, the Fluent GPU solver “primarily utilizes the GPU’s FP32 cores” and does not use the FP64 cores even where a card has them. The same guide reports that the Fluent GPU solver generally “converges better in single precision than the Fluent CPU solver.”
Fluent can still run its double-precision mode on cards without native FP64 cores. The guide says that on such cards “double-precision runs at roughly half the speed”, and that “When native FP64 cores are present, as in the H100, this emulation is unnecessary”. The H200 NVL uses the same Hopper architecture. The guide counts cards “up to the NVIDIA RTX 6000 Ada” among those without FP64 cores, but does not say which path Fluent takes on the RTX PRO 6000, with its two FP64 cores per SM. Ansys’s blog of 22 July 2026 adds a third mode, gpu_hybrid_precision, which it says “can provide nearly double-precision accuracy with lower RAM demand and reduced computation time.” For a team that runs Fluent in double precision today, that mode is the first thing to validate before choosing the RTX PRO 6000.
Fluent GPU solver: memory per million cells
The Fluent GPU solver needs the whole case in GPU memory, and Ansys publishes the memory it plans per million cells. The buying guide assumes “1.2GB GPU RAM per million cells for a typical Single precision Steady State SIMPLE model” and states that actual requirements are case specific. The Fluent user’s guide for release 2024 R2 gives about 1 GB per million hexahedral cells in single precision with turbulence, 50 per cent more GPU memory for double precision and 20 to 40 per cent more for a polyhedral mesh.
| FLUENT CASE | GB PER MILLION CELLS | CELLS ON 96 GB | CELLS ON 141 GB |
|---|---|---|---|
| Hex, single, segregated | 1.2 | about 80 million | about 117 million |
| Hex, single, coupled | 2.2 | about 43 million | about 64 million |
| Hex, double, segregated | 1.9 | about 50 million | about 74 million |
| Hex, double, coupled | 3.6 | about 26 million | about 39 million |
| Poly, single, segregated | 1.8 | about 53 million | about 78 million |
| Poly, double, coupled | 5.6 | about 17 million | about 25 million |
GB per million cells from Ansys’s Fluent GPU solver hardware buying guide (Ansys Innovation Space, last modified 9 April 2026). Cell counts are our division of the nominal card memory, rounded down, and are upper limits before process overhead.
Each Fluent compute process also takes about 95 MB of GPU memory and 210 MB of system memory, according to the guide. Several GPUs, in one server or across several, pool their memory for one case: “The sum of the memory of all GPU cards must be able to hold the model and the computation overhead.” The guide also warns that multi-GPU performance is limited by the slowest card and the smallest memory in the set, and recommends matching cards. A 60 million cell hexahedral case in double precision with the coupled solver therefore needs about 216 GB across matched cards, three RTX PRO 6000 or two H200 NVL by our arithmetic, with headroom still to add.
We build GPU servers with RTX PRO 6000 Server Edition or H200 NVL cards around the case size and precision you run. Send us your largest case, its cell count, mesh type and precision, and the configuration and quote follow within one business day.
Fluent licences and the SM count
Ansys licenses the Fluent GPU solver by streaming multiprocessors, not by card. The buying guide states that “40 SMs/CUs are included with the CFD Enterprise license”, that additional SMs need Ansys HPC licences, and that a CFD HPC Ultimate licence gives access to unlimited SMs. NVIDIA’s whitepaper lists 188 SMs for the RTX PRO 6000 Workstation Edition, and the guide’s table gives 188 for the RTX PRO 6000 Blackwell. NVIDIA’s H200 product page does not list an SM count for the H200 NVL, so ask your Ansys channel for the licence count per card before you compare configurations.
Siemens STAR-CCM+ GPU-native solver
Siemens describes a GPU-native solver for STAR-CCM+. Its blog of 18 February 2022 states that “NVIDIA GPUs with Volta architecture or newer are supported” and that cards with HBM2 “are preferred over Graphics Double Data Rate 6 (GDDR6)”. At that point the GPU solver covered constant-density flows with the segregated solver, and a Siemens post of 21 June 2023 announced the coupled solver on the GPU. Siemens’s release post for version 2602, dated 24 February 2026, adds GPU-native Volume of Fluid and Mixture Multiphase solvers, better multi-GPU scaling and wider AMD GPU support. It also states that “the same solver is used for both CPU and GPU executions”. Volupe, a Siemens solutions partner, wrote on 27 February 2026 that 2602 “officially supports the Nvidia Blackwell GPU series”, the generation of the RTX PRO 6000.
The Siemens posts we read do not state which floating-point precision the GPU-native solver uses. HP’s technical paper of January 2025 ran the “mixed precision version” of STAR-CCM+ 2310 on up to four RTX 6000 Ada cards with 48 GB each. In that test a 57 million cell case needed at least two cards and a 106 million cell case at least three. Confirm with Siemens or your reseller whether your physics models run on the GPU in the precision you validate with.
On licensing, Siemens wrote on 20 June 2024 that “a single Power Session Plus licence allows you to run on unlimited CPUs or GPUs”, without variation by GPU type. Volupe states that GPU use needs a Power Session Plus or Power-on-Demand licence.
OpenFOAM and in-house codes
The OpenFOAM High Performance Computing Technical Committee, on its page last edited on 26 November 2025, lists “GPU enabling of OpenFOAM” as a current priority. It names GPU support for the PETSc4FOAM library as a planned activity and lists AmgX GPU solver development among its tasks. We found no OpenFOAM document that names a GPU class for it, so the card follows from the GPU library you use and the precision it runs in. For in-house or research codes, ask the developers whether the kernels use FP64, because that alone separates the two cards by a factor of about 16.
| SOLVER | GPU SUPPORT | PRECISION ON GPU | CARD FIT |
|---|---|---|---|
| Ansys Fluent GPU solver | NVIDIA, and AMD on Linux from 2025 R1; the guide lists A100, H100, L40 and RTX A6000 | single, double, hybrid | RTX PRO 6000 for single or hybrid; H200 NVL for double at full rate |
| Siemens STAR-CCM+ | NVIDIA Volta or newer, Blackwell from 2602; AMD | mixed in HP’s test; not stated by Siemens | RTX PRO 6000 for mixed; H200 NVL for HBM bandwidth |
| OpenFOAM | GPU enabling listed as a priority by the HPC committee | not stated | follows the GPU library and its precision |
Ansys Fluent GPU buying guide (9 April 2026) and Ansys blog (22 July 2026); Siemens blogs (2022, 2024, February 2026), Volupe (February 2026), HP paper C09079055 (January 2025); OpenFOAM HPC Technical Committee page. Card fit is our reading of these sources.
Memory bandwidth and multi-GPU scaling
Ansys’s guide states that memory bandwidth “is important to transport the data to the SMs/CUs” and that “In most cases you benefit from a higher memory bandwidth.” The H200 NVL reads its memory at 4.8 TB/s, three times the 1,597 GB/s of the RTX PRO 6000 Server Edition. In single precision the RTX PRO 6000 has twice the FP32 rate, so which card finishes a single-precision case first depends on the solver, and we found no published Fluent or STAR-CCM+ result that sets the two cards against each other.
Ansys’s blog of 22 July 2026 measured Fluent 2026 R1 on four RTX PRO 6000 Blackwell Server cards against CPU baselines. On the DrivAer 50 million cell case, the four cards delivered “approximately a ~4.6X speedup” over a baseline of about 516 CPU cores. Ansys places these configurations at smaller jobs and “more cost-sensitive server deployments”, and the blog did not test the H200 NVL. NVIDIA’s HPC claims for the H200 NVL, and the test conditions we could not find for them, are covered in our H200 NVL vs H100 NVL comparison.
We found no Ansys statement that the Fluent GPU solver uses the H200 NVL’s NVLink bridges, so ask the solver vendor before you plan bridges for a CFD-only server.
GPU simulation server: configuration for CFD
Ansys recommends “at least as much system random-access memory (RAM) as total GPU memory”, and its guide expects more for polyhedral meshes. Four RTX PRO 6000 cards therefore call for at least 384 GB of system memory, and four H200 NVL cards for at least 564 GB. At up to 600 W per card, four cards draw 2.4 kW before processors and fans, and our article on how many GPUs fit in one server explains the PCIe lanes, power feeds and airflow that decide the card count.
Mixed use also points to a card. NVIDIA presents the RTX PRO 6000 Server Edition for “data analytics, engineering simulation, and visual computing”, so the same server can also render results. If it also hosts language models, our RTX PRO 6000 vs H200 NVL inference comparison covers that side.
We check the rack, power and airflow before we quote. Write to us with the solver, its version, your largest case and the rack position’s power feed, and we reply within one business day.
What we supply
We supply the H200 NVL and the RTX PRO 6000 in its Server, Workstation and Max-Q editions as professional NVIDIA GPUs, and build GPU servers to order around them, assembled and burn-in tested, with manufacturer warranty on every component. A double-precision solver server gets H200 NVL cards, with NVLink bridges where the workload uses them, and a single-precision or mixed-precision server gets RTX PRO 6000 Server Edition cards. We also build on your own chassis or with parts you already own, after a compatibility check of the platform, power and cooling. Operating system, drivers and CUDA are installed on request, and everything comes on one EU contract and invoice with delivery anywhere in the EU.
FAQ
Which GPU is better for CFD, the H200 NVL or the RTX PRO 6000?
Does the Ansys Fluent GPU solver need FP64?
How much GPU memory does Fluent need per million cells?
Does STAR-CCM+ support GPUs?
Is the RTX PRO 6000 suitable for FP64 simulation?
How much system RAM does a GPU server for CFD need?
Send us the solver and its version, the precision you run, the cell count and mesh type of your largest case, and the rack position’s power feed. We reply within one business day with a configuration and a quote in writing, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day