BLOG · COMPARISON ·

GPU for CFD simulation: H200 NVL or RTX PRO 6000 for double- and single-precision solvers

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • The solver’s precision decides the card; NVIDIA rates the H200 NVL at 30 TFLOPS of FP64, while the RTX PRO 6000 runs FP64 at 1/64 of its FP32 rate, about 1.9 TFLOPS on the Server Edition by our arithmetic
  • For single-precision and mixed-precision solvers the RTX PRO 6000 Server Edition offers 120 TFLOPS of FP32 and 96 GB per card; the H200 NVL offers 60 TFLOPS of FP32, 141 GB and 4.8 TB/s of memory bandwidth
  • Ansys’s Fluent GPU hardware guide plans 1.2 GB per million hexahedral cells in single precision with the segregated solver and 3.6 GB in double precision with the coupled solver, so one 96 GB card holds about 80 or 26 million cells
  • STAR-CCM+ 2602 runs VOF and mixture multiphase on its GPU-native solver; HP’s paper of January 2025 used the mixed-precision version of release 2310, and Siemens wrote in 2024 that one Power Session Plus licence covers unlimited CPUs or GPUs
  • Fluent counts GPU licences by streaming multiprocessors, with 40 SMs included in CFD Enterprise, and Ansys asks for system RAM of at least the total GPU memory

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Which GPU for CFD: H200 NVL or RTX PRO 6000

Choose a GPU for CFD by the precision your solver runs in. A case that has to run in double precision belongs on the H200 NVL, which NVIDIA rates at 30 TFLOPS of FP64. The RTX PRO 6000 runs FP64 at 1/64 of its FP32 rate, about 1.9 TFLOPS on the Server Edition by our arithmetic. It suits solvers that run in single or mixed precision, where its 120 TFLOPS of FP32 and 96 GB per card count. After precision, card memory sets the mesh size per GPU and memory bandwidth sets much of the speed.

GPUFP64FP32MEMORYBANDWIDTH
NVIDIA H200 NVL30 TFLOPS, 60 on Tensor Cores60 TFLOPS141 GB HBM3e4.8 TB/s
RTX PRO 6000 Server Editionabout 1.9 TFLOPS120 TFLOPS96 GB GDDR7, ECC1,597 GB/s
RTX PRO 6000 Workstationabout 2.0 TFLOPS125 TFLOPS96 GB GDDR7, ECC1,792 GB/s

NVIDIA H200, RTX PRO 6000 Server Edition and RTX PRO 6000 Workstation Edition product pages, read on 10 October 2026; the 1/64 FP64 rate from NVIDIA’s RTX Blackwell PRO architecture whitepaper v1.1. RTX PRO 6000 FP64 values are our arithmetic, since NVIDIA’s product pages list none.

Both are PCIe Gen5 cards rated up to 600 W, so the same server class takes either. The H200 NVL also joins 2 or 4 cards over an NVLink bridge at 900 GB/s per GPU, and the RTX PRO 6000 has no NVLink. Our comparison of the two as A100 replacements covers the power cable, driver and MIG differences, which apply to a simulation server in the same way.

Double precision on the RTX PRO 6000 and the H200 NVL

NVIDIA’s RTX Blackwell PRO architecture whitepaper (version 1.1) says the GB202 chip has two FP64 cores per SM and that “The FP64 TFLOP rate is 1/64th the TFLOP rate of FP32 operations.” NVIDIA gives the reason too, that the few FP64 cores “are included to ensure any programs with FP64 code operate correctly”. A double-precision code therefore runs on the RTX PRO 6000, but at about 1.9 TFLOPS on the Server Edition. The H200 NVL is a Hopper card with native FP64 cores, rated at 30 TFLOPS and at 60 TFLOPS on its Tensor Cores.

Whether a case needs double precision is a solver and validation question for the simulation team. Ansys’s Fluent GPU solver hardware buying guide, last modified on 9 April 2026, calls it “The most important question to consider”. Where double precision is needed, the guide finds the high-end server products “more appealing”. In single precision, it says, the Fluent GPU solver “primarily utilizes the GPU’s FP32 cores” and does not use the FP64 cores even where a card has them. The same guide reports that the Fluent GPU solver generally “converges better in single precision than the Fluent CPU solver.”

Fluent can still run its double-precision mode on cards without native FP64 cores. The guide says that on such cards “double-precision runs at roughly half the speed”, and that “When native FP64 cores are present, as in the H100, this emulation is unnecessary”. The H200 NVL uses the same Hopper architecture. The guide counts cards “up to the NVIDIA RTX 6000 Ada” among those without FP64 cores, but does not say which path Fluent takes on the RTX PRO 6000, with its two FP64 cores per SM. Ansys’s blog of 22 July 2026 adds a third mode, gpu_hybrid_precision, which it says “can provide nearly double-precision accuracy with lower RAM demand and reduced computation time.” For a team that runs Fluent in double precision today, that mode is the first thing to validate before choosing the RTX PRO 6000.

Fluent GPU solver: memory per million cells

The Fluent GPU solver needs the whole case in GPU memory, and Ansys publishes the memory it plans per million cells. The buying guide assumes “1.2GB GPU RAM per million cells for a typical Single precision Steady State SIMPLE model” and states that actual requirements are case specific. The Fluent user’s guide for release 2024 R2 gives about 1 GB per million hexahedral cells in single precision with turbulence, 50 per cent more GPU memory for double precision and 20 to 40 per cent more for a polyhedral mesh.

FLUENT CASEGB PER MILLION CELLSCELLS ON 96 GBCELLS ON 141 GB
Hex, single, segregated1.2about 80 millionabout 117 million
Hex, single, coupled2.2about 43 millionabout 64 million
Hex, double, segregated1.9about 50 millionabout 74 million
Hex, double, coupled3.6about 26 millionabout 39 million
Poly, single, segregated1.8about 53 millionabout 78 million
Poly, double, coupled5.6about 17 millionabout 25 million

GB per million cells from Ansys’s Fluent GPU solver hardware buying guide (Ansys Innovation Space, last modified 9 April 2026). Cell counts are our division of the nominal card memory, rounded down, and are upper limits before process overhead.

Each Fluent compute process also takes about 95 MB of GPU memory and 210 MB of system memory, according to the guide. Several GPUs, in one server or across several, pool their memory for one case: “The sum of the memory of all GPU cards must be able to hold the model and the computation overhead.” The guide also warns that multi-GPU performance is limited by the slowest card and the smallest memory in the set, and recommends matching cards. A 60 million cell hexahedral case in double precision with the coupled solver therefore needs about 216 GB across matched cards, three RTX PRO 6000 or two H200 NVL by our arithmetic, with headroom still to add.

We build GPU servers with RTX PRO 6000 Server Edition or H200 NVL cards around the case size and precision you run. Send us your largest case, its cell count, mesh type and precision, and the configuration and quote follow within one business day.

Fluent licences and the SM count

Ansys licenses the Fluent GPU solver by streaming multiprocessors, not by card. The buying guide states that “40 SMs/CUs are included with the CFD Enterprise license”, that additional SMs need Ansys HPC licences, and that a CFD HPC Ultimate licence gives access to unlimited SMs. NVIDIA’s whitepaper lists 188 SMs for the RTX PRO 6000 Workstation Edition, and the guide’s table gives 188 for the RTX PRO 6000 Blackwell. NVIDIA’s H200 product page does not list an SM count for the H200 NVL, so ask your Ansys channel for the licence count per card before you compare configurations.

Siemens STAR-CCM+ GPU-native solver

Siemens describes a GPU-native solver for STAR-CCM+. Its blog of 18 February 2022 states that “NVIDIA GPUs with Volta architecture or newer are supported” and that cards with HBM2 “are preferred over Graphics Double Data Rate 6 (GDDR6)”. At that point the GPU solver covered constant-density flows with the segregated solver, and a Siemens post of 21 June 2023 announced the coupled solver on the GPU. Siemens’s release post for version 2602, dated 24 February 2026, adds GPU-native Volume of Fluid and Mixture Multiphase solvers, better multi-GPU scaling and wider AMD GPU support. It also states that “the same solver is used for both CPU and GPU executions”. Volupe, a Siemens solutions partner, wrote on 27 February 2026 that 2602 “officially supports the Nvidia Blackwell GPU series”, the generation of the RTX PRO 6000.

The Siemens posts we read do not state which floating-point precision the GPU-native solver uses. HP’s technical paper of January 2025 ran the “mixed precision version” of STAR-CCM+ 2310 on up to four RTX 6000 Ada cards with 48 GB each. In that test a 57 million cell case needed at least two cards and a 106 million cell case at least three. Confirm with Siemens or your reseller whether your physics models run on the GPU in the precision you validate with.

On licensing, Siemens wrote on 20 June 2024 that “a single Power Session Plus licence allows you to run on unlimited CPUs or GPUs”, without variation by GPU type. Volupe states that GPU use needs a Power Session Plus or Power-on-Demand licence.

OpenFOAM and in-house codes

The OpenFOAM High Performance Computing Technical Committee, on its page last edited on 26 November 2025, lists “GPU enabling of OpenFOAM” as a current priority. It names GPU support for the PETSc4FOAM library as a planned activity and lists AmgX GPU solver development among its tasks. We found no OpenFOAM document that names a GPU class for it, so the card follows from the GPU library you use and the precision it runs in. For in-house or research codes, ask the developers whether the kernels use FP64, because that alone separates the two cards by a factor of about 16.

SOLVERGPU SUPPORTPRECISION ON GPUCARD FIT
Ansys Fluent GPU solverNVIDIA, and AMD on Linux from 2025 R1; the guide lists A100, H100, L40 and RTX A6000single, double, hybridRTX PRO 6000 for single or hybrid; H200 NVL for double at full rate
Siemens STAR-CCM+NVIDIA Volta or newer, Blackwell from 2602; AMDmixed in HP’s test; not stated by SiemensRTX PRO 6000 for mixed; H200 NVL for HBM bandwidth
OpenFOAMGPU enabling listed as a priority by the HPC committeenot statedfollows the GPU library and its precision

Ansys Fluent GPU buying guide (9 April 2026) and Ansys blog (22 July 2026); Siemens blogs (2022, 2024, February 2026), Volupe (February 2026), HP paper C09079055 (January 2025); OpenFOAM HPC Technical Committee page. Card fit is our reading of these sources.

Memory bandwidth and multi-GPU scaling

Ansys’s guide states that memory bandwidth “is important to transport the data to the SMs/CUs” and that “In most cases you benefit from a higher memory bandwidth.” The H200 NVL reads its memory at 4.8 TB/s, three times the 1,597 GB/s of the RTX PRO 6000 Server Edition. In single precision the RTX PRO 6000 has twice the FP32 rate, so which card finishes a single-precision case first depends on the solver, and we found no published Fluent or STAR-CCM+ result that sets the two cards against each other.

Ansys’s blog of 22 July 2026 measured Fluent 2026 R1 on four RTX PRO 6000 Blackwell Server cards against CPU baselines. On the DrivAer 50 million cell case, the four cards delivered “approximately a ~4.6X speedup” over a baseline of about 516 CPU cores. Ansys places these configurations at smaller jobs and “more cost-sensitive server deployments”, and the blog did not test the H200 NVL. NVIDIA’s HPC claims for the H200 NVL, and the test conditions we could not find for them, are covered in our H200 NVL vs H100 NVL comparison.

We found no Ansys statement that the Fluent GPU solver uses the H200 NVL’s NVLink bridges, so ask the solver vendor before you plan bridges for a CFD-only server.

GPU simulation server: configuration for CFD

Ansys recommends “at least as much system random-access memory (RAM) as total GPU memory”, and its guide expects more for polyhedral meshes. Four RTX PRO 6000 cards therefore call for at least 384 GB of system memory, and four H200 NVL cards for at least 564 GB. At up to 600 W per card, four cards draw 2.4 kW before processors and fans, and our article on how many GPUs fit in one server explains the PCIe lanes, power feeds and airflow that decide the card count.

Mixed use also points to a card. NVIDIA presents the RTX PRO 6000 Server Edition for “data analytics, engineering simulation, and visual computing”, so the same server can also render results. If it also hosts language models, our RTX PRO 6000 vs H200 NVL inference comparison covers that side.

We check the rack, power and airflow before we quote. Write to us with the solver, its version, your largest case and the rack position’s power feed, and we reply within one business day.

What we supply

We supply the H200 NVL and the RTX PRO 6000 in its Server, Workstation and Max-Q editions as professional NVIDIA GPUs, and build GPU servers to order around them, assembled and burn-in tested, with manufacturer warranty on every component. A double-precision solver server gets H200 NVL cards, with NVLink bridges where the workload uses them, and a single-precision or mixed-precision server gets RTX PRO 6000 Server Edition cards. We also build on your own chassis or with parts you already own, after a compatibility check of the platform, power and cooling. Operating system, drivers and CUDA are installed on request, and everything comes on one EU contract and invoice with delivery anywhere in the EU.

FAQ

Which GPU is better for CFD, the H200 NVL or the RTX PRO 6000?
It depends on the precision your solver runs in. For double-precision work the H200 NVL, at 30 TFLOPS of FP64, is the card to choose, while the RTX PRO 6000 runs FP64 at 1/64 of its FP32 rate, about 1.9 TFLOPS on the Server Edition. For single-precision or mixed-precision solvers the RTX PRO 6000 offers 120 TFLOPS of FP32 and 96 GB per card.
Does the Ansys Fluent GPU solver need FP64?
Not in single precision, where Ansys says it primarily uses the GPU’s FP32 cores and does not use FP64 cores even where a card has them. Fluent can also run in double precision on cards without native FP64 cores, at roughly half the speed according to Ansys’s buying guide, which does not say which path it takes on the RTX PRO 6000. Ansys’s hybrid precision mode, which it says gives nearly double-precision accuracy with less memory, is a third option.
How much GPU memory does Fluent need per million cells?
Ansys’s buying guide plans about 1.2 GB per million hexahedral cells in single precision with the segregated solver and 3.6 GB in double precision with the coupled solver. Polyhedral meshes need more, up to 5.6 GB per million cells in double precision with the coupled solver. One 96 GB RTX PRO 6000 therefore holds about 80 or 26 million hexahedral cells, before process overhead.
Does STAR-CCM+ support GPUs?
Yes, STAR-CCM+ has a GPU-native solver for NVIDIA GPUs from the Volta architecture on and for AMD GPUs, and version 2602 added Volume of Fluid and Mixture Multiphase on the GPU. HP’s paper of January 2025 used the mixed-precision version of release 2310 on RTX 6000 Ada cards. Siemens wrote in June 2024 that one Power Session Plus licence covers unlimited CPUs or GPUs.
Is the RTX PRO 6000 suitable for FP64 simulation?
Not for codes that compute mostly in FP64. NVIDIA’s whitepaper says its two FP64 cores per SM are there so that FP64 programs run correctly, at 1/64 of the FP32 rate. Double-precision codes belong on the H200 NVL, which NVIDIA rates at 30 TFLOPS of FP64.
How much system RAM does a GPU server for CFD need?
Ansys recommends at least as much system memory as the total GPU memory in the server, and more for polyhedral meshes. A server with four RTX PRO 6000 cards therefore needs at least 384 GB, and one with four H200 NVL cards at least 564 GB. Each Fluent compute process also takes about 210 MB of system memory.

Send us the solver and its version, the precision you run, the cell count and mesh type of your largest case, and the rack position’s power feed. We reply within one business day with a configuration and a quote in writing, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna