Medical imaging AI server: GPU memory for MONAI, 3D segmentation and radiology models
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A CT or MRI study is a 3D volume of hundreds of slices, so medical imaging models train on 3D patches and need far more GPU memory per sample than 2D image models; inference with sliding windows needs less
- MONAI’s whole-body CT segmentation bundle asks for 48 GB of GPU memory for training; on a 512 × 512 × 397 CT its 1.5 mm model used 28.73 GB for inference and its 3.0 mm model 5.89 GB
- NVIDIA’s synthetic CT models (NV-Generate-CT, built on MAISI) peak at 15.0 GB for a 256 × 256 × 128 volume and 49.7 GB for 512 × 512 × 768, measured on an A100 80 GB
- The NIM support matrices for VISTA-3D and MAISI list the A100, H100, L40S and RTX 6000 Ada in FP32; of the cards we supply, the L40S and the RTX 6000 Ada are on them and no Blackwell card is, as of October 2026
- Software that its manufacturer intends for diagnosis can be a medical device under Article 2(1) of the EU MDR, and data concerning health is a special category under Article 9 of the GDPR; both are assessments for the hospital’s legal and regulatory functions
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
GPU requirements for medical imaging AI
A medical imaging AI server is sized by the 3D volume it processes more than by the number of users. A CT or MRI study holds hundreds of slices, so segmentation models train on 3D patches, and each sample fills far more GPU memory than a 2D photograph does. MONAI’s whole-body CT segmentation bundle asks for 48 GB of GPU memory for training. Inference is lighter: on a CT of 512 × 512 × 397 voxels, the same bundle’s 3.0 mm model used 5.89 GB and its 1.5 mm model 28.73 GB. Training therefore points to cards of 48 GB and more, such as the RTX PRO 6000 with 96 GB or the H200 NVL with 141 GB, while most inference fits a 24 GB L4 or a 48 GB L40S.
| WORKLOAD | MEMORY DRIVER | STATED MEMORY | CARD (OUR RANGE) |
|---|---|---|---|
| Training 3D segmentation | patch size, batch, channels | 48 GB whole-body CT; at least 32 GB Swin UNETR | RTX PRO 6000, H200 NVL |
| Segmentation inference | volume, output channels, stitching | 5.89 GB at 3.0 mm, 28.73 GB at 1.5 mm | L4 (3.0 mm), L40S, RTX PRO 5000 |
| Synthetic CT, inference | output volume size | 15.0 to 49.7 GB peak | L40S for smaller volumes, RTX PRO 6000 |
| Synthetic CT, training | latent size | 39 GB at 512³, 58 GB at 512 × 512 × 768 | RTX PRO 6000, H200 NVL |
| MONAI Deploy app | the packaged model | at least 8 GB video RAM | L4, RTX PRO 4000 |
MONAI Model Zoo READMEs (wholeBody_ct_segmentation, swin_
Why 3D volumes fill GPU memory
A CT of 512 × 512 × 397 voxels holds about 104 million values, 416 MB in FP32 before the network has computed anything, by our arithmetic. Networks do not take such a volume whole. The whole-body bundle and the Swin UNETR bundle both train on patches of 96 × 96 × 96 voxels, which is 884,736 voxels, more than three times a 512 × 512 slice. Every 3D convolution keeps its activations per voxel and per channel for the backward pass, so training memory grows with the patch volume, the number of channels and the batch size. The Swin UNETR bundle trains with mixed precision (AMP) and asks for at least 32 GB.
For inference, MONAI’s SlidingWindowInferer moves the window across the volume and runs sw_batch_size windows per forward pass. The output is stitched on the device of the input unless device is set otherwise, and MONAI’s docstring notes that with the output on the CPU “the gpu memory consumption is less”. The output accounts for about half of the 28.73 GB measured for the 1.5 mm model, by our arithmetic: 105 output channels over the resampled 287 × 287 × 397 volume take 13.7 GB in FP32. The bundle’s README puts a CT of 300 slices at about 27 GB and advises more GPU memory, the 3.0 mm model or CPU inference for larger scans.
System memory matters as well. MONAI’s CacheDataset keeps preprocessed volumes in RAM, and the whole-body bundle reports 83 GB of system RAM on one GPU for 1,000 volumes at a cache rate of 0.4, and 666 GB with eight GPUs. A training server for 3D models therefore needs system RAM sized for the cache as well as for the GPUs, and NVMe capacity for the datasets.
MONAI, VISTA-3D and NVIDIA’s open imaging models
NVIDIA’s Clara pages at nvidia.com/clara describe MONAI as “the open medical imaging AI framework”, PyTorch-based, with “50+ high-quality pretrained models”. The MONAI Model Zoo bundles above are published under the Apache 2.0 licence by the MONAI Consortium. NVIDIA’s own imaging models now sit in the NVIDIA-Medtech organisation on GitHub, which describes itself as “Open foundation models for physical AI and medical imaging”. The MAISI folder in MONAI’s tutorials has been marked deprecated since October 2025 and points to NV-Generate-CTMR.
VISTA-3D segments CT automatically or from point clicks. Its repository says it was trained on 11,454 volumes covering 127 types of human anatomical structures and various lesions, and NVIDIA offers it as a NIM, model version 0.5.7 in the support matrix as read on 10 October 2026. NV-Segment-CTMR starts from the NV-Segment-CT checkpoint, shares the architecture of VISTA3D-CT and was fine-tuned on over 30,000 CT and MRI scans for more than 300 classes. Outside NVIDIA, MedSAM2, posted on arXiv on 4 April 2025, fine-tunes the Segment Anything Model 2 on over 455,000 3D image-mask pairs and 76,000 video frames to segment 3D medical images and videos.
| MODEL | TASK | WEIGHTS LICENCE | MEMORY AS STATED |
|---|---|---|---|
| Whole-body CT (MONAI bundle) | 104 structures, SegResNet | Apache 2.0 | 48 GB training; 5.89 or 28.73 GB inference |
| VISTA-3D | CT, automatic and point-click | commercial friendly, per the repository | NIM lists 48 and 80 GB GPUs |
| NV-Segment-CTMR | CT and MRI, 345+ classes | Non-Commercial, per the repository | not stated |
| NV-Generate-CT | synthetic CT with 132-class masks | NVIDIA Open Model | 15.0 to 49.7 GB peak inference |
| NV-Reason-CXR-3B | chest X-ray reasoning (VLM) | NVIDIA OneWay Non-Commercial | 3B per the card, BF16; no figure |
MONAI Model Zoo, VISTA, NV-Segment-CTMR and NV-Generate-CTMR repositories on GitHub, NIM for VISTA-3D support matrix (no update date shown) and the NV-Reason-CXR-3B model card on Hugging Face, read on 10 October 2026.
Three of NVIDIA’s imaging models carry non-commercial weights as of October 2026: NV-Segment-CTMR (“Non-Commercial” in its repository), the MRI generator NV-Generate-MR (“NVIDIA Non-Commercial”) and NV-Reason-CXR-3B (“NVIDIA OneWay Non-Commercial License for academic research purposes”). Whether a planned use falls within such a licence is for the hospital’s legal department to assess. NV-Reason-CXR-3B is fine-tuned from Qwen2.5-VL-3B-Instruct and was tested on the A100, H100 and L40S, according to its model card, which gives 3B parameters while its file metadata shows 4B.
Training and inference: which GPUs
For training, the 48 GB that the whole-body bundle asks for fit both cards we supply for this work with room for larger patches or batches. The RTX PRO 6000 Server Edition has 96 GB of GDDR7 at 1,597 GB/s and draws up to 600 W; the H200 NVL has 141 GB of HBM3e at 4.8 TB/s and joins 2 or 4 cards over an NVLink bridge. Diffusion training at 512 × 512 × 768, at 58 GB peak, fits either. Our comparison of the H200 NVL and the RTX PRO 6000 for training covers tensor rates and the link between cards. Fine-tuning a vision-language model such as NV-Reason-CXR-3B follows the rules for language models, set out in our guide to GPU memory for LoRA and QLoRA fine-tuning.
For inference, a 48 GB L40S holds the 1.5 mm whole-body model at 28.73 GB, and a 24 GB L4 at 72 W holds the 3.0 mm model and meets the 8 GB that the MONAI Deploy App SDK asks for; our L4 and L40S comparison covers power and density per server. MIG splits an RTX PRO 6000 into four 24 GB or two 48 GB instances and an H200 NVL into up to seven, so several inference models can share one card with isolated memory.
The NIM support matrices for VISTA-3D and MAISI list the A100, H100, L40S and RTX 6000 Ada, all in FP32, and require CUDA compute capability 7.0 or higher. Of the cards we supply, the L40S and the RTX 6000 Ada, both with 48 GB, are on these lists as of October 2026. Both pages name Ampere and Hopper as the supported architectures, which takes in the H200 NVL as a Hopper card, but they do not name Blackwell, so NVIDIA gives no support statement for the RTX PRO cards; run the NIM on the target card before you plan on it. The MAISI page, last updated on 5 June 2026, marks image size limits for the 48 GB cards and gives a peak of 55 GB for a 512 × 512 × 768 volume.
We supply the L4, L40S, RTX 6000 Ada, RTX PRO 6000 and H200 NVL as cards or in AI servers built to order. Tell us which models you run and at what resolution, and whether you train or only run inference.
Connecting to PACS with MONAI Deploy
In radiology, studies reach the inference server from the PACS and the results go back to it. The MONAI Deploy App SDK, in its README’s words, offers “a framework and associated tools to design, develop and verify AI-driven applications” and packages an inference application “with a single command” into a MONAI Application Package. It has built-in operators to load DICOM data, and version 3.0.0 of 22 April 2025 extended its DICOM Segmentation operator to fill DICOM tags with information on the AI model. Version 4.0.0, published on PyPI on 13 June 2026, asks for Ubuntu 22.04 with glibc 2.35 or later, CUDA 13.0 or above and an NVIDIA GPU with at least 8 GB of video RAM, and depends on NVIDIA’s Holoscan SDK for CUDA 13. The project’s GitHub releases page still shows 3.0.0, based on Holoscan SDK v3, as the latest release, so check which version a packaged application was built with.
The MONAI Deploy Informatics Gateway, by its repository’s description, “facilitates integration with DICOM compliant systems, enables ingestion of imaging data” and pushes results to PACS systems, using DICOM and FHIR. For building training sets, MONAI Label is an “intelligent open source image labeling and learning tool” that works with 3D Slicer, OHIF and other viewers and connects to a PACS over DICOMweb. A server that receives studies from the PACS sits in the clinical network, so its BMC, drivers and inference endpoints need the hardening in our guide to securing a GPU server.
EU MDR and GDPR for imaging AI
Under Article 2(1) of the Medical Device Regulation (EU) 2017/745, software can be a medical device when the manufacturer intends it for purposes that include “diagnosis, prevention, monitoring, prediction, prognosis, treatment or alleviation of disease”. Recital 19 adds that “software for general purposes, even when used in a healthcare setting,” is not a medical device. NVIDIA states that VISTA-3D “is for research purposes and not for clinical usage”, and the NV-Reason-CXR-3B card says the model “should not be used for clinical diagnosis or treatment decisions”. Article 9(1) of the GDPR lists data concerning health among the special categories of personal data whose processing is prohibited unless an exception in Article 9(2) applies. Whether a given use is a medical device and on which legal basis images are processed is an assessment for the hospital’s legal and regulatory functions; the choice of server does not change it.
General information on EU law as of October 2026: Regulation (EU) 2017/745, Article 2(1) and recital 19, and Regulation (EU) 2016/679, Article 9(1), official texts on eur-lex.europa.eu.
Example configurations for hospitals and research labs
The configurations below are our estimates from the memory figures above, not vendor recommendations. Each assumes 3D CT or MRI at the resolutions the models document.
| SETUP | WORKLOAD | ESTIMATED SETUP |
|---|---|---|
| Radiology inference | CT segmentation through MONAI Deploy | 1 to 2 L40S in a 2U server; L4 for 3.0 mm models |
| Annotation workstation | MONAI Label with 3D Slicer or OHIF | RTX PRO 5000 (72 GB) or RTX PRO 6000 Workstation Edition |
| Research lab training | MONAI bundles, Swin UNETR, whole-body CT | 2 to 4 RTX PRO 6000 Server Edition, RAM sized for the data cache |
| Synthetic data | NV-Generate-CT up to 512 × 512 × 768 | 1 to 2 RTX PRO 6000 Server Edition |
| Large-scale training | foundation models on large datasets | 4 or 8 H200 NVL with NVLink bridges |
Our estimates from the stated memory figures in the tables above; system RAM from the whole-body bundle’s CacheDataset figures.
The two-card starter on our AI servers page pairs two RTX PRO 6000 Server Edition cards with one processor and 256 GB of RAM, which by our estimate suits inference and smaller training sets. For training with a large in-memory cache, plan system RAM from the 83 GB the whole-body bundle reports for one GPU, NVMe capacity for the raw and preprocessed volumes, and a network link to the PACS or research archive.
We check the rack, power and airflow before we quote. Describe your scan volumes, models and site in the form below, and we reply within one business day with a configuration and quote.
What we supply
We supply the GPUs that suit medical imaging work, the L4, L40S, RTX 6000 Ada, RTX PRO 5000, RTX PRO 6000 and H200 NVL, as cards or in AI servers built to order, which are assembled and burn-in tested, with manufacturer warranty and delivery anywhere in the EU, on one EU contract and invoice. NVIDIA AI Enterprise and vGPU licences come on the same invoice. Operating system, drivers, CUDA and a container runtime are installed on request. Model deployment, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How much GPU memory does MONAI need?
What GPU do I need for 3D medical image segmentation?
Can a hospital run radiology AI on premise?
What is NVIDIA Clara for medical imaging?
Which GPUs does the VISTA-3D NIM support?
Is medical imaging AI software a medical device under the MDR?
Send us the models you plan to run, the scan types and their resolution, the number of studies per day and whether you train models or only run inference. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day