GPU servers for pharma and life sciences: AlphaFold 3, BioNeMo, LLMs and GxP validation
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- AlphaFold 3 needs Linux, a GPU of compute capability 8.0 or higher and up to 1 TB of SSD for its genetic databases; DeepMind verified inputs of up to 5,120 tokens on one 80 GB A100 or H100, and DeepMind’s terms limit its weights to non-commercial use by or on behalf of non-commercial organisations
- The Boltz-2 NIM asks for at least 48 GB per GPU and lists the RTX PRO 6000 Workstation Edition and the L40S; the OpenFold3 NIM lists the H200 NVL, the RTX PRO 6000 Server Edition and the L40S
- BioNeMo Recipes train in FP8 from compute capability 9.0, which the H200 NVL and the RTX PRO 6000 both meet; MXFP8 on the RTX PRO 6000’s compute capability 12.0 is still marked as pending
- The H200 NVL (141 GB, 30 TFLOPS of FP64, NVLink) suits training, large inputs and double-precision codes; the RTX PRO 6000 Server Edition (96 GB, 120 TFLOPS of FP32) suits structure screening, mixed-precision molecular dynamics and mid-size LLMs
- EU GMP Annex 11 asks for validated applications and qualified IT infrastructure, and the draft Annex 22 put to consultation in 2025 says LLMs should not be used in critical GMP applications; validation stays the company’s quality decision
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What a GPU server for pharma and life sciences has to run
In pharma, biotech and CRO companies, the AI infrastructure for life sciences usually carries three kinds of work on GPU servers: structure prediction and biomolecular models such as AlphaFold 3, Boltz-2, OpenFold3 and models trained with NVIDIA BioNeMo, language models for documents and regulatory writing, and imaging, which our medical imaging AI server guide covers. The H200 NVL (141 GB, 30 TFLOPS of FP64) takes large inputs, training across bridged cards and double-precision codes. The RTX PRO 6000 Server Edition (96 GB, 120 TFLOPS of FP32) takes structure screening, mixed-precision molecular dynamics and mid-size language models.
| WORKLOAD | SOFTWARE | WHAT DOCS STATE | CARDS WE SUPPLY |
|---|---|---|---|
| Structure prediction | AlphaFold 3 | 5,120 tokens verified on one 80 GB A100 or H100; compute capability 8.0 or higher | H200 NVL; RTX PRO 6000 meets the minimum, not in DeepMind’s tests |
| Structure and affinity | Boltz-2 NIM 1.9.0 | at least 48 GB per GPU | L40S, RTX 6000 Ada, RTX PRO 6000 Workstation Edition listed |
| Structure prediction | OpenFold3 NIM | 40 GB cards and up in the list | H200 NVL, RTX PRO 6000 Server Edition, L40S listed |
| Protein model training | BioNeMo Recipes | FP8 from compute capability 9.0 | H200 NVL with NVLink; RTX PRO 6000 in BF16 or FP8 |
| Molecular dynamics | GROMACS 2026.4 | no figure; mixed precision, compute capability 5.0 or higher | RTX PRO 6000, L40S |
AlphaFold 3 installation and performance guides on GitHub; NVIDIA NIM support matrices for Boltz-2 (updated 30 September 2026) and OpenFold3 (updated 8 September 2026); BioNeMo Recipes README; GROMACS 2026.4 installation guide; NVIDIA CUDA GPUs page; all read on 10 October 2026. The card column is our mapping to the cards we supply.
AlphaFold 3 hardware requirements and the weights licence
DeepMind’s installation guide asks for Linux, “an NVIDIA GPU with Compute Capability 8.0 or greater” and up to 1 TB of disk for the genetic databases, preferably on SSD. It states that inputs “with up to 5,120 tokens can fit on a single NVIDIA A100 80 GB, or a single NVIDIA H100 80 GB”, and that GPUs with more memory can predict larger structures. For larger inputs or smaller cards, the performance guide turns on unified memory, which spills GPU memory to host memory and makes the run slower. The guide recommends at least 64 GB of RAM, since the genetic search of long targets uses a lot of it.
The repository describes the data pipeline as “CPU-only, time consuming and could be run on a machine without a GPU”, while inference needs the GPU. Predictions therefore need CPU cores and fast SSD as well as GPU memory. DeepMind has verified numerical accuracy on the A100 and H100, and calls other devices “believed to be numerically accurate” without large-scale tests. The H200 NVL is a Hopper card like the H100; the RTX PRO 6000 (compute capability 12.0) meets the minimum but is not among the tested cards.
For a pharma company, the licence terms of each model are part of the choice. The source code is under Apache 2.0, but the terms of use for the weights, last modified on 9 November 2024, make the parameters and output “only available for non-commercial use by, or on behalf of, non-commercial organizations” and exclude use “in connection with any commercial activities, including research on behalf of commercial organizations”. Google grants access to the parameters at its sole discretion. The Boltz repository states that code and weights are under the MIT licence for academic and commercial use. OpenFold3, a research preview that aims to reproduce AlphaFold 3, states that its repository is “freely available for academic and commercial use under the Apache 2.0 license” and names no separate licence for the weights. Which models and which NVIDIA software terms fit a planned use is a legal assessment for the company’s legal department.
Boltz-2, OpenFold3 and BioNeMo hardware requirements
NVIDIA packages Boltz-2 and OpenFold3 as NIM microservices whose support matrices name cards. The Boltz-2 NIM, release 1.9.0, “requires NVIDIA GPUs with at least 48 GB of GPU Memory”, plus at least 12 CPU cores, 64 GB of RAM and 80 GB of NVMe storage. Its list, updated on 30 September 2026, includes the H200, the L40S, the RTX 6000 Ada and the RTX PRO 6000 Blackwell Workstation Edition, but not the H200 NVL or the RTX PRO 6000 Server Edition. NVIDIA adds that other GPUs with at least 48 GB may work but have not been officially tested. The OpenFold3 NIM list, updated on 8 September 2026, names both, together with the L40S, and asks for at least 8 cores and 64 GB of RAM. From release 1.6.0, one OpenFold3 container given several GPUs serves requests from all of them at once.
The H200 NVL splits into up to seven MIG instances of 16.5 GB, below the 48 GB that the Boltz-2 NIM asks for, so structure models are planned on whole cards. NVIDIA’s NIM FAQ says “Using NIM in production requires an NVIDIA AI Enterprise license.” Members of the NVIDIA Developer Program may self-host NIM for research, development and experimentation on up to 16 GPUs. NVIDIA’s H200 page lists AI Enterprise as “Included” with the H200 NVL. Its RTX PRO 6000 Server Edition page lists no included subscription, so for that card the licence is bought separately.
NVIDIA describes the BioNeMo Framework as “a collection of programming tools, libraries, and models for computational drug discovery.” Its repository holds BioNeMo Recipes, with models such as ESM-2, Geneformer and AMPLIFY that scale out with FSDP. The README says FP8 “Requires compute capability 9.0 and above (Hopper+)”, which the H200 NVL (Hopper, 9.0 as NVIDIA lists the H200) and the RTX PRO 6000 (12.0) both meet. MXFP8 runs on compute capability 10.0 and 10.3, with “12.0 support pending”, so BioNeMo training on the RTX PRO 6000 runs in BF16 or FP8 for now. The older framework documentation lists neither the H200 nor any Blackwell card, so test a recipe on the target card first.
H200 NVL or RTX PRO 6000: double precision and memory
GROMACS’s installation guide says its simulations normally run in mixed floating-point precision and calls a double-precision build “slower, and not normally useful”. In FP32 the RTX PRO 6000 Server Edition offers 120 TFLOPS, the H200 NVL 60. Codes that compute in FP64 belong on the H200 NVL, which NVIDIA rates at 30 TFLOPS of FP64 and 60 on its Tensor Cores. NVIDIA’s RTX Blackwell PRO architecture whitepaper (v1.1) states that “The FP64 TFLOP rate is 1/64th the TFLOP rate of FP32 operations.” On the Server Edition that is about 1.9 TFLOPS by our arithmetic, as our CFD comparison of the H200 NVL and RTX PRO 6000 sets out.
The H200 NVL has 141 GB of HBM3e at 4.8 TB/s and joins 2 or 4 cards over an NVLink bridge at 900 GB/s per GPU. The RTX PRO 6000 Server Edition has 96 GB of GDDR7 at 1,597 GB/s and no NVLink. A single large AlphaFold-style input gains from the extra memory, while screening many mid-size complexes gains more from the number of cards. Our guide to running the H200 NVL and RTX PRO 6000 in one platform covers scheduling both pools.
We build GPU servers with H200 NVL or RTX PRO 6000 Server Edition cards around the models you run. Tell us which structure models and input sizes you plan, and whether you train with BioNeMo or only run inference.
LLMs for regulatory documents and literature
Language models in pharma draft documents, summarise literature and answer questions over internal document stores. They are sized like any company assistant, as our private LLM server sizing by company size shows for 500 and 2,000 employees.
The draft Annex 22 on artificial intelligence, which the Commission put to consultation in 2025, “does not apply to Generative AI and Large Language Models (LLM)” and says such models should not be used in critical GMP applications. In non-critical uses, the draft keeps qualified personnel responsible for the output through a human-in-the-loop approach. The EMA’s reflection paper on AI warns that “Large language models, often containing billions of parameters, are at particular risk of memorisation due to their size.” That matters when a model is fine-tuned on documents that contain personal data. A platform for these uses needs a query log and access rights per user.
GxP: EU GMP Annex 11 and the EMA reflection paper on AI
Annex 11 of EudraLex Volume 4 (revision January 2011, to come into operation by 30 June 2011) “applies to all forms of computerised systems used as part of a GMP regulated activities.” Its principle reads “The application should be validated; IT infrastructure should be qualified.” Its glossary counts hardware and operating systems as IT infrastructure, so a GPU server used in a GMP-regulated activity is infrastructure to be qualified, and the application on it is to be validated, under the regulated company’s responsibility. Annex 11 bases the extent of validation on a documented risk assessment.
| ANNEX 11 CLAUSE | WHAT IT ASKS | PLATFORM PROVIDES |
|---|---|---|
| Principle | “IT infrastructure should be qualified” | recorded hardware, firmware, driver and CUDA versions; a load test before shipping, with a test report on request |
| Suppliers (3) | formal agreements stating third-party responsibilities | a contract that says who supplies, installs and maintains what |
| Back-ups (7.2) | regular back-ups; restore checked during validation | back-ups of model weights, container images, inputs and results, with restore tests |
| Audit trails (9) | a record of GMP-relevant changes and deletions, based on risk | job and model-version logs; for LLMs, logged queries and answers |
| Change management (10) | changes only in a controlled manner | pinned driver, CUDA and container versions changed by a defined procedure |
| Security (12.1) | access restricted to authorised persons | role-based access; management controllers on a separate network |
| Business continuity (16) | support for processes when a system breaks down | a second server or spare capacity for the critical workload |
EudraLex Volume 4, Annex 11 (revision January 2011), clauses as numbered there; the right-hand column is our reading of what a GPU platform contributes, and the company’s quality unit decides what applies.
The EMA reflection paper EMA/CHMP/CVMP/83833/2023, adopted in September 2024, places responsibility with the clinical trial sponsor, the marketing authorisation applicant or holder, or the manufacturer. They are to ensure that “all algorithms, models, datasets, and data processing pipelines used are fit for purpose”. As of October 2026, the Commission’s EudraLex page lists the 2011 Annex 11 and no Annex 22.
General information on EU law as of October 2026: EudraLex Volume 4, Annex 11 and the consultation drafts of Annex 11 and Annex 22 (European Commission, health.ec.europa.eu), and the EMA reflection paper EMA/CHMP/CVMP/83833/2023 (ema.europa.eu).
The validation plan, the risk assessment and the release of a system for GxP use stay with the company’s quality unit. Our article on AI governance with ISO/IEC 42001 for a private LLM maps query logs, access rights and evaluation records to management-system controls.
Example platform for a mid-size pharma company or CRO
Our estimate, not a vendor recommendation, is for a company of about 1,000 staff with a computational team and an internal assistant. One server holds four H200 NVL cards on a four-way NVLink bridge for BioNeMo training in FP8, structure inputs too large for 96 GB, double-precision codes and the larger language model. A second server holds eight RTX PRO 6000 Server Edition cards for screening with the OpenFold3 NIM across all eight cards, GROMACS runs and mid-size models for the assistant. Boltz-2 joins it after a test, since its NIM list does not name the Server Edition.
Plan the CPU side from the documented minimums, 12 cores and 64 GB of RAM for the Boltz-2 NIM and 8 cores and 64 GB for the OpenFold3 NIM. Eight cards at up to 600 W each draw up to 4.8 kW before processors and fans, so the rack’s power feeds and cooling are planned for that load. If each server can run the assistant’s model, one keeps it running when the other fails. For workloads in GMP scope, such spare capacity is one input to the continuity provisions of Annex 11, clause 16.
We check the rack, power and airflow before we quote. Describe your teams, models and GxP scope in the form below, and the configuration and quote follow within one business day.
What we supply
We supply the H200 NVL, the RTX PRO 6000 in its Server and Workstation editions and the L40S as professional NVIDIA GPUs, and build AI servers to order around them, assembled and burn-in tested, with manufacturer warranty on every component, on one EU contract and invoice. NVIDIA AI Enterprise licences for NIM in production come on the same invoice. Operating system, drivers, CUDA and a container runtime are installed on request. Models, query logging and access management on top are our Private AI/ML service, with engineering by our partner Vixen.UNO. Qualification of the infrastructure and validation of the applications stay with your quality unit.
FAQ
What are the AlphaFold 3 hardware requirements?
Can a pharma company use AlphaFold 3?
What are the BioNeMo hardware requirements?
What GPU server does a pharma company need for protein structure prediction?
Does EU GMP Annex 11 apply to an AI server?
Can LLMs be used in GxP processes?
Send us the structure and language models you plan to run, their input sizes, the number of users, whether you train or only run inference, and which uses are GxP-relevant. We reply within one business day with a configuration and a written quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day