GPU server tender specification: measurable requirements for an AI server RFP
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A GPU server tender should state measurable minimums and outcomes instead of product names: GPU memory per card and per server, memory bandwidth, precisions, hardware partitioning, host link, power per rack position, management, warranty, acceptance tests and licence lines
- Each threshold decides which cards can bid: FP4 Tensor Core support excludes the H200 NVL, whose NVIDIA specification table lists FP8 but no FP4, while 96 GB per GPU and 1,500 GB/s admit both the RTX PRO 6000 Server Edition and the H200 NVL
- Article 42(4) of Directive 2014/24/EU bars a reference to a make, trade mark or type unless the subject matter of the contract justifies it, or “on an exceptional basis” where a sufficiently precise and intelligible description is not possible; such a reference “shall be accompanied by the words ‘or equivalent’”
- Article 67(2) bases the award on price or cost and allows a price-quality ratio with criteria such as after-sales service and technical assistance; Article 67(4) requires that tenderers’ information can be verified
- A published MLPerf® Inference result ID (closed division, Available category) can serve as a reference for the offered GPU model and count; the acceptance test on the delivered servers, with your model and load, is what the contract can check
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What a GPU server tender specification should contain
A GPU server tender specification states measurable minimums and outcomes, not brands. It gives GPU memory per card and in total, memory bandwidth, the precisions the models need, hardware partitioning and virtualisation, the PCIe generation, the power available at each rack position, the management interface, warranty term and response time, burn-in and acceptance tests, and the licence lines.
The sizing belongs before the tender: models, precision, context length and peak concurrent requests set the memory and card count, as our guide to what to specify when buying an AI server explains. This article turns that sizing into wording that bidders can answer line by line, for a public tender or a private company’s request for proposal.
Technical requirements as measurable lines
Write each requirement as a threshold with a unit and a source of proof, usually the manufacturer’s datasheet. A threshold that only one product meets narrows the field as much as a brand name, so check which cards pass each line before publishing.
| REQUIREMENT | MEASURABLE WORDING | WHY IT MATTERS |
|---|---|---|
| GPU memory per card | at least 96 GB per GPU, ECC enabled | the largest model plus its KV cache must fit one card or one set of cards |
| GPU memory per server | at least 768 GB in at most 8 GPUs | number of model copies, or one large model per server |
| Memory bandwidth | at least 1,500 GB/s per GPU, per datasheet | sets generation speed per user; admits RTX PRO 6000 SE and H200 NVL |
| Precisions | Tensor Core support for FP8; FP4 only if planned | an FP4 line excludes the H200 NVL |
| Partitioning | hardware partitioning into at least 4 isolated instances per GPU | RTX PRO 6000 SE up to four, H200 NVL up to seven |
| GPU-to-GPU link | only if one model spans cards: GB/s per GPU | bridged H200 NVL 900 GB/s per GPU; RTX PRO 6000 SE uses PCIe |
| Host link | PCIe 5.0 x16 per GPU | NVIDIA’s configuration guide, for both cards |
| CPU and system memory | at least 6 physical cores per GPU; RAM as sized | NVIDIA’s guide gives 2 × total GPU memory as “a starting point” |
| Management and security | out-of-band controller, Redfish 1.0 or later, TPM 2.0 | remote operation and secure boot, as NVIDIA’s guide lists them |
| Power | maximum input power in kW at full GPU load, redundant supplies, outlet type | must match the feed at the rack position |
Wording and thresholds are our example for a server with eight 96 GB cards. Card values from NVIDIA’s product pages for the RTX PRO 6000 Blackwell Server Edition and the H200 NVL; CPU, memory, PCIe, Redfish and TPM lines from NVIDIA’s configuration guide for NVIDIA-Certified Systems, updated 30 September 2026; all read on 10 October 2026.
NVIDIA lists 96 GB GDDR7 with ECC and 1,597 GB/s for the RTX PRO 6000 Blackwell Server Edition. For the H200 NVL it lists 141 GB at 4.8 TB/s, so a floor of 96 GB and 1,500 GB/s admits both. The Server Edition page lists 4 PFLOPS of FP4 Tensor Core performance, while the H200 NVL table lists FP8 and no FP4. For partitioning, NVIDIA gives the Server Edition “up to four fully isolated instances” and the H200 NVL “Up to 7 MIGs @16.5GB each”, so four instances per GPU is a floor both reach.
A GPU-to-GPU line belongs in the specification only when one model has to span several cards and the sizing shows that PCIe limits it. The H200 NVL joins 2 or 4 cards with an NVLink bridge at 900 GB/s per GPU, and the RTX PRO 6000 Server Edition has no NVLink, so on a server that runs one model copy per card the line only removes bidders.
We build AI servers to order and check the rack, power and airflow before we quote. Send us the workload behind your specification, and we reply within one business day with a configuration and quote.
Brand or equivalent under Article 42 of Directive 2014/24/EU
Directive 2014/24/EU sets the EU rules for contracting authorities’ contracts at or above its thresholds. Article 42(2) requires that specifications “afford equal access of economic operators to the procurement procedure”. Article 42(3)(a), the form of the table above, allows them “in terms of performance or functional requirements, including environmental characteristics”, provided the parameters are precise enough for tenderers to determine the subject matter of the contract.
Under Article 42(4), unless justified by the subject matter of the contract, a specification may not refer to “a specific make or source”, to trade marks or to types with the effect of favouring or eliminating certain undertakings or products. Such a reference is permitted “on an exceptional basis”, where a sufficiently precise and intelligible description under paragraph 3 is not possible, and “Such reference shall be accompanied by the words ‘or equivalent’.” Where a tender uses a line such as “RTX PRO 6000 Blackwell Server Edition or equivalent”, the measurable lines of the table give bidders a way to show equivalence and the buyer a way to check it.
The Court of Justice ruled on 16 January 2025 in DYKA Plastics (C-424/23) that the list of methods for formulating technical specifications in Article 42(3) is exhaustive. It added that equal access is “necessarily infringed” where a specification that breaks Article 42(3) and (4) eliminates certain undertakings or products.
| PROVISION | WHAT IT SAYS | FOR A GPU SERVER |
|---|---|---|
| Art. 42(2) | equal access, no unjustified obstacles to competition | thresholds no higher than the sized workload |
| Art. 42(3)(a) | performance or functional requirements, incl. environmental characteristics | memory, bandwidth, partitions, input power |
| Art. 42(4) | make or trade mark only exceptionally, with “or equivalent” | criteria that define the equivalent |
| Art. 67(2) | price or cost, life-cycle costing, price-quality ratio | warranty, after-sales service, energy |
| Art. 67(4) | criteria verifiable from tenderers’ information | each criterion checkable from the bids |
| Reg. (EU) 2025/2152 | thresholds of Art. 4 for 2026 and 2027 | whether the EU rules apply to the contract |
Directive 2014/24/EU, Articles 42 and 67, as reproduced in the judgments C-424/23 (16 January 2025) and C-546/16 (20 September 2018) on eur-lex.europa.eu; Commission Delegated Regulation (EU) 2025/2152 of 22 October 2025, applicable from 1 January 2026.
General information on EU law as of October 2026: Directive 2014/24/EU, Articles 42 and 67, and Commission Delegated Regulation (EU) 2025/2152, official texts on eur-lex.europa.eu.
Which procedure applies, and whether the contract value reaches the thresholds of Article 4 as amended by Regulation (EU) 2025/2152, is decided by the procurement office, and the legal assessment of each clause is for the organisation’s legal department.
Award criteria and MLPerf® benchmark results
Article 67(2) identifies the most economically advantageous tender “on the basis of the price or cost, using a cost-effectiveness approach, such as life-cycle costing”, and allows “the best price-quality ratio”. Its examples include “after-sales service and technical assistance”, under which a tender can score warranty response and the handling of GPU replacements. Article 67(4) requires criteria accompanied by “specifications that allow the information provided by the tenderers to be effectively verified”. “GPU memory per server above the 768 GB floor, scored per 96 GB” can be verified; a criterion such as “higher AI performance” cannot.
MLCommons released the MLPerf® Inference v6.0 results on 1 April 2026. Its closed division “is intended to compare hardware platforms or software frameworks”, and systems in the Available category contain “only components that are available for purchase or for rent in the cloud”. A tender can ask bidders to name the result ID of a published MLPerf® Inference result for a system with the offered GPU model and count, closed division, Available category, from any round where one exists. MLCommons’ messaging guidelines require that any use of a result include its result ID, and state that results “may only be compared against compatible MLPerf results”, meaning the same benchmark and scenario from compatible versions.
A result shows what the benchmark’s model and software achieved on that system, not what your model will do. The contractual performance check is a load test on the delivered servers with your model, context length and concurrency, stated in the tender with the tool and settings.
Power, energy and take-back lines
State the power at each rack position in amps and phases, the PDU outlet type and the inlet temperature, and ask bidders for the maximum input power of their server at full GPU load. NVIDIA rates both the H200 NVL and the RTX PRO 6000 Server Edition at “Up to 600W (configurable)”, so eight cards account for 4.8 kW before processors and fans.
Energy can enter as an environmental characteristic under Article 42(3)(a) or through life-cycle costing under Article 67(2). Ask for idle and full-load input power read from the management controller during the acceptance run. Our article on GPU energy consumption per token shows how to measure it against the work done. Where the environmental policy asks for it, add take-back of packaging and replaced parts as a delivery line.
Warranty, burn-in, acceptance and licence lines
Specify the warranty term per component, on-site service or parts only, the response time and a single point of contact for GPU claims. Ask for a burn-in report per server showing that it ran under load before shipping, and name the on-site acceptance tests in the tender, for example dcgmi diag -r 3 with current DCGM on every GPU and a sustained power run without thermal slowdown. Our guide to GPU server acceptance testing and burn-in gives the checks and pass criteria to copy into the tender. How quotes differ on these items is covered in our article on comparing GPU server quotes line by line.
Licences need lines of their own. NVIDIA AI Enterprise is licensed per GPU, and NVIDIA states that the “H200 NVL comes with a five-year NVIDIA AI Enterprise subscription”, so ask bidders to list included subscriptions and their term separately from purchased ones, with the support level. Our article on NVIDIA AI Enterprise licensing explains when the licence is needed. If virtual machines will share the cards, add vGPU licence lines, licensed per concurrent user for graphics profiles and per GPU for compute profiles, as our article on MIG and vGPU per GPU explains. A further line can ask whether the offered server model appears in NVIDIA’s Qualified System Catalog for the offered GPU model and count, since NVIDIA lists H200 NVL server options as “NVIDIA-Certified Systems with up to 8 GPUs”.
Our servers are assembled and burn-in tested, with a test report on request. Tell us which acceptance tests your tender names and how many servers it covers.
Worked example: two inference servers for 1,500 staff
Take a company of 1,500 staff whose sizing calls for two inference servers with eight 96 GB cards each, at two rack positions so that one runs while the other is serviced. The specification asks per server for at least 768 GB of GPU memory in at most eight GPUs, at least 1,500 GB/s per GPU, FP8 support, four instances per GPU for embedding and speech models, and PCIe 5.0 x16 per GPU. It adds at least 48 physical cores, the sized system memory, two 25 GbE ports, a management port and redundant power supplies on two feeds.
The award section scores price or cost with warranty response, input power at full load and GPU memory above the floor, each with a weight and a verification method. The acceptance section names the burn-in report, the DCGM run, a sustained power run and a load test with the company’s model at its declared context and peak concurrency.
From workload to published tender
- Record the workload: models, precision, context length, peak concurrent requests and expected growth over the contract term.
- Size it into a reference configuration and check which GPU models meet each threshold.
- Turn each component into a measurable minimum with a unit and a source of proof, and use a make or type name only with “or equivalent” and the criteria for the equivalent.
- Add the site data, meaning rack positions, feeds in amps and phases, outlets, inlet temperature and network ports.
- Define the award criteria and weights, each verifiable from the information in the bids.
- Write the burn-in, acceptance and load tests and the remedy when a server fails one.
- Add the licence and warranty lines, pass the documents to the procurement office or legal department, and publish.
What we supply
We build AI servers to order with the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S and the L4, assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU. The cards also come on their own from our professional NVIDIA GPU range for existing servers. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware, under one EU contract. We check the rack, power and airflow before we quote, and the configuration and quote follow within one business day.
FAQ
What should a GPU server tender specification include?
How do I write an RFP for an AI server?
Can a public tender name a specific GPU such as the RTX PRO 6000?
What does “brand or equivalent” mean in a technical specification?
Can MLPerf results be used as tender criteria for GPU servers?
Which award criteria fit a GPU server procurement?
Send us the workload behind your specification: the models, peak concurrent requests, number of servers and the power at each rack position. We check the rack, power and airflow and reply within one business day with a configuration and quote.
Talk to an expertWe reply within one business day