AI server for manufacturing: visual inspection, engineering and assistants on one GPU platform
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A manufacturer with 500 to 2,000 staff runs three AI workloads in two places: vision models for inspection on edge nodes at the line, and simulation, digital twins, model training and assistants on central GPU servers
- Line-side cards are small: the passive L4 (24 GB, 72 W, single-slot low profile) for edge servers with server airflow, the actively cooled RTX PRO 4000 SFF (24 GB, 70 W) and RTX PRO 2000 (16 GB, 70 W) for industrial PCs and cabinets
- Digital twins on Omniverse and Isaac Sim need RTX GPUs with RT cores; NVIDIA’s Omniverse server recommendation is eight RTX PRO 6000 Server Edition, while double-precision CFD points to the H200 NVL with 30 TFLOPS of FP64
- Assistants over manuals, quality reports and service tickets are sized by requests in flight: with the example values of our company-size guide, 1,000 staff need two servers with four RTX PRO 6000 each, three of them for model copies, and 2,000 staff two servers with two H200 NVL each
- IEC 62443-3-2 sets requirements for partitioning a system into zones and conduits with a target security level for each; edge GPU nodes usually sit in the cell zone, central servers in IT, and model updates pass through a defined conduit in planned production stops
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
AI servers for manufacturing: three workloads and where each runs
AI servers for a manufacturing company with 500 to 2,000 staff carry three kinds of workload, and each belongs in a different place. Vision models for visual inspection run at the line on small edge cards such as the L4, the RTX PRO 2000 or the RTX PRO 4000 SFF, so that the check keeps running during a WAN or data centre outage. Engineering work, simulation and digital twins run on RTX PRO 6000 cards, with the H200 NVL for double-precision solvers. Assistants and RAG over manuals, quality reports and service tickets run on central GPU servers, sized by the requests in flight.
| WORKLOAD | WHERE IT RUNS | CARD | REASON |
|---|---|---|---|
| Inspection at the line | edge node per line or cell | L4 with server airflow; RTX PRO 4000 SFF or RTX PRO 2000 without | runs without the WAN; 70 to 72 W per card |
| Camera analytics on site | plant server room | L4, RTX PRO 6000 Server Edition | many streams meet one server; decoders set the limit |
| Training inspection models | central server | RTX PRO 6000 Server Edition, whole card or MIG instance | images stay on site; retraining after product changes |
| CFD in single precision | central server or workstation | RTX PRO 6000, Server Edition in a server | 96 GB and 120 TFLOPS of FP32 per Server Edition card |
| CFD in double precision | central server | H200 NVL | 30 TFLOPS of FP64 |
| Digital twin, Isaac Sim | central server or workstation | RTX PRO 6000 or another RTX card | RT cores required |
| Assistant and RAG for staff | central server room or EU data centre | RTX PRO 6000 Server Edition or H200 NVL; L4 for RAG models | one service for all plants |
Cards and figures from NVIDIA’s product pages, the L4 product brief (March 2023), Isaac Sim requirements (updated 18 September 2026) and Omniverse technical requirements (updated 8 October 2026), read on 10 October 2026; placement is our design summary.
Visual inspection at the line: edge GPUs and NVIDIA Metropolis
An inspection station takes images at line speed, a detection or segmentation model decides pass or fail, and the result drives a reject gate or a signal to the line controller. On a node next to the line, that loop does not depend on the data centre or the WAN.
The cards for such a node are small. The L4 has 24 GB, four NVDEC video decoders, four JPEG decoders and a 72 W maximum in a single-slot, low-profile card. NVIDIA’s product brief describes it as passively cooled and “requiring system airflow to operate”, so it belongs in an edge server with server airflow. In an industrial PC or a control cabinet without that airflow, the actively cooled RTX PRO 4000 SFF (24 GB, 70 W) and RTX PRO 2000 (16 GB, 70 W), both dual-slot cards of 2.7 by 6.6 inches, fit where the PC’s maker lists them for that chassis and the cabinet’s temperature. For heavier models in a full-height slot, the single-slot RTX PRO 4000 offers 24 GB and 672 GB/s at 145 W.
NVIDIA’s DeepStream stream counts are for camera pipelines with light detectors, not inspection models. Size a station’s card by your own model’s inference time per image at line speed, measured on the card before you order for every line. Camera analytics for halls, yards and loading bays, where tens or hundreds of streams meet one server, is a sizing task of its own, covered in our guide to GPU servers for video analytics.
NVIDIA calls Metropolis “a vision AI application platform and partner ecosystem”, and its Metropolis page says that customised vision AI models can help “pinpoint visual defects in industrial visual inspection applications”. The same page presents vision foundation models customised with NVIDIA TAO and deployed as NIM microservices. NVIDIA’s NIM FAQ states that “Using NIM in production requires an NVIDIA AI Enterprise license”, so a plan that runs NIM at the line counts the edge cards in the licence plan as well.
We supply the L4 and the RTX PRO 2000, 4000 and 4000 SFF for edge systems you already run, after a compatibility check of platform, power and cooling. Tell us your inspection stations and the edge hardware at the lines.
Engineering, simulation and digital twins on RTX PRO 6000
Engineering teams bring CFD and structural runs, digital twins of lines and robot cells, and rendering, and the solver or application decides the card. NVIDIA rates the H200 NVL at 30 TFLOPS of FP64, which suits double-precision solvers, while single-precision and mixed-precision solvers make use of the 120 TFLOPS of FP32 and 96 GB of the RTX PRO 6000 Server Edition. Our CFD comparison of the H200 NVL and RTX PRO 6000 works through memory per million cells and solver licences.
Digital twins built on Omniverse and Isaac Sim need RTX GPUs with RT cores. Isaac Sim’s requirements page, updated on 18 September 2026, states that “GPUs without RT Cores (A100, H100) are not supported” and lists the RTX PRO 6000 Blackwell as its ideal GPU. NVIDIA’s Omniverse technical requirements, updated on 8 October 2026, recommend one RTX PRO 6000 Blackwell with 128 to 256 GB of RAM for a Kit workstation, and “8x RTX Pro 6000 Blackwell Server Edition w/ 768 GB GPU Memory” with 64 or more CPU cores for a server. Our guide to Omniverse server requirements covers drivers, streaming and licences for digital twins.
A central RTX PRO 6000 Server Edition server can serve simulation, rendering and the training of inspection models in turn. NVIDIA’s page for the card lists up to four fully isolated MIG instances, so smaller training and test jobs can share one card, and the H200 NVL splits into up to seven of 16.5 GB each.
Assistants and RAG on manuals, quality reports and service tickets
The third workload is an assistant that answers from machine manuals, work instructions, quality reports, 8D reports and service tickets. Scanned forms and drawings in these sources need OCR or a vision-language model before they can be indexed, and the RAG pipeline adds an embedding model, a reranker and a vector index.
The server is sized by the requests in flight at the peak, not by headcount, as our guide to sizing a private ChatGPT server by company size explains. With its example values, 500 employees produce about 20 requests in flight and 2,000 about 80. With gpt-oss-120b at a declared 32K context, one RTX PRO 6000 holds about 19 conversations and one H200 NVL about 55, by that guide’s estimate. The example values assume office use, and where many staff work on the shop floor without a desk, the busy-hour share can differ, so take it from a pilot’s gateway logs. The 0.6B embedding and reranker models in that guide fit a 24 GB L4 or a MIG instance of a larger card.
vLLM, an open-source engine under the Apache 2.0 licence, also supports multimodal and embedding models. NVIDIA describes its NIM microservices as “Part of NVIDIA AI Enterprise”, and each H200 NVL comes with a five-year AI Enterprise subscription according to NVIDIA’s product page.
One central GPU server or edge nodes per plant
An edge node runs inference only, with a fixed model version, and keeps working when the link to the central site fails. The central servers train and validate models, publish approved versions for the lines, run the engineering workloads and host the assistant for all plants. With several plants, the assistant sits in one place, either the main server room or an EU data centre, and each plant keeps its edge nodes and, where it runs camera analytics, a server in its own server room.
Every edge node is a system to patch, monitor and replace. Keep one or two card types and one edge chassis across all plants, so that one spare node covers many lines and a failed node is swapped in a short stop. Size the plant uplink for retraining images sent in batches and model packages coming back, not for live video.
OT and IT separation for GPU servers under IEC 62443
The IEC 62443 series, in IEC’s words, “was developed to secure industrial automation and control systems (IACS) throughout their lifecycle”. Part 3-2 of 2020, “Security risk assessment for system design”, establishes requirements for “partitioning the SUC into zones and conduits” and for “establishing the target security level (SL-T) for each zone and conduit”, and part 3-3 sets the system security requirements and security levels. A GPU edge node at the line usually belongs to the zone of its cell and the central servers to IT, and the traffic between them, model packages down and images and results up, is a conduit with its own rules. NIST SP 800-82 Rev. 3 (September 2023) names a DMZ network architecture with firewalls among the ways to keep traffic from “passing directly between the corporate and OT networks”. The plant’s risk assessment decides the zones and target levels. Our guide to network segmentation and microsegmentation describes the OT zone and its DMZ in an IT network. Supplier access for model updates should also pass through the conduit, not through a direct route into the cell.
Production downtime and maintenance windows
A GPU node at the line is part of production, so its changes follow production’s calendar. Driver, CUDA and container updates, and new model versions, go in during planned production stops, after a test on a staging node with the same card and the same model. Keep the previous model version and driver on hand, so a change can be rolled back within the same stop.
- Pin the driver branch and container image per card type, and record both for every edge node.
- Validate a new model version on the staging node with line images, against acceptance criteria signed off by quality.
- Roll it out to one line first and compare its pass and fail rates with the previous version.
- Keep the previous version on each node until the next planned stop confirms the new one.
Plants that run three shifts have no quiet hours for the central assistant either. Two servers that each carry the whole peak, as in the company-size guide, let one server be updated while the other serves all users.
Configurations for 500 to 2,000 employees
The table applies this placement to three plant profiles. With the company-size guide’s example values, 1,000 employees give about 40 requests in flight at the peak, and three copies of gpt-oss-120b on one server hold about 57.
| PLANT PROFILE | LINE EDGE | ENGINEERING | ASSISTANT AND RAG |
|---|---|---|---|
| 1 plant, 500 staff | one L4 or RTX PRO 4000 SFF node per inspection station, one spare | RTX PRO 6000 workstations for simulation and digital twins | two servers, each 2 × RTX PRO 6000 Server Edition and 1 × L4 |
| 2 to 3 plants, 1,000 staff | the same per station, one spare node per card type | one server, 4 × RTX PRO 6000 Server Edition, for simulation and model training | two servers, each 4 × RTX PRO 6000 Server Edition: three model copies and one card in MIG for RAG models |
| 4 plants, 2,000 staff | the same per station, one spare node per plant | one server, 8 × RTX PRO 6000 Server Edition; 4 × H200 NVL if CFD runs in double precision, with NVLink bridges where the solver uses them | two servers, each 2 × H200 NVL and 1 × L4 |
Our estimates, not measurements. Assistant sizes from our company-size guide (gpt-oss-120b, 32K context, 16-bit KV cache; 40 per cent busy-hour users, 6 requests per hour, 30 seconds, peak factor 2); the eight-card server follows NVIDIA’s Omniverse server recommendation; edge nodes assume a few cameras per station and a test with your model.
Each assistant layout keeps the whole peak on one server if the other fails: 38 conversations for 500 employees, 57 for 1,000 and 110 for 2,000. The engineering column follows the applications more than headcount.
We build the central servers to order, supply the edge cards and check the rack, power and airflow before we quote. Send us your plants, stations, engineering applications and headcount through the form below, and the configuration and quote follow within one business day.
What we supply
We supply the cards for every layer of this platform: the L4, RTX PRO 2000, RTX PRO 4000 and RTX PRO 4000 SFF for the lines, the RTX PRO 6000 in its Workstation, Max-Q and Server editions for engineering, and the H200 NVL with NVLink bridges and the L40S for central servers. They come as GPUs for systems you already run or in AI servers built to order, assembled and burn-in tested, with manufacturer warranty and NVIDIA AI Enterprise and vGPU licences on one EU contract and invoice. Operating system, drivers, CUDA and a container runtime are installed on request. The assistants, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What AI server does a manufacturing company need?
Which GPU is suitable for visual inspection at the production line?
Should industrial AI run on-premise or in the cloud?
Can one GPU server run simulation, digital twins and LLMs together?
How does IEC 62443 affect where AI servers sit in a factory?
How many GPUs does an LLM assistant for 1,000 factory employees need?
Send us your plants, inspection stations and cameras, engineering applications, headcount and the rack positions you plan to use. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day