BLOG · GUIDE ·

AI server for manufacturing: visual inspection, engineering and assistants on one GPU platform

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A manufacturer with 500 to 2,000 staff runs three AI workloads in two places: vision models for inspection on edge nodes at the line, and simulation, digital twins, model training and assistants on central GPU servers
  • Line-side cards are small: the passive L4 (24 GB, 72 W, single-slot low profile) for edge servers with server airflow, the actively cooled RTX PRO 4000 SFF (24 GB, 70 W) and RTX PRO 2000 (16 GB, 70 W) for industrial PCs and cabinets
  • Digital twins on Omniverse and Isaac Sim need RTX GPUs with RT cores; NVIDIA’s Omniverse server recommendation is eight RTX PRO 6000 Server Edition, while double-precision CFD points to the H200 NVL with 30 TFLOPS of FP64
  • Assistants over manuals, quality reports and service tickets are sized by requests in flight: with the example values of our company-size guide, 1,000 staff need two servers with four RTX PRO 6000 each, three of them for model copies, and 2,000 staff two servers with two H200 NVL each
  • IEC 62443-3-2 sets requirements for partitioning a system into zones and conduits with a target security level for each; edge GPU nodes usually sit in the cell zone, central servers in IT, and model updates pass through a defined conduit in planned production stops

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

AI servers for manufacturing: three workloads and where each runs

AI servers for a manufacturing company with 500 to 2,000 staff carry three kinds of workload, and each belongs in a different place. Vision models for visual inspection run at the line on small edge cards such as the L4, the RTX PRO 2000 or the RTX PRO 4000 SFF, so that the check keeps running during a WAN or data centre outage. Engineering work, simulation and digital twins run on RTX PRO 6000 cards, with the H200 NVL for double-precision solvers. Assistants and RAG over manuals, quality reports and service tickets run on central GPU servers, sized by the requests in flight.

WORKLOADWHERE IT RUNSCARDREASON
Inspection at the lineedge node per line or cellL4 with server airflow; RTX PRO 4000 SFF or RTX PRO 2000 withoutruns without the WAN; 70 to 72 W per card
Camera analytics on siteplant server roomL4, RTX PRO 6000 Server Editionmany streams meet one server; decoders set the limit
Training inspection modelscentral serverRTX PRO 6000 Server Edition, whole card or MIG instanceimages stay on site; retraining after product changes
CFD in single precisioncentral server or workstationRTX PRO 6000, Server Edition in a server96 GB and 120 TFLOPS of FP32 per Server Edition card
CFD in double precisioncentral serverH200 NVL30 TFLOPS of FP64
Digital twin, Isaac Simcentral server or workstationRTX PRO 6000 or another RTX cardRT cores required
Assistant and RAG for staffcentral server room or EU data centreRTX PRO 6000 Server Edition or H200 NVL; L4 for RAG modelsone service for all plants

Cards and figures from NVIDIA’s product pages, the L4 product brief (March 2023), Isaac Sim requirements (updated 18 September 2026) and Omniverse technical requirements (updated 8 October 2026), read on 10 October 2026; placement is our design summary.

Visual inspection at the line: edge GPUs and NVIDIA Metropolis

An inspection station takes images at line speed, a detection or segmentation model decides pass or fail, and the result drives a reject gate or a signal to the line controller. On a node next to the line, that loop does not depend on the data centre or the WAN.

The cards for such a node are small. The L4 has 24 GB, four NVDEC video decoders, four JPEG decoders and a 72 W maximum in a single-slot, low-profile card. NVIDIA’s product brief describes it as passively cooled and “requiring system airflow to operate”, so it belongs in an edge server with server airflow. In an industrial PC or a control cabinet without that airflow, the actively cooled RTX PRO 4000 SFF (24 GB, 70 W) and RTX PRO 2000 (16 GB, 70 W), both dual-slot cards of 2.7 by 6.6 inches, fit where the PC’s maker lists them for that chassis and the cabinet’s temperature. For heavier models in a full-height slot, the single-slot RTX PRO 4000 offers 24 GB and 672 GB/s at 145 W.

NVIDIA’s DeepStream stream counts are for camera pipelines with light detectors, not inspection models. Size a station’s card by your own model’s inference time per image at line speed, measured on the card before you order for every line. Camera analytics for halls, yards and loading bays, where tens or hundreds of streams meet one server, is a sizing task of its own, covered in our guide to GPU servers for video analytics.

NVIDIA calls Metropolis “a vision AI application platform and partner ecosystem”, and its Metropolis page says that customised vision AI models can help “pinpoint visual defects in industrial visual inspection applications”. The same page presents vision foundation models customised with NVIDIA TAO and deployed as NIM microservices. NVIDIA’s NIM FAQ states that “Using NIM in production requires an NVIDIA AI Enterprise license”, so a plan that runs NIM at the line counts the edge cards in the licence plan as well.

We supply the L4 and the RTX PRO 2000, 4000 and 4000 SFF for edge systems you already run, after a compatibility check of platform, power and cooling. Tell us your inspection stations and the edge hardware at the lines.

Engineering, simulation and digital twins on RTX PRO 6000

Engineering teams bring CFD and structural runs, digital twins of lines and robot cells, and rendering, and the solver or application decides the card. NVIDIA rates the H200 NVL at 30 TFLOPS of FP64, which suits double-precision solvers, while single-precision and mixed-precision solvers make use of the 120 TFLOPS of FP32 and 96 GB of the RTX PRO 6000 Server Edition. Our CFD comparison of the H200 NVL and RTX PRO 6000 works through memory per million cells and solver licences.

Digital twins built on Omniverse and Isaac Sim need RTX GPUs with RT cores. Isaac Sim’s requirements page, updated on 18 September 2026, states that “GPUs without RT Cores (A100, H100) are not supported” and lists the RTX PRO 6000 Blackwell as its ideal GPU. NVIDIA’s Omniverse technical requirements, updated on 8 October 2026, recommend one RTX PRO 6000 Blackwell with 128 to 256 GB of RAM for a Kit workstation, and “8x RTX Pro 6000 Blackwell Server Edition w/ 768 GB GPU Memory” with 64 or more CPU cores for a server. Our guide to Omniverse server requirements covers drivers, streaming and licences for digital twins.

A central RTX PRO 6000 Server Edition server can serve simulation, rendering and the training of inspection models in turn. NVIDIA’s page for the card lists up to four fully isolated MIG instances, so smaller training and test jobs can share one card, and the H200 NVL splits into up to seven of 16.5 GB each.

Assistants and RAG on manuals, quality reports and service tickets

The third workload is an assistant that answers from machine manuals, work instructions, quality reports, 8D reports and service tickets. Scanned forms and drawings in these sources need OCR or a vision-language model before they can be indexed, and the RAG pipeline adds an embedding model, a reranker and a vector index.

The server is sized by the requests in flight at the peak, not by headcount, as our guide to sizing a private ChatGPT server by company size explains. With its example values, 500 employees produce about 20 requests in flight and 2,000 about 80. With gpt-oss-120b at a declared 32K context, one RTX PRO 6000 holds about 19 conversations and one H200 NVL about 55, by that guide’s estimate. The example values assume office use, and where many staff work on the shop floor without a desk, the busy-hour share can differ, so take it from a pilot’s gateway logs. The 0.6B embedding and reranker models in that guide fit a 24 GB L4 or a MIG instance of a larger card.

vLLM, an open-source engine under the Apache 2.0 licence, also supports multimodal and embedding models. NVIDIA describes its NIM microservices as “Part of NVIDIA AI Enterprise”, and each H200 NVL comes with a five-year AI Enterprise subscription according to NVIDIA’s product page.

One central GPU server or edge nodes per plant

An edge node runs inference only, with a fixed model version, and keeps working when the link to the central site fails. The central servers train and validate models, publish approved versions for the lines, run the engineering workloads and host the assistant for all plants. With several plants, the assistant sits in one place, either the main server room or an EU data centre, and each plant keeps its edge nodes and, where it runs camera analytics, a server in its own server room.

Every edge node is a system to patch, monitor and replace. Keep one or two card types and one edge chassis across all plants, so that one spare node covers many lines and a failed node is swapped in a short stop. Size the plant uplink for retraining images sent in batches and model packages coming back, not for live video.

OT and IT separation for GPU servers under IEC 62443

The IEC 62443 series, in IEC’s words, “was developed to secure industrial automation and control systems (IACS) throughout their lifecycle”. Part 3-2 of 2020, “Security risk assessment for system design”, establishes requirements for “partitioning the SUC into zones and conduits” and for “establishing the target security level (SL-T) for each zone and conduit”, and part 3-3 sets the system security requirements and security levels. A GPU edge node at the line usually belongs to the zone of its cell and the central servers to IT, and the traffic between them, model packages down and images and results up, is a conduit with its own rules. NIST SP 800-82 Rev. 3 (September 2023) names a DMZ network architecture with firewalls among the ways to keep traffic from “passing directly between the corporate and OT networks”. The plant’s risk assessment decides the zones and target levels. Our guide to network segmentation and microsegmentation describes the OT zone and its DMZ in an IT network. Supplier access for model updates should also pass through the conduit, not through a direct route into the cell.

Production downtime and maintenance windows

A GPU node at the line is part of production, so its changes follow production’s calendar. Driver, CUDA and container updates, and new model versions, go in during planned production stops, after a test on a staging node with the same card and the same model. Keep the previous model version and driver on hand, so a change can be rolled back within the same stop.

  1. Pin the driver branch and container image per card type, and record both for every edge node.
  2. Validate a new model version on the staging node with line images, against acceptance criteria signed off by quality.
  3. Roll it out to one line first and compare its pass and fail rates with the previous version.
  4. Keep the previous version on each node until the next planned stop confirms the new one.

Plants that run three shifts have no quiet hours for the central assistant either. Two servers that each carry the whole peak, as in the company-size guide, let one server be updated while the other serves all users.

Configurations for 500 to 2,000 employees

The table applies this placement to three plant profiles. With the company-size guide’s example values, 1,000 employees give about 40 requests in flight at the peak, and three copies of gpt-oss-120b on one server hold about 57.

PLANT PROFILELINE EDGEENGINEERINGASSISTANT AND RAG
1 plant, 500 staffone L4 or RTX PRO 4000 SFF node per inspection station, one spareRTX PRO 6000 workstations for simulation and digital twinstwo servers, each 2 × RTX PRO 6000 Server Edition and 1 × L4
2 to 3 plants, 1,000 staffthe same per station, one spare node per card typeone server, 4 × RTX PRO 6000 Server Edition, for simulation and model trainingtwo servers, each 4 × RTX PRO 6000 Server Edition: three model copies and one card in MIG for RAG models
4 plants, 2,000 staffthe same per station, one spare node per plantone server, 8 × RTX PRO 6000 Server Edition; 4 × H200 NVL if CFD runs in double precision, with NVLink bridges where the solver uses themtwo servers, each 2 × H200 NVL and 1 × L4

Our estimates, not measurements. Assistant sizes from our company-size guide (gpt-oss-120b, 32K context, 16-bit KV cache; 40 per cent busy-hour users, 6 requests per hour, 30 seconds, peak factor 2); the eight-card server follows NVIDIA’s Omniverse server recommendation; edge nodes assume a few cameras per station and a test with your model.

Each assistant layout keeps the whole peak on one server if the other fails: 38 conversations for 500 employees, 57 for 1,000 and 110 for 2,000. The engineering column follows the applications more than headcount.

We build the central servers to order, supply the edge cards and check the rack, power and airflow before we quote. Send us your plants, stations, engineering applications and headcount through the form below, and the configuration and quote follow within one business day.

What we supply

We supply the cards for every layer of this platform: the L4, RTX PRO 2000, RTX PRO 4000 and RTX PRO 4000 SFF for the lines, the RTX PRO 6000 in its Workstation, Max-Q and Server editions for engineering, and the H200 NVL with NVLink bridges and the L40S for central servers. They come as GPUs for systems you already run or in AI servers built to order, assembled and burn-in tested, with manufacturer warranty and NVIDIA AI Enterprise and vGPU licences on one EU contract and invoice. Operating system, drivers, CUDA and a container runtime are installed on request. The assistants, RAG and MLOps on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

What AI server does a manufacturing company need?
Most manufacturers need two kinds of system: small edge nodes at the lines for visual inspection and central GPU servers for model training, simulation, digital twins and assistants. Edge nodes take low-power cards such as the L4, RTX PRO 2000 or RTX PRO 4000 SFF, while the central servers take RTX PRO 6000 Server Edition or H200 NVL cards. The central configuration follows from the engineering applications and from the requests in flight on the assistant.
Which GPU is suitable for visual inspection at the production line?
For an edge server with server airflow, the passively cooled L4 has 24 GB, four video decoders and four JPEG decoders at 72 W in a single-slot, low-profile card. In an industrial PC or cabinet without that airflow, the actively cooled RTX PRO 4000 SFF with 24 GB or RTX PRO 2000 with 16 GB, both at 70 W, fit where the PC’s maker lists them. Test your inspection model at the line’s frame rate and resolution on the card before ordering for all stations.
Should industrial AI run on-premise or in the cloud?
Inspection at the line should run on a node at the line, so that a WAN or data centre outage does not stop the check or let parts pass unchecked. Assistants, training and engineering workloads can run in a central server room or an EU data centre, since they serve all plants. The split keeps production independent of the link while the shared workloads stay in one place.
Can one GPU server run simulation, digital twins and LLMs together?
An RTX PRO 6000 Server Edition server can serve simulation, Omniverse and the training of inspection models in turn, and MIG splits each card into up to four isolated instances for smaller jobs. Digital twins need RT cores, which the RTX PRO 6000 has and the H100 and A100 do not, according to Isaac Sim’s requirements page. For an assistant used by all staff, separate servers that each carry the peak keep the service available during updates and failures.
How does IEC 62443 affect where AI servers sit in a factory?
IEC 62443-3-2 requires partitioning the system into zones and conduits and setting a target security level for each. A GPU node at the line usually belongs to its cell’s zone and central servers to IT, and the transfer of model packages and images between them is a conduit with its own rules, usually through a DMZ. The zones and target levels themselves come from the plant’s risk assessment.
How many GPUs does an LLM assistant for 1,000 factory employees need?
It depends on the requests in flight at the peak, not on headcount. With the example values of our company-size guide, 1,000 employees produce about 40 requests in flight, and with gpt-oss-120b at 32K one RTX PRO 6000 holds about 19 conversations, so three copies hold about 57. Two servers with four RTX PRO 6000 each, three for model copies and one in MIG for the RAG models, keep the whole peak if one server fails.

Send us your plants, inspection stations and cameras, engineering applications, headcount and the rack positions you plan to use. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna