AI servers for energy and utility companies: on-premise LLMs, forecasting and NIS2
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Energy (electricity, district heating and cooling, oil, gas, hydrogen), drinking water and waste water are sectors of high criticality in Annex I of NIS2, so for a utility in scope the AI platform is one more system under its Article 21 risk-management measures
- A utility with 500 to 2,000 staff runs three kinds of AI workload: LLM assistants over technical documentation, regulations and maintenance records, forecasting models for load and generation, and image models for inspecting lines, substations and plants
- With the example values of our company-size guide, 2,000 staff peak at 80 requests in flight; two inference servers with two H200 NVL or five RTX PRO 6000 Server Edition each carry that peak alone, plus one L40S per server for retrieval in this article’s example
- Forecasting models are small next to LLMs: Chronos-2 has 120M parameters and TimesFM 2.5 0.2B, and Amazon’s model card states GPU and CPU inference, so an L4 or a MIG instance of an RTX PRO 6000 is the size to test first
- The AI servers sit in the IT zone and read grid data from a copy in a DMZ, over a path that IEC 62443-3-2 terms a conduit between zones; supply chain security under Article 21(2)(d) covers the platform’s direct suppliers, such as the server supplier and the engineering partner
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
AI servers for energy and utility companies under NIS2
Energy, drinking water and waste water are sectors of high criticality in Annex I of the NIS2 Directive, so a utility in its scope treats an on-premise AI platform as one more network and information system covered by its Article 21 risk-management measures. The platform usually carries three kinds of workload: LLM assistants over technical documentation, regulations and maintenance records, forecasting models for load and generation, and image models for inspecting lines, substations and plants. All three run in the IT zone, apart from the control networks, and read operational data through a defined path.
For 500 to 2,000 staff, the platform fits on two to four GPU servers. Two inference servers each carry the assistants’ peak alone, so that one can fail or be patched. A third server with RTX PRO 6000 Server Edition cards runs forecasting, weather and vision models, and a fourth can hold a copy at a second site. The sections below size each part and map the NIS2 measures that apply to the platform.
Which energy and water companies NIS2 covers
Annex I of Directive (EU) 2022/2555 lists, under energy, electricity suppliers, distribution and transmission system operators, producers, operators of district heating or district cooling, oil and gas operators and operators of hydrogen production, storage and transmission, among other types. Under points 6 and 7 it lists suppliers and distributors of water intended for human consumption and undertakings collecting, disposing of or treating urban, domestic or industrial waste water. Waste management is in Annex II.
Under Article 3(1)(a), entities of an Annex I type that exceed the ceilings for medium-sized enterprises in Recommendation 2003/361/EC are essential entities, and medium-sized ones are, as a rule, important entities under Article 3(2). In that recommendation, the category of micro, small and medium-sized enterprises is made up of enterprises “which employ fewer than 250 persons”, with financial ceilings alongside the headcount. Some entities are covered regardless of size under Article 2(2), and entities identified as critical entities under Directive (EU) 2022/2557 are essential entities whatever their size (Articles 2(3) and 3(1)(f)). Whether a particular utility is essential, important or outside the scope is a legal assessment for its legal department. Our guide to the NIS2 Article 21 security measures covers all ten measures, and this article only applies them to an AI platform.
Workloads, model classes and GPUs for a utility
The three workloads differ by orders of magnitude in model size, so they need different cards, and putting them on one large server mixes jobs with different maintenance and availability needs.
| WORKLOAD | MODEL CLASS | EXAMPLES | GPU |
|---|---|---|---|
| Assistant over documents | LLM with RAG, open weights | gpt-oss-120b, 117B total, 5.1B active | H200 NVL or RTX PRO 6000 Server Edition |
| Embedding and reranking | retrieval models of about 1B or less | the platform’s embedder and reranker | L40S per inference server, or L4 |
| Load and PV forecasts | time-series foundation models | Chronos-2 (120M), TimesFM 2.5 (0.2B) | L4 or a MIG instance; CPU for small runs |
| Weather for wind and PV | AI weather and downscaling models | FourCastNet 3, CorrDiff in Earth2Studio | RTX PRO 6000 or H200 NVL, tested first |
| Asset image inspection | detection and segmentation models | drone and substation camera images | L4 for streams, RTX PRO 6000 for training |
Model facts from the Hugging Face model cards of openai/gpt-oss-120b, amazon/chronos-2 and google/timesfm-2.5-200m-pytorch and from the Earth2Studio documentation, read on 10 October 2026; the GPU column is our sizing.
On-premise LLMs for technical documentation and maintenance records
A utility’s assistant answers from manufacturers’ manuals for transformers, switchgear and turbines, internal operating instructions, network codes and regulatory decisions, maintenance records, work orders and incident reports. Much of this describes the grid and its weak points, so the documents and the index built from them stay on hardware the company controls. A RAG platform indexes the document stores and passes each user’s access rights into retrieval, so a field technician and a grid planner see answers from different sources.
OpenAI publishes gpt-oss-120b under the Apache 2.0 licence, with 117B parameters in total and 5.1B active, and its model card says MXFP4 quantisation lets it “run on a single 80GB GPU”. Our company-size guide sizes it at a declared 32K context with a 16-bit cache: one RTX PRO 6000 holds about 19 conversations and one H200 NVL about 55, by that guide’s estimate. With its example values of 40 per cent busy-hour users, 6 requests an hour, 30 seconds per request and a peak factor of 2, 2,000 staff peak at 80 requests in flight. The private AI platform architecture for 2,000 users shows the gateway, vector database and control plane roles around these servers.
Load, generation and weather forecasting on GPUs
Load forecasts per substation or feeder and generation forecasts for wind and photovoltaic plants are time series with weather and calendar covariates. Time-series foundation models forecast such series without training on each one. Amazon describes Chronos-2 as “a 120M-parameter, encoder-only time series foundation model for zero-shot forecasting” under the Apache 2.0 licence, and its model card says “Chronos-2 supports all covariate types natively”. The card reports “over 300 time series forecasts per second on a single A10G GPU” and support for “both GPU and CPU inference”. Google lists TimesFM 2.5 at 0.2B parameters, also under Apache 2.0.
In 32-bit precision, 120M parameters take about 0.5 GB, so memory does not limit these models on any card in our range. What decides the hardware is the number of series, the forecast interval and the time a run may take. A run over a few hundred substations is a candidate for CPU inference, measured on your own data before any card is bought, while forecasts for tens of thousands of meters or fine-tuning on years of history justify an L4 or a MIG instance of an RTX PRO 6000 Server Edition, a card that MIG divides into “up to four fully isolated instances”, in NVIDIA’s words.
Weather models are larger. NVIDIA describes Earth2Studio as an “Open-source deep-learning framework for exploring, building and deploying AI weather and climate workflows”, with medium-range forecast models such as FourCastNet 3 and downscaling models such as CorrDiff. Its overview page, read on 10 October 2026, names no GPU requirement, so test the model and resolution you need on an RTX PRO 6000 or H200 NVL before sizing a server around it.
Image inspection of lines, substations and plants
Drone flights over overhead lines, fixed cameras in substations and thermal images of plants produce two kinds of load. Batches of still images from a flight are processed after landing, and a server works through them at its own pace. Fixed cameras send continuous streams, which are limited first by the card’s video decoders, as our guide to the GPU server for video analytics shows for the L4, L40S and RTX PRO cards. Training the detection and segmentation models on the company’s own defect images is the heavier job and fits the RTX PRO 6000 Server Edition, whose MIG instances let smaller training runs share one card.
Two to four servers for 500 to 2,000 staff
The inference servers are sized so that each carries the assistant’s peak alone. The third server takes forecasting, weather and vision work, which tolerates a pause during maintenance better than the assistant does.
| STAFF | PEAK IN FLIGHT | EACH OF 2 SERVERS | THIRD SERVER |
|---|---|---|---|
| 500 | 20 | 2 × RTX PRO 6000 (38 held) or 1 × H200 NVL (55), plus 1 × L4 | 2 × RTX PRO 6000 Server Edition |
| 1,000 | 40 | 3 × RTX PRO 6000 (57) or 1 × H200 NVL (55), plus 1 × L40S | 4 × RTX PRO 6000 Server Edition |
| 2,000 | 80 | 5 × RTX PRO 6000 (95) or 2 × H200 NVL (110), plus 1 × L40S | 4 × RTX PRO 6000 Server Edition, fourth server at a second site |
Peaks and conversations held from the example values and gpt-oss-120b estimates of our company-size guide (32K context, 16-bit cache); retrieval cards and the third server are our example layout.
The H200 NVL and the RTX PRO 6000 Server Edition each run at up to 600 W, so five RTX PRO 6000 cards draw up to 3 kW before processors and fans.
We build AI servers to order with H200 NVL, RTX PRO 6000 Server Edition, L40S and L4 cards and check the rack, power and airflow before we quote. Send us your headcount, use cases and rack positions through the form below.
Keeping the AI platform apart from OT networks
The AI servers belong in the IT data centre, outside the control network. Forecasting needs measurements from the SCADA historian, and the usual route is a copy of that data in a DMZ between the OT and IT networks, from which the platform reads; forecasts return to control-room systems through the same defined path. In IEC’s words, the IEC 62443 series “was developed to secure industrial automation and control systems (IACS) throughout their lifecycle”, and part 3-2 of 2020, “Security risk assessment for system design”, sets requirements for “partitioning the SUC into zones and conduits” (SUC is the system under consideration) and for “establishing the target security level (SL-T) for each zone and conduit”. In those terms, the path between historian and AI platform is a conduit with its own rules, and the operator’s risk assessment sets the zones and levels. Implementing Regulation (EU) 2024/2690 binds only the providers listed in its Article 1, such as cloud, data centre and managed service providers, but two of its points describe this layout. Point 6.8.2(d) asks those providers to “deploy a demilitarised zone within their communication networks”, and point 6.7.2(h) describes supplier access that fits model and driver updates: “allow connections of service providers only after an authorisation request and for a set time period”. Our guide to network segmentation and microsegmentation covers the OT zone and its DMZ.
NIS2 measures applied to the AI platform
Article 21(1) has the measures take into account, where applicable, “relevant European and international standards”. For the AI platform, most of the ten measures translate into a few concrete settings and records.
| ARTICLE 21(2) | NIS2 WORDING | FOR THE AI PLATFORM |
|---|---|---|
| Point (c) | “backup management and disaster recovery” | back up model files, vector index, prompts and configuration; second server or site for the assistant |
| Point (d) | “supply chain security” | server, licence and engineering suppliers in the supplier register; open models and serving engine with version, source and licence in the inventory |
| Point (e) | “vulnerability handling and disclosure” | drivers, CUDA, serving engine and containers patched in maintenance windows, versions pinned |
| Point (h) | “use of cryptography and, where appropriate, encryption” | TLS on the gateway and between services, encrypted disks for documents and index |
| Point (i) | “access control policies and asset management” | access rights per document carried into retrieval; servers, cards and models in the asset inventory |
| Point (j) | “multi-factor authentication” | MFA through the directory service for users and administrators |
Wording from Article 21(2) of Directive (EU) 2022/2555, CELEX 32022L2555, text copied on 6 October 2026; the right column is our reading for an AI platform.
Under Article 21(3), entities take into account “the vulnerabilities specific to each direct supplier and service provider and the overall quality of products and cybersecurity practices of their suppliers and service providers, including their secure development procedures”. For the AI platform the direct suppliers and service providers are typically the server supplier and the engineering partner. An open-source serving engine and an open model downloaded from its publisher usually come without a contract, so they go into the asset inventory with version, source and licence and are assessed like other third-party software. Our guide to NIS2 supply chain security lists the contract clauses and evidence suppliers are asked for.
Our Private AI/ML service, with engineering by our partner Vixen.UNO, includes logging of queries and answers and assistants that respect each user’s access rights. Describe your directory service, logging and backup setup in the form below for an answer within one business day.
What we supply
We build AI servers to order for this platform, with H200 NVL, RTX PRO 6000 Server Edition, L40S and L4 cards from our professional GPU range, assembled and burn-in tested, with manufacturer warranty and delivery anywhere in the EU on one EU contract and invoice. NVIDIA AI Enterprise and vGPU licences come on the same invoice as the hardware. We check the rack, power and airflow before we quote. The platform on top, with models, RAG and MLOps, is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What AI server does an energy company need?
Can a utility company run private AI on premise?
Which LLM can a utility run for technical documentation?
Do I need a GPU for grid load forecasting?
How does NIS2 apply to AI at energy companies?
Should AI servers sit in the OT network of a utility?
Send us your headcount, the use cases (documents, forecasting, asset images), the models you are considering, how your IT and OT networks are separated and the rack positions available. We reply within one business day with a configuration and quote for the GPU servers, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day