BLOG · GUIDE ·

Air-gapped LLM server: installing NVIDIA drivers, containers and models offline

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Everything the server would download, from the driver, kernel headers and Container Toolkit to images, model weights and Python wheels, is fetched at exact versions on a connected staging host, verified, carried across with a SHA-256 manifest and installed from internal mirrors
  • NVIDIA’s driver installation guide of 23 September 2026 offers a local repository package for Ubuntu; the kernel headers come from the distribution’s repository, and the branch pinning package goes in before the driver
  • The GPU Operator’s air-gapped guide for v26.7.1 needs a local image registry and, for its driver container, a local package repository; with the driver installed on the hosts, driver.enabled=false leaves that container out
  • Hugging Face weights are downloaded at a fixed revision and checked with hf cache verify on the connected side; on the server, HF_HUB_OFFLINE=1 stops all HTTP calls to the Hub and vLLM serves the model from a local path
  • NVIDIA AI Enterprise software runs with or without a licence server; only vGPU for Compute needs the NVIDIA License System, offline through an on-premises DLS instance with manually uploaded licences or a node-locked licence file

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

What an air-gapped LLM server needs

An air-gapped LLM server needs everything it would otherwise download to be fetched on a connected staging host, checked there, carried across and installed from local copies. That covers the NVIDIA driver and the kernel headers it builds against, the container runtime, container images, model weights, Python packages and, for vGPU for Compute only, licences. Versions and dates in this article are as read on 10 October 2026.

COMPONENTOFFLINE METHODDOCUMENTED IN
NVIDIA driver (Ubuntu)local repository package and branch pinning packageNVIDIA driver installation guide, 23 September 2026
Kernel headers, OS packagesmirror of the distribution repository, for example with apt-mirrorNVIDIA driver installation guide; GPU Operator air-gapped guide, 23 September 2026
Container Toolkitfour packages at one pinned versionNVIDIA Container Toolkit install guide, 18 September 2026
Container imagesinternal registry, filled with skopeo sync or docker pull, tag and pushskopeo-sync manual; GPU Operator air-gapped guide
GPU Operatorchart from helm fetch, local registry and package repository in values.yamlGPU Operator air-gapped guide
Model weightshf download at a fixed revision, local path, HF_HUB_OFFLINE=1Hugging Face Hub documentation; vLLM engine arguments
NIM containersmodel cache or model store prepared on the connected sideNVIDIA NIM air-gap guide, 6 October 2026
Python packagespip download, then pip install with --no-index and --find-linkspip user guide, version 26.2.1
Licencesno licence server on bare metal; DLS or node-locked file for vGPU for ComputeNVIDIA licensing guide; NVIDIA License System guide 3.6.1

Sources: NVIDIA, Hugging Face, vLLM, pip and skopeo documentation, read on 10 October 2026; dates are those shown on each page.

Staging host, transfer path and checksums

The staging host sits outside the air gap and holds the mirrors: a copy of the distribution repository, NVIDIA’s driver and toolkit packages, an image registry or image directory, the model store and a wheel directory. Give it the same operating system release and processor architecture as the servers inside, because packages, kernel headers and Python wheels are built per release and architecture. Files cross through one path, removable media or a one-way transfer system as your security policy allows, and are checked on both sides.

  1. Download each component on the staging host at an exact version: driver version, toolkit version, image digest and model commit.
  2. Verify it there against what the publisher provides: signed metadata for packages from a network repository, hf cache verify for Hugging Face weights.
  3. Write a manifest with the SHA-256 value of every file that will cross.
  4. Copy the files and the manifest to the transfer medium.
  5. On the inside, check every file against the manifest with sha256sum -c before it enters a mirror.
  6. Import the files into the internal mirrors and install or update the servers from those mirrors only.

Hugging Face describes hf cache verify as a way to “validate local files against their checksums on the Hub”. It needs the Hub, so it runs on the staging host. On a mismatch it prints a list of the files and exits with a non-zero status, but a missing file only produces a warning unless you add --fail-on-missing-files. Keep every manifest, since it records which versions reached the inside and when, and a restore needs the same record, as our article on backing up an AI server explains.

Installing the NVIDIA driver offline on Ubuntu

NVIDIA’s driver installation guide, updated on 23 September 2026 and written for branch 615, gives Ubuntu two methods: local repository enablement and network repository enablement. The steps below are for Ubuntu only; Red Hat Enterprise Linux, SUSE, Debian and other distributions have their own pages in the same guide. The repository comes as one Debian package whose name starts with nvidia-driver-local-repo- and continues with the distribution, the driver version and the architecture. The guide installs it with dpkg -i, runs apt update and enrols the “ephemeral” GPG key it carries by copying its keyring file to /usr/share/keyrings/; apt install nvidia-open then installs the open kernel modules. Since the key comes inside the package, record the package’s SHA-256 value on the staging host when you download it from NVIDIA.

The kernel headers come from outside that package. The guide’s preparation step installs them for the running kernel with apt install linux-headers-$(uname -r) from the distribution’s own repository, which also supplies DKMS, the framework the nvidia-dkms-open package uses to build the modules. The internal mirror must therefore hold the headers for every kernel your servers boot. Install the pinning package, nvidia-driver-pinning-<branch>, before the driver; NVIDIA writes that it suggests “installing the pinning package prior to installing the driver”. From branch 590 the Ubuntu package names no longer carry the branch, so an unpinned server moves to any newer branch you later import into the mirror. Whether to run R580 or R595 is covered in our guide to NVIDIA driver branches and CUDA versions.

On request we install the operating system, drivers, CUDA and a container runtime on AI servers we build, so the server arrives ready for your team. Tell us your distribution, kernel and driver branch in the form below.

Container Toolkit, images and Python packages

The NVIDIA Container Toolkit install guide of 18 September 2026 has no offline section. It pins four packages to one version, 1.20.1-1 at that date: nvidia-container-toolkit, nvidia-container-toolkit-base, libnvidia-container-tools and libnvidia-container1. Fetch all four from NVIDIA’s repository on the staging host and keep them at one version inside. The guide lists the NVIDIA driver as a prerequisite, so the driver goes in first.

Images go into a registry inside the gap. The skopeo sync command copies images between registries and local directories, which its manual calls “useful when synchronizing a local container registry mirror or for populating registries running inside of air-gapped environments”. Its --all option copies every image of a multi-architecture list instead of only the one for the current platform, and --preserve-digests makes the copy fail if a digest cannot be preserved. Deploy by digest rather than by tag; our article on securing a GPU server explains why images should come only from a registry you control.

A serving container such as vLLM’s, published on Docker Hub as vllm/vllm-openai, carries its Python packages inside the image. Python installed directly on the host needs a wheel directory instead. The pip user guide, version 26.2.1, describes installing “from local packages only, with no traffic to PyPI” in two steps: pip download with --destination-directory DIR and the requirements file on the staging host, then pip install with --no-index and --find-links=DIR inside.

GPU Operator in an air-gapped Kubernetes cluster

NVIDIA’s air-gapped guide for the GPU Operator, updated on 23 September 2026 for v26.7.1, covers four cases, from an HTTP proxy with full internet access to “Full Air-Gapped (w/o HTTP Proxy)”. The full case needs a local image registry and a local package repository inside the gap; NVIDIA adds that clusters able to run Precompiled Driver Containers do not need the package repository. The images are pulled on the connected side, tagged for the local registry and pushed there, the chart is fetched as an archive with helm fetch, and values.yaml sets the repository field of each component to the local registry.

The package repository exists because the Operator’s driver container, in NVIDIA’s words, “requires certain packages to be available”. NVIDIA names apt-mirror for copying them, and the repository list goes into a ConfigMap in the Operator’s namespace, referenced through driver.repoConfig.configMapName. A shorter route is to install the driver on the hosts from the local repository, as above, and set driver.enabled=false. The getting-started guide says this setting “prevents the Operator from installing the GPU driver on any nodes in the cluster”. The air-gapped guide does not discuss this case; by our reading, nothing in the cluster then needs the package repository. Our article on a private LLM platform on Kubernetes describes the layers above the Operator.

Model weights offline: Hugging Face, vLLM and NIM

Download the weights on the staging host at a fixed commit with hf download and --revision, adding --local-dir if you want a plain directory rather than the Hugging Face cache. Check them with hf cache verify, add them to the manifest and copy them into the model store inside. On the server, set HF_HUB_OFFLINE=1. Hugging Face’s documentation says that with this variable “no HTTP calls will be made to the Hugging Face Hub”; only cached files are used, and if a file is missing from the cache the library raises an error instead of trying to download it.

vLLM accepts a directory as its model, since its --model argument is the “Name or path of the Hugging Face model to use”. Point it at the model store and set --served-model-name, so that clients keep calling the same name when a new revision arrives under a new path. Leave --trust-remote-code off unless the model requires it, because trusted code from a model repository runs on the server.

NVIDIA NIM has its own procedure. Its air-gap guide, updated on 6 October 2026, splits the work into a network-connected phase, in which download-to-cache or create-model-store prepares the model files, and an air-gapped phase, in which the isolated host mounts them and starts the NIM with NIM_MODEL_PROFILE set to the profile ID, or with NIM_MODEL_PATH pointing to the model store. NVIDIA states: “In the air-gapped phase, do not set NGC_API_KEY or HF_TOKEN.” On Kubernetes, every image the NIM Helm chart references must be in a registry the cluster can reach. Model profiles and disk space are in our article on NVIDIA NIM requirements.

Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise. Describe the models you plan to run and how files enter your network in the form below.

Licences without internet: AI Enterprise and the DLS

On bare metal and in containers, no licence server takes part. NVIDIA’s licensing guide, updated on 2 September 2026, states that “NVIDIA AI Enterprise software runs with or without a valid license server connection” and that the NVIDIA License System is only required for vGPU for Compute drivers. The licence terms, for NIM in production for example, still apply.

Virtual machines with C-series compute profiles need a licence, and two methods work without internet. A Delegated License Service (DLS) instance is hosted on-premises. NVIDIA’s License System guide, version 3.6.1 of 29 June 2026, says that because a DLS instance is fully disconnected from the NVIDIA Licensing Portal, you “must download licenses from the NVIDIA Licensing Portal and upload them to the instance manually”. For a client with no network connection, the same guide describes node-locked licensing, with which such a client “can obtain a node-locked NVIDIA vGPU software license from a file installed locally”, supported since vGPU software 15.0. How licences are counted per GPU is in our guide to NVIDIA AI Enterprise licensing.

Updates and security patches on a disconnected server

On a server without internet access, every fix takes the staging path and needs a schedule. NVIDIA gives a production driver branch “Quarterly (or as-needed) bug and security releases for 1 year”, and a long term support branch the same for 3 years, on its driver lifecycle page of 9 September 2026. NVIDIA’s product security team publishes its bulletins on GitHub in Markdown, CSAF and CVE formats and, in parallel, on its Product Security website. NVIDIA advises customers to subscribe to notifications, and whoever runs the staging host should read them.

SOURCE TO WATCHWHAT IT CHANGESPUBLISHED IN
NVIDIA security bulletinsfixes for drivers, Container Toolkit and GPU OperatorNVIDIA Product Security page and GitHub
NVIDIA driver lifecyclesecurity releases and end of life per branchNVIDIA data-centre driver documentation
GPU Operator release notesbundled toolkit, device plugin and DCGM versionsNVIDIA cloud-native documentation
Distribution advisorieskernel, headers and base packagesthe distribution’s security notices
Model repositoriesnew revisions and changed access termseach model’s page on Hugging Face

NVIDIA Product Security page and driver lifecycle page, read on 10 October 2026; the other rows name where each component’s maintainers publish changes.

Import a new kernel together with its headers. Keep the previous driver repository, image digests and model revision in the mirrors until the new set has run on a test node, so that a rollback needs no new transfer. Update driver, toolkit and images as one tested set, since the images expect a CUDA version the driver supports.

What we supply

We build AI servers to order with the RTX PRO 6000 Server Edition, the H200 NVL, the L40S or the L4, assembled and burn-in tested, with manufacturer warranty on every component, on one EU contract and invoice. NVIDIA AI Enterprise and vGPU licences come on the same invoice. Models, serving and the platform on top are our Private AI/ML service, with engineering by our partner Vixen.UNO and support under an agreed SLA.

FAQ

How do you install an LLM on an air-gapped server?
Fetch every component on a connected staging host at an exact version: the NVIDIA driver repository, kernel headers, Container Toolkit packages, container images, model weights and any Python wheels. Verify them there, record a SHA-256 manifest, carry them across and check them against the manifest before importing them into internal mirrors. The server then installs only from those mirrors and serves the model from a local path, with HF_HUB_OFFLINE=1 set.
How do I install the NVIDIA driver offline?
On Ubuntu, NVIDIA’s driver installation guide offers a local repository package, installed with dpkg -i, followed by apt update, copying its keyring to /usr/share/keyrings/ and apt install nvidia-open; other distributions have their own pages in the guide. The kernel headers and DKMS come from the distribution’s repository, so they must be in your internal mirror. Install the pinning package for your branch before the driver, as NVIDIA suggests.
Can the NVIDIA GPU Operator run in an air-gapped Kubernetes cluster?
Yes. NVIDIA’s air-gapped guide for GPU Operator v26.7.1 uses a local image registry, a chart fetched with helm fetch and values.yaml pointing at that registry, plus a local package repository for the driver container. If the driver is installed on the hosts instead, driver.enabled=false stops the Operator from installing a driver on any node.
How do I use Hugging Face models offline?
Download the model on a connected machine with hf download at a fixed revision, verify it with hf cache verify and its fail-on-missing-files option, and copy it to the offline server. Set HF_HUB_OFFLINE=1 there, with which the library makes no HTTP calls to the Hub and raises an error if a file is not cached. vLLM can then serve the model from its local directory.
Does NVIDIA AI Enterprise need internet access for licensing?
Not on bare metal or in containers: NVIDIA states that AI Enterprise software runs with or without a valid licence server connection. Only vGPU for Compute drivers need the NVIDIA License System, which works offline through an on-premises DLS instance whose licences are downloaded from the portal and uploaded manually, or through a node-locked licence file on the client.
How do you keep an air-gapped GPU server updated?
Every update takes the same staging path as the installation, so someone has to follow NVIDIA’s security bulletins, the driver branch lifecycle, the GPU Operator release notes, the distribution’s advisories and new model revisions. NVIDIA issues bug and security releases quarterly or as needed for each supported branch. Test each new set of driver, toolkit and images on one node and keep the previous set in the mirrors for a rollback.

Send us your Linux distribution and kernel, the driver branch, the container platform, the models you plan to run and how files reach the isolated network today. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna