Air-gapped LLM server: installing NVIDIA drivers, containers and models offline
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Everything the server would download, from the driver, kernel headers and Container Toolkit to images, model weights and Python wheels, is fetched at exact versions on a connected staging host, verified, carried across with a SHA-256 manifest and installed from internal mirrors
- NVIDIA’s driver installation guide of 23 September 2026 offers a local repository package for Ubuntu; the kernel headers come from the distribution’s repository, and the branch pinning package goes in before the driver
- The GPU Operator’s air-gapped guide for v26.7.1 needs a local image registry and, for its driver container, a local package repository; with the driver installed on the hosts, driver.enabled=false leaves that container out
- Hugging Face weights are downloaded at a fixed revision and checked with hf cache verify on the connected side; on the server, HF_HUB_OFFLINE=1 stops all HTTP calls to the Hub and vLLM serves the model from a local path
- NVIDIA AI Enterprise software runs with or without a licence server; only vGPU for Compute needs the NVIDIA License System, offline through an on-premises DLS instance with manually uploaded licences or a node-locked licence file
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
What an air-gapped LLM server needs
An air-gapped LLM server needs everything it would otherwise download to be fetched on a connected staging host, checked there, carried across and installed from local copies. That covers the NVIDIA driver and the kernel headers it builds against, the container runtime, container images, model weights, Python packages and, for vGPU for Compute only, licences. Versions and dates in this article are as read on 10 October 2026.
| COMPONENT | OFFLINE METHOD | DOCUMENTED IN |
|---|---|---|
| NVIDIA driver (Ubuntu) | local repository package and branch pinning package | NVIDIA driver installation guide, 23 September 2026 |
| Kernel headers, OS packages | mirror of the distribution repository, for example with apt-mirror | NVIDIA driver installation guide; GPU Operator air-gapped guide, 23 September 2026 |
| Container Toolkit | four packages at one pinned version | NVIDIA Container Toolkit install guide, 18 September 2026 |
| Container images | internal registry, filled with skopeo sync or docker pull, tag and push | skopeo-sync manual; GPU Operator air-gapped guide |
| GPU Operator | chart from helm fetch, local registry and package repository in values.yaml | GPU Operator air-gapped guide |
| Model weights | hf download at a fixed revision, local path, HF_ | Hugging Face Hub documentation; vLLM engine arguments |
| NIM containers | model cache or model store prepared on the connected side | NVIDIA NIM air-gap guide, 6 October 2026 |
| Python packages | pip download, then pip install with --no-index and --find-links | pip user guide, version 26.2.1 |
| Licences | no licence server on bare metal; DLS or node-locked file for vGPU for Compute | NVIDIA licensing guide; NVIDIA License System guide 3.6.1 |
Sources: NVIDIA, Hugging Face, vLLM, pip and skopeo documentation, read on 10 October 2026; dates are those shown on each page.
Staging host, transfer path and checksums
The staging host sits outside the air gap and holds the mirrors: a copy of the distribution repository, NVIDIA’s driver and toolkit packages, an image registry or image directory, the model store and a wheel directory. Give it the same operating system release and processor architecture as the servers inside, because packages, kernel headers and Python wheels are built per release and architecture. Files cross through one path, removable media or a one-way transfer system as your security policy allows, and are checked on both sides.
- Download each component on the staging host at an exact version: driver version, toolkit version, image digest and model commit.
- Verify it there against what the publisher provides: signed metadata for packages from a network repository,
hf cache verifyfor Hugging Face weights. - Write a manifest with the SHA-256 value of every file that will cross.
- Copy the files and the manifest to the transfer medium.
- On the inside, check every file against the manifest with
sha256sum -cbefore it enters a mirror. - Import the files into the internal mirrors and install or update the servers from those mirrors only.
Hugging Face describes hf cache verify as a way to “validate local files against their checksums on the Hub”. It needs the Hub, so it runs on the staging host. On a mismatch it prints a list of the files and exits with a non-zero status, but a missing file only produces a warning unless you add --fail-on-missing-files. Keep every manifest, since it records which versions reached the inside and when, and a restore needs the same record, as our article on backing up an AI server explains.
Installing the NVIDIA driver offline on Ubuntu
NVIDIA’s driver installation guide, updated on 23 September 2026 and written for branch 615, gives Ubuntu two methods: local repository enablement and network repository enablement. The steps below are for Ubuntu only; Red Hat Enterprise Linux, SUSE, Debian and other distributions have their own pages in the same guide. The repository comes as one Debian package whose name starts with nvidia-driver-local-repo- and continues with the distribution, the driver version and the architecture. The guide installs it with dpkg -i, runs apt update and enrols the “ephemeral” GPG key it carries by copying its keyring file to /usr/share/keyrings/; apt install nvidia-open then installs the open kernel modules. Since the key comes inside the package, record the package’s SHA-256 value on the staging host when you download it from NVIDIA.
The kernel headers come from outside that package. The guide’s preparation step installs them for the running kernel with apt install linux-headers-$(uname -r) from the distribution’s own repository, which also supplies DKMS, the framework the nvidia-dkms-open package uses to build the modules. The internal mirror must therefore hold the headers for every kernel your servers boot. Install the pinning package, nvidia-driver-pinning-<branch>, before the driver; NVIDIA writes that it suggests “installing the pinning package prior to installing the driver”. From branch 590 the Ubuntu package names no longer carry the branch, so an unpinned server moves to any newer branch you later import into the mirror. Whether to run R580 or R595 is covered in our guide to NVIDIA driver branches and CUDA versions.
On request we install the operating system, drivers, CUDA and a container runtime on AI servers we build, so the server arrives ready for your team. Tell us your distribution, kernel and driver branch in the form below.
Container Toolkit, images and Python packages
The NVIDIA Container Toolkit install guide of 18 September 2026 has no offline section. It pins four packages to one version, 1.20.1-1 at that date: nvidia-container-toolkit, nvidia-container-toolkit-base, libnvidia-container-tools and libnvidia-container1. Fetch all four from NVIDIA’s repository on the staging host and keep them at one version inside. The guide lists the NVIDIA driver as a prerequisite, so the driver goes in first.
Images go into a registry inside the gap. The skopeo sync command copies images between registries and local directories, which its manual calls “useful when synchronizing a local container registry mirror or for populating registries running inside of air-gapped environments”. Its --all option copies every image of a multi-architecture list instead of only the one for the current platform, and --preserve-digests makes the copy fail if a digest cannot be preserved. Deploy by digest rather than by tag; our article on securing a GPU server explains why images should come only from a registry you control.
A serving container such as vLLM’s, published on Docker Hub as vllm/vllm-openai, carries its Python packages inside the image. Python installed directly on the host needs a wheel directory instead. The pip user guide, version 26.2.1, describes installing “from local packages only, with no traffic to PyPI” in two steps: pip download with --destination-directory DIR and the requirements file on the staging host, then pip install with --no-index and --find-links=DIR inside.
GPU Operator in an air-gapped Kubernetes cluster
NVIDIA’s air-gapped guide for the GPU Operator, updated on 23 September 2026 for v26.7.1, covers four cases, from an HTTP proxy with full internet access to “Full Air-Gapped (w/o HTTP Proxy)”. The full case needs a local image registry and a local package repository inside the gap; NVIDIA adds that clusters able to run Precompiled Driver Containers do not need the package repository. The images are pulled on the connected side, tagged for the local registry and pushed there, the chart is fetched as an archive with helm fetch, and values.yaml sets the repository field of each component to the local registry.
The package repository exists because the Operator’s driver container, in NVIDIA’s words, “requires certain packages to be available”. NVIDIA names apt-mirror for copying them, and the repository list goes into a ConfigMap in the Operator’s namespace, referenced through driver.. A shorter route is to install the driver on the hosts from the local repository, as above, and set driver.enabled=false. The getting-started guide says this setting “prevents the Operator from installing the GPU driver on any nodes in the cluster”. The air-gapped guide does not discuss this case; by our reading, nothing in the cluster then needs the package repository. Our article on a private LLM platform on Kubernetes describes the layers above the Operator.
Model weights offline: Hugging Face, vLLM and NIM
Download the weights on the staging host at a fixed commit with hf download and --revision, adding --local-dir if you want a plain directory rather than the Hugging Face cache. Check them with hf cache verify, add them to the manifest and copy them into the model store inside. On the server, set HF_HUB_OFFLINE=1. Hugging Face’s documentation says that with this variable “no HTTP calls will be made to the Hugging Face Hub”; only cached files are used, and if a file is missing from the cache the library raises an error instead of trying to download it.
vLLM accepts a directory as its model, since its --model argument is the “Name or path of the Hugging Face model to use”. Point it at the model store and set --served-model-name, so that clients keep calling the same name when a new revision arrives under a new path. Leave --trust-remote-code off unless the model requires it, because trusted code from a model repository runs on the server.
NVIDIA NIM has its own procedure. Its air-gap guide, updated on 6 October 2026, splits the work into a network-connected phase, in which download-to-cache or create-model-store prepares the model files, and an air-gapped phase, in which the isolated host mounts them and starts the NIM with NIM_MODEL_PROFILE set to the profile ID, or with NIM_MODEL_PATH pointing to the model store. NVIDIA states: “In the air-gapped phase, do not set NGC_API_KEY or HF_TOKEN.” On Kubernetes, every image the NIM Helm chart references must be in a registry the cluster can reach. Model profiles and disk space are in our article on NVIDIA NIM requirements.
Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise. Describe the models you plan to run and how files enter your network in the form below.
Licences without internet: AI Enterprise and the DLS
On bare metal and in containers, no licence server takes part. NVIDIA’s licensing guide, updated on 2 September 2026, states that “NVIDIA AI Enterprise software runs with or without a valid license server connection” and that the NVIDIA License System is only required for vGPU for Compute drivers. The licence terms, for NIM in production for example, still apply.
Virtual machines with C-series compute profiles need a licence, and two methods work without internet. A Delegated License Service (DLS) instance is hosted on-premises. NVIDIA’s License System guide, version 3.6.1 of 29 June 2026, says that because a DLS instance is fully disconnected from the NVIDIA Licensing Portal, you “must download licenses from the NVIDIA Licensing Portal and upload them to the instance manually”. For a client with no network connection, the same guide describes node-locked licensing, with which such a client “can obtain a node-locked NVIDIA vGPU software license from a file installed locally”, supported since vGPU software 15.0. How licences are counted per GPU is in our guide to NVIDIA AI Enterprise licensing.
Updates and security patches on a disconnected server
On a server without internet access, every fix takes the staging path and needs a schedule. NVIDIA gives a production driver branch “Quarterly (or as-needed) bug and security releases for 1 year”, and a long term support branch the same for 3 years, on its driver lifecycle page of 9 September 2026. NVIDIA’s product security team publishes its bulletins on GitHub in Markdown, CSAF and CVE formats and, in parallel, on its Product Security website. NVIDIA advises customers to subscribe to notifications, and whoever runs the staging host should read them.
| SOURCE TO WATCH | WHAT IT CHANGES | PUBLISHED IN |
|---|---|---|
| NVIDIA security bulletins | fixes for drivers, Container Toolkit and GPU Operator | NVIDIA Product Security page and GitHub |
| NVIDIA driver lifecycle | security releases and end of life per branch | NVIDIA data-centre driver documentation |
| GPU Operator release notes | bundled toolkit, device plugin and DCGM versions | NVIDIA cloud-native documentation |
| Distribution advisories | kernel, headers and base packages | the distribution’s security notices |
| Model repositories | new revisions and changed access terms | each model’s page on Hugging Face |
NVIDIA Product Security page and driver lifecycle page, read on 10 October 2026; the other rows name where each component’s maintainers publish changes.
Import a new kernel together with its headers. Keep the previous driver repository, image digests and model revision in the mirrors until the new set has run on a test node, so that a rollback needs no new transfer. Update driver, toolkit and images as one tested set, since the images expect a CUDA version the driver supports.
What we supply
We build AI servers to order with the RTX PRO 6000 Server Edition, the H200 NVL, the L40S or the L4, assembled and burn-in tested, with manufacturer warranty on every component, on one EU contract and invoice. NVIDIA AI Enterprise and vGPU licences come on the same invoice. Models, serving and the platform on top are our Private AI/ML service, with engineering by our partner Vixen.UNO and support under an agreed SLA.
FAQ
How do you install an LLM on an air-gapped server?
How do I install the NVIDIA driver offline?
Can the NVIDIA GPU Operator run in an air-gapped Kubernetes cluster?
How do I use Hugging Face models offline?
Does NVIDIA AI Enterprise need internet access for licensing?
How do you keep an air-gapped GPU server updated?
Send us your Linux distribution and kernel, the driver branch, the container platform, the models you plan to run and how files reach the isolated network today. We reply within one business day with a configuration and quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day