Batch LLM inference on-premise: processing document archives overnight with vLLM
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Batch LLM inference is sized by throughput and by the hour the job must finish, not by time to first token: the engine takes every request at once, runs large batches and fills GPU memory with KV cache
- vLLM runs batch jobs through its LLM class (
llm.generateover a list of prompts) or throughvllm run-batch, which reads an OpenAI batch-format JSONL file and supports the chat completions, embeddings and score endpoints - In MLPerf® Inference v6.0 (entry 6.0-0005), eight RTX PRO 6000 Server Edition cards produced 48,613.8 tokens per second on Llama 3.1 8B offline, which at an assumed 600 tokens per summary is about 7 hours, or 55 GPU-hours, for 2 million documents by our arithmetic
- Under the interactive latency limits the same server produced 6,262.6 tokens per second on Llama 2 70B against 27,730.1 offline, less than a quarter, which is why batch work runs apart from chat
- Kueue lets one Kubernetes cluster run chat by day and batch Jobs at night on the same GPUs: a low WorkloadPriorityClass for batch,
LowerPrioritypreemption in the ClusterQueue and input split into shards so an evicted job loses one shard
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Batch LLM inference: throughput instead of latency
Batch LLM inference runs a model over a fixed set of inputs, such as a document archive, with no user waiting for each answer. It is sized by throughput, the tokens or documents processed per hour, and by the hour by which the job has to be done, not by time to first token. The engine therefore receives all requests at once, runs the largest batches that GPU memory allows and keeps the cards busy until the queue is empty. On-premise, these jobs fit the night hours of GPUs that serve an assistant during the day.
Published results show what latency limits cost in throughput. In MLPerf® Inference v6.0, entry 6.0-0005, a server with eight RTX PRO 6000 Server Edition cards produced 27,730.1 tokens per second on Llama 2 70B in the offline scenario, where all queries are sent at once. Under the latency limits of the interactive scenario the same server produced 6,262.6, less than a quarter. Our guide to LLM latency targets sets targets for chat, RAG and agents; a batch job’s target is the number of documents finished by morning.
Batch and interactive settings in vLLM
| SETTING | INTERACTIVE CHAT | OVERNIGHT BATCH | SOURCE |
|---|---|---|---|
| Target | TTFT and TPOT at the 95th percentile | documents per hour, end time of the job | our examples |
| Entry point | vllm serve, OpenAI-compatible API | LLM. | vLLM quickstart, run-batch docs |
| max_ | small, such as 2048, for lower ITL | above 8192 for throughput | vLLM optimisation guide |
| gpu_ | default 0.92 | higher for more KV cache, if no other model shares the card | vLLM engine arguments, optimisation guide |
| Parallelism | tensor parallel when one card is too small or too slow | data parallel copies when the model fits one card | vLLM optimisation guide, NIM profiles |
| Output | streamed text | JSON schema or a fixed list of choices | vLLM structured outputs |
| NIM profile | latency | throughput | NIM for LLMs 1.14.0 profiles |
vLLM documentation (latest): quickstart and run-batch example, both dated 10 October 2026; optimisation guide, 20 August 2026; engine arguments; structured outputs, 19 May 2026. NVIDIA NIM for LLMs 1.14.0, model profiles, updated 23 July 2026. The interactive targets are the examples from our latency guide.
vLLM’s optimisation guide states that smaller values of max_num_batched_tokens “achieve better ITL because there are fewer prefills slowing down decodes”, and recommends values above 8192 “for optimal throughput … especially for smaller models on large GPUs”. The guide also explains that when KV cache space runs short, vLLM preempts requests, and that vLLM V1 recomputes them by default. Recomputation costs throughput, and the guide’s remedies are a higher gpu_memory_utilization or a lower max_num_seqs.
Set --max-model-len to the longest document plus the answer. If it is not set, vLLM derives it “from the model config”, which can be far longer than a document needs. With --kv-cache-dtype left at “auto”, the cache uses the model’s data type; an FP8 cache takes half the bytes of a 16-bit one and fits about twice as many documents in flight.
Put the instructions that every request shares at the start of the prompt and the document at the end. vLLM’s automatic prefix caching “caches the KV cache of existing queries”, so a new query with the same prefix can “skip the computation of the shared part”. The documentation adds that it only shortens prefill and “does not reduce the time of generating new tokens”, so it helps long shared instructions, not long answers. Some vLLM recipes switch it off with no-enable-prefix-caching for benchmarks, so check that a configuration copied from a recipe keeps it on.
vLLM offline inference, the batch runner and NIM
vLLM offers two ways to run a batch without a client. With its Python class LLM, the quickstart describes “offline batch inferencing” over a list of prompts, where llm.generate “adds the input prompts to the vLLM engine’s waiting queue” and generates the outputs “with high throughput”. For a Ray cluster of several servers, vLLM’s offline inference page points to Ray Data LLM, which states that “Continuous batching keeps vLLM replicas saturated and maximizes GPU utilization.”
The second needs no code. vllm run-batch reads a JSONL file in the OpenAI batch format, and the vLLM example calls it “batch inference using the OpenAI batch file format, not the complete Batch (REST) API”. “Each line represents a separate request”, with a custom_id, the method, the endpoint as url and the request body. The supported endpoints are /v1/chat/completions, /v1/embeddings and /v1/score. Input and output can be local files or HTTP(S) URLs, the output written by HTTP PUT. A run reads vllm run-batch -i in.jsonl -o out.jsonl plus --model and the engine settings above.
NVIDIA NIM for LLMs selects an engine configuration by profile. Its profile page for release 1.14.0 says throughput profiles are designed to maximise total throughput per GPU and that “Latency variants use additional GPUs to decrease request latency at the cost of decreased total throughput per GPU relative to the throughput variant.” A profile is chosen with NIM_MODEL_PROFILE. The page does not mention batch files, so a batch job sends its requests to the NIM API from a client that keeps a fixed number of requests in flight.
Classification, extraction, summaries and metadata for RAG
The task decides how many tokens each document produces, and with them how long the job runs. vLLM’s structured outputs constrain the answer: with choice, “the output will be exactly one of the choices”, and a JSON schema can be passed as response_format through the API or as StructuredOutputsParams offline. Classification then returns a label of a few tokens, so the cost of each document lies mostly in reading it. Extraction returns a JSON object with dates, parties, amounts or ticket fields, tens to a few hundred tokens. A summary is the longest output, several hundred tokens per document.
Metadata enrichment for RAG combines these. The model assigns a document type, a title, a date and keywords to each document or chunk, and the retrieval index stores them as filters. Where the enrichment adds a title or summary to the chunk text, the chunks are embedded again, which vllm run-batch covers through its embeddings endpoint, while sizing the embedding server is the subject of our guide to embedding and reranker servers. Scanned pages need text first, and the GPU cost of reading pages with a vision-language model is covered in our article on document AI with vision-language models.
GPU-hours for a document archive: method and example
The GPU-hours of a batch job follow from the number of documents, the output tokens per document and a throughput figure for the model and the cards.
- Count the documents and take a sample of about a thousand that represents the archive.
- Fix the task and the output length per document with
max_tokensand, where it fits, a schema. - Take a published throughput for a comparable model and server, or time the sample as a batch job;
vllm bench throughputwith--input-lenand--output-lenset to the sample’s averages gives a first estimate. - Divide documents × output tokens by tokens per second, then by 3,600 for hours, and multiply by the number of cards for GPU-hours.
- Compare the hours with the night window and keep a margin before you decide on the number of servers.
MLPerf Inference v6.0 publishes offline throughput for two models that suit this method. The Llama 3.1 8B benchmark summarises news articles from the CNN/DailyMail set, 13,368 of them in the data-centre set, “with an average input length of 778 tokens and output length of 73 tokens”, as MLCommons describes the dataset, and the reference implementation caps each summary at 128 tokens. Summaries of business documents are longer, so the example below assumes 600 output tokens per summary. The Llama 2 70B benchmark answers questions with about 274 tokens per answer in the same entry. The example below is an archive of 2 million documents.
| JOB, 2M DOCUMENTS | OUTPUT TOKENS | 8 CARDS, TOK/S | HOURS, 1 SERVER | GPU-HOURS |
|---|---|---|---|---|
| Summaries, 8B, RTX PRO 6000 | 600 per document, assumed | 48,613.8 | 6.9 | 55 |
| Summaries, 8B, H200 NVL | 600 per document, assumed | 53,193.6 | 6.3 | 50 |
| Answers, 70B, RTX PRO 6000 | 274 per document | 27,730.1 | 5.5 | 44 |
| Answers, 70B, H200 NVL | 274 per document | 32,004.0 | 4.8 | 38 |
| Summaries, 70B, RTX PRO 6000 | 600, rate assumed | 27,730.1 | 12.0 | 96 |
MLPerf Inference v6.0, data-centre suite, closed division, available systems, offline scenario, Llama 3.1 8B and Llama 2 70B (99 per cent accuracy variant). Entry 6.0-0005, eight RTX PRO 6000 Server Edition, FP4 weights, TensorRT 10.14, retrieved from MLCommons’ v6.0 results repository on 10 October 2026; entry 6.0-0021, Dell PowerEdge XE7740 with eight H200 NVL, FP8 weights, retrieved from mlcommons.org on 24 September 2026 for our H200 NVL benchmark review; results verified by MLCommons Association. The 274 tokens per answer are from entry 6.0-0005 and applied to both cards; the 600 tokens per summary are our assumption. Hours and GPU-hours are our arithmetic, not MLPerf metrics; the last row assumes the 70B rate holds for longer answers, which the benchmark does not show. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information.
These rates come from tuned TensorRT runs on the benchmark’s own inputs. A 30-page contract is many times longer than a news article, so prefill takes a larger share of the time, and vLLM on your documents gives a different figure. This example plans with half the published rate until a test run on the sample gives a measured one, a planning margin, not a vendor figure. At that margin, 2 million summaries with an 8B model take about 14 hours on one server with eight RTX PRO 6000, which is two nights or one night on two servers. The H200 NVL figures and the system behind them are in our H200 NVL benchmark review.
We build GPU servers to order around the workload and check the rack, power and airflow before we quote. Send us your document count, task and night window through the form below, and you receive a configuration within one business day.
Night windows on the chat GPUs with Kubernetes and Kueue
Kueue describes itself as “a kubernetes-native system that manages quotas and how jobs consume them”, and it decides when a job waits, when it starts and when it is preempted. Its stable integrations include the batch Job and the Deployment, so the chat service and the night job can share one ClusterQueue whose quota covers the GPUs. A CronJob starts the batch Job each evening; Kueue’s documentation asks for the kueue.x-k8s.io/queue-name label in the job template and says Kueue “automatically manages the Job’s suspension” and decides when it starts. The same page notes that concurrencyPolicy defaults to Allow and startingDeadlineSeconds to no deadline, so set both to fit the window.
Priorities decide which of the two gives way. A WorkloadPriorityClass, set on a Job with the label kueue.x-k8s.io/priority-class, is “independent from Pod’s priority” and is used for ordering and for deciding whether a workload can preempt others. Give the chat service the higher class and the batch Job the lower one, and set withinClusterQueue to LowerPriority in the ClusterQueue, since the default Never preempts nothing. When the chat deployment scales up in the morning and needs its GPUs back, the batch workload is evicted with the reason Preempted. Kueue preempts the whole workload, all pods of the Job at once.
Split the archive into shard files and record which shards are finished, so an evicted or failed job restarts with the next shard and loses at most one. With the annotation kueue.x-k8s.io/job-min-parallelism, Kueue can admit a Job with fewer parallel pods than requested when GPUs are short, without changing the number of completions. For example, two servers with eight RTX PRO 6000 each run an assistant model as several one-card copies by day; at night a scheduled scale-down, which Kueue does not perform, leaves one copy per server and the batch Job takes the other fourteen cards. Quotas per department and showback of GPU-hours are the subject of our guide to a shared GPU platform for departments.
Our inference servers carry 2 to 8 GPUs per node, sized by model size and concurrent users. Describe your daytime chat load and the night batch job in the form below, and we size one platform for both.
What we supply
We build AI servers to order for private LLMs and RAG in production, with 2 to 8 GPUs per node, assembled and burn-in tested, with manufacturer warranty on every component and delivery anywhere in the EU, on one EU contract and invoice. The cards include the RTX PRO 6000 Server Edition, the H200 NVL with NVLink bridges, the L40S and the L4, listed on our professional GPU page, and NVIDIA AI Enterprise and vGPU licences come on the same invoice. Operating system, drivers, CUDA and a container runtime are installed on request. The batch pipeline, the models and RAG on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What is batch inference for LLMs?
How do I run offline batch inference with vLLM?
llm.generate from vLLM’s Python class LLM with a list of prompts, which queues them all and generates the outputs with high throughput, or run vllm run-batch with a JSONL file in the OpenAI batch format. The batch runner supports the chat completions, embeddings and score endpoints and reads and writes local files or HTTP(S) URLs. For throughput, set max_num_batched_tokens above 8192 and a max-model-len that fits your longest document.How long does it take to process a million documents with an LLM?
Can the same GPUs serve chat by day and batch jobs at night?
withinClusterQueue: LowerPriority lets the chat service take its GPUs back in the morning by evicting the batch workload. Splitting the input into shards limits the loss to one shard per eviction.Which GPU is better for overnight LLM batch processing?
What settings raise vLLM throughput for batch jobs?
max_num_batched_tokens above 8192 for throughput and a higher gpu_memory_utilization for more KV cache, and data-parallel copies when the model fits one card. An FP8 KV cache fits about twice as many documents in flight as a 16-bit one, and automatic prefix caching shortens prefill when every prompt starts with the same instructions. Constraining the output with a JSON schema or a list of choices keeps answers short.Send us the number and kind of documents, the task per document, the model you have in mind, your daytime chat load and the hours of your night window. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day