BLOG · GUIDE ·

Vision-language models for document AI: GPUs for OCR, invoices and contracts on-premise

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A vision-language model reads each page as image tokens: Qwen3.8-27B makes one token per 32 × 32 pixels, about 2,145 for an A4 page at 150 dpi, while Gemma 4 uses fixed budgets of 70 to 1,120 tokens and DeepSeek-OCR-2 256 to 1,120
  • Open OCR models of 0.9B to 8.3B parameters have weights of 2.7 to 10.6 GB in their published files and fit a 24 GB L4 or RTX PRO 4000; general models such as Qwen3.8-27B (30.9 GB in FP8) need a 96 GB RTX PRO 6000 or an H200 NVL for many pages in flight
  • Published throughput is scarce: Datalab gives 1.44 pages per second for its 5B chandra-ocr-2 on one H100 80GB with vLLM, and Z.ai 1.86 pages per second for GLM-OCR on PDFs without naming the GPU
  • By our estimate, scaled from that H100 figure by the lower of memory bandwidth and compute, 10,000 pages in an 8-hour window need one RTX PRO 5000 with a 5B OCR model, and 100,000 pages need three H200 NVL, or two RTX PRO 6000 over 24 hours
  • Licences differ by model: GLM-OCR and Unlimited-OCR are MIT, DeepSeek-OCR-2, olmOCR-2, Qwen3.8-27B and Gemma 4 Apache 2.0, chandra-ocr-2 a modified OpenRAIL-M with a revenue threshold, and Mistral lists its OCR models as Premier with no open licence

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

What a vision-language model needs from a GPU for documents

The GPU a vision-language model needs for document AI on-premise depends on its weights and on the image tokens of every page in flight. A vision-language model (VLM) turns each page image into image tokens with a vision encoder, and its language model then writes the text, Markdown or fields. GPU memory therefore holds the weights, including the encoder, plus a cache for the image tokens and the generated text of every page in flight. Document-specialised OCR models of 0.9B to 8.3B parameters fit a 24 GB card such as the L4 or RTX PRO 4000. General VLMs such as Qwen3.8-27B or Gemma 4 31B need a 96 GB RTX PRO 6000 or a 141 GB H200 NVL once many pages run at the same time.

The model facts below come from the model cards and files on Hugging Face and the vLLM documentation, as read on 10 October 2026. The configurations are our estimates, with the method shown. For text-only models, our LLM hardware requirements by model covers the same arithmetic.

Image tokens per page, as each model counts them

The cost of a page depends on how many tokens the encoder makes from it, and each model family counts differently.

Qwen3.8-27B, which Qwen’s card describes as “a native vision-language model that understands images and videos”, uses the Qwen3-VL processor. Its preprocessor_config.json sets a patch size of 16 and a merge size of 2, so one token covers a 32 × 32 pixel block. The same file allows 65,536 to 16,777,216 pixels per image, which is 64 to 16,384 tokens. An A4 page rendered at 150 dpi (1,240 × 1,754 pixels) comes to about 2,145 tokens by our arithmetic, and at 200 dpi to about 3,796. The render resolution you choose is therefore the main setting for cost per page.

Gemma 4 uses fixed budgets instead. Its card states that “The supported token budgets are: 70, 140, 280, 560, and 1120”, and advises higher budgets “for tasks like OCR, document parsing, or reading small text”. DeepSeek-OCR-2 makes 256 tokens for a 1,024 × 1,024 overview plus 144 for each of up to six 768 × 768 crops, so 256 to 1,120 per page. olmOCR-2 renders pages so that “the longest dimension is 1288 pixels”, and its Qwen2.5-VL processor (patch 14, merge 2) makes about 1,500 tokens from an A4 page by our arithmetic.

MODELPARAMETERSIMAGE TOKENS/PAGEWEIGHTSLICENCE
GLM-OCR (Z.ai)1.3B in files, 0.9B by cardnot stated2.7 GB, BF16MIT
Nemotron Parse 2.00.9Bnot stated; input 1,024 × 1,280 to 1,664 × 2,0483.6 GB, F32 filesOpenMDW 1.1
Unlimited-OCR (Baidu)3.3Bnot stated6.7 GB, BF16MIT
DeepSeek-OCR-23.4B256 to 1,1206.8 GB, BF16Apache 2.0
chandra-ocr-2 (Datalab)5.3Bnot stated10.6 GB, BF16modified OpenRAIL-M
olmOCR-2-7B FP8 (Ai2)8.3Babout 1,500 (our arithmetic)10.1 GB, FP8 and BF16Apache 2.0
Qwen3.8-27B27B64 to 16,384; A4 at 150 dpi about 2,14530.9 GB FP8; 55.6 GB BF16Apache 2.0
Gemma 4 31B30.7Bbudget of 70 to 1,12062.6 GB BF16; 23.3 GB QATApache 2.0

Hugging Face model cards, preprocessor_config.json files and the safetensors parameter counts of each repository, read on 10 October 2026; weights computed from the parameters per data type. Qwen3.8-27B and Gemma 4 sizes as in our Qwen and Gemma guides.

NVIDIA’s Nemotron Nano 12B v2 VL, released on 28 October 2025, is a general VLM listed at 13B parameters on Hugging Face. Its card reports 85.4% on OCRBench for the FP8 version and names the H100 SXM 80GB as tested hardware. Mistral’s model overview lists OCR 4.1 and OCR 3 as Premier models with no open licence, so we leave them out of on-premise sizing.

Vision encoder, image cache and pages in flight

The vision encoder is part of the weights and small next to the language model. Google’s card gives about 550M parameters for the encoder of Gemma 4 31B and 26B A4B, and about 150M for E2B and E4B.

vLLM sizes memory for images in three places. Its documentation, updated on 9 October 2026, says that “Encoder cache size is determined by the actual inputs at runtime”. limit_mm_per_prompt sets how many images or videos one request may carry, and for an image-only service “there is no need to allocate any memory for videos”. A processor cache in host memory, set with mm_processor_cache_gb and 4 GiB by default, holds processed inputs; the vLLM recipes for DeepSeek-OCR and Unlimited-OCR set it to 0 and turn prefix caching off, which the DeepSeek-OCR recipe recommends “to avoid unnecessary hashing and caching”. For Qwen2-VL models the documentation shows a max_pixels setting in mm_processor_kwargs that caps the tokens per image.

The KV cache then holds the image tokens and the generated text of every page in flight. Our Qwen hardware guide puts Qwen3.8-27B at 64 KiB per token in 16-bit plus a fixed 144 MiB per sequence. A page of 2,145 image tokens plus an assumed 1,500 output tokens, our estimate for a dense text page in Markdown, takes about 0.36 GiB. With the FP8 weights on one RTX PRO 6000, that is room for about 150 pages in flight. Qwen3.8 runs in thinking mode by default, so switch thinking off for transcription, or the reasoning tokens add to every page.

Gemma 4 31B keeps 80 KiB per token in its full-attention layers and a fixed 0.78 GiB in its sliding-window layers, as our Gemma hardware guide sets out. A page at the 1,120-token budget with the same output takes about 1 GiB, so one RTX PRO 6000 holds about 25 pages with BF16 weights and about 53 with FP8 weights, both with a 16-bit cache, by our estimate. The small OCR models need far less, and a 24 GB card holds their weights with room for many pages.

Published throughput in pages per second

Few model authors publish pages per second, and none of the figures we found names a card we supply.

Datalab’s card for chandra-ocr-2 reports 1.44 pages per second on one H100 80GB with vLLM and 96 concurrent sequences, on documents from the olmOCR benchmark set. It adds an estimate of about 2 pages per second for typical documents, because that set is slower. The card names neither the H100 version, SXM or PCIe, nor the vLLM version, nor a date for the measurement. Z.ai’s card states that GLM-OCR “achieves a throughput of 1.86 pages/second for PDF documents and 0.67 images/second for images”, at a single concurrency, without naming the GPU. DeepSeek’s repository lists about 2,500 tokens per second for DeepSeek-OCR’s PDF script on an A100-40G.

As of 10 October 2026 we found no published pages-per-second figure for Qwen3.8-27B, Gemma 4 or Nemotron Parse 2.0. Measure your own documents in a pilot before you size more than one server.

GPUs for 1,000, 10,000 and 100,000 pages a day

Our method has three steps. Divide the pages per day by the processing window, 8 hours (28,800 seconds) or 24 hours, to get the rate the server must sustain. Scale Datalab’s 1.44 pages per second to each card by the lower of two ratios to the H100: memory bandwidth, which limits writing the text, and BF16 Tensor Core throughput, which limits reading the image tokens. Datalab does not say which H100 80GB it used, so we take the SXM version with 3.35 TB/s and 1,979 BF16 teraFLOPS with sparsity, which gives the lower estimates. For a 27B VLM on every page we divide the result by five, the ratio of parameters to the 5.3B of chandra-ocr-2.

By this method one L4 handles about 0.13 pages per second with a 5B OCR model and one RTX PRO 6000 Server Edition about 0.69, both limited by memory bandwidth. The RTX PRO 5000 and the H200 NVL are limited by compute. NVIDIA lists 1,671 BF16 teraFLOPS with sparsity for the H200 NVL, below the H100 SXM, so despite its 4.8 TB/s it comes to about 1.2 pages per second, and the RTX PRO 5000 to about 0.38. With Qwen3.8-27B, one RTX PRO 6000 handles about 0.14 pages per second and one H200 NVL about 0.24.

PAGES PER DAYRATE NEEDED5B OCR MODEL27B VLM, EVERY PAGE
1,0000.035 pages/s in 8 hone L4 or RTX PRO 4000one RTX PRO 6000, or one DGX Spark over 24 h
10,0000.35 pages/s in 8 h; 0.12 in 24 hone RTX PRO 5000, or one L4 over 24 htwo H200 NVL, or one RTX PRO 6000 over 24 h
100,0003.5 pages/s in 8 h; 1.2 in 24 hthree H200 NVL, or two RTX PRO 6000 over 24 hfive H200 NVL over 24 h; about 15 H200 NVL in 8 h

Our estimates, not measurements: chandra-ocr-2’s 1.44 pages per second on one H100 80GB (Datalab’s model card, undated), scaled by the lower of the memory bandwidth and BF16 compute ratios to the H100 SXM from NVIDIA’s product pages and datasheets, read on 10 October 2026; one copy of the model per card, RTX PRO 6000 Server Edition, DGX Spark (128 GB); add headroom for peaks and reprocessing.

Small OCR models scale as one copy per card behind a queue, so more cards add throughput without splitting a model. A DGX Spark (273 GB/s) suits a developer comparing models on sample documents, at about 0.12 pages per second with a 5B model by the same method. These estimates cover throughput and not accuracy, and a model that is fast on invoices may fail on handwriting or tables, so test each document type.

We supply these cards for servers you already run, or in AI servers built to order with 2 to 8 GPUs. Send us your pages per day, the processing window and the document types through the form below for a configuration and a quote.

Invoice extraction and contracts: OCR first, fields second

For invoice extraction, a common design runs two stages. An OCR model turns every page into Markdown or text with tables, and a text LLM then extracts the fields into a fixed schema. vLLM supports this with structured outputs, where with the json option “the output will follow the JSON schema”. The OCR stage carries the page volume on small cards, and the text model sees a few thousand tokens per invoice instead of the image.

A general VLM that reads the page image and returns the fields in one step fits low volumes and difficult layouts, such as stamps, handwriting or scanned forms. At high volume it costs about five times the compute per page by our method above, so route only the pages the OCR stage cannot handle with confidence to it.

Contracts run to many pages, and the questions span clauses on different pages. Convert the whole contract with the OCR stage first, then give the text to a long-context LLM or index it for retrieval. Our article on RAG on company data covers the index.

Deploying open models with vLLM, with process automation on top, is part of our Private AI/ML service, with engineering by our partner Vixen.UNO. Describe your document types, languages and the fields you extract in the form below; the first call is free of charge.

Licences and data protection for document models

Licence terms belong to each checkpoint and differ by model. GLM-OCR and Unlimited-OCR are published under MIT. DeepSeek-OCR-2, olmOCR-2, Qwen3.8-27B and Gemma 4 come under Apache 2.0. Nemotron Parse 2.0 comes under the OpenMDW License Agreement 1.1. The weights of chandra-ocr-2 use a modified OpenRAIL-M licence, which permits use without a separate licence for “research, personal use, and startups under” a funding or revenue threshold its card states, and the card adds that they “Cannot be used competitively with our API”.

Llama 4 is a natively multimodal model whose Acceptable Use Policy excludes companies with their principal place of business in the EU from the Section 1(a) rights for its multimodal models, with an exception for end users. Our Llama 4 hardware guide quotes the clause. Whether a clause applies to your company is a legal assessment for your legal department.

Invoices, contracts and personnel files usually contain personal data, so the GDPR applies to their processing wherever the model runs. On your own servers the page images and the extracted text stay in your network.

What we supply

We supply the GPUs for document AI across this range, from the L4 and RTX PRO 4000 for an OCR model in an existing server to the RTX PRO 6000 Server Edition and the H200 NVL in AI servers built to order. Every card comes with manufacturer warranty, on one EU contract and invoice, and we check the rack, power and airflow before we quote. From your pages per day and document types we return a configuration and a quote within one business day. Deploying the models on-premise, with RAG and MLOps on top, is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

How much GPU memory does a vision language model need?
It needs the weights, including the vision encoder, plus a KV cache for the image tokens and the generated text of every page in flight. Open OCR models of 0.9B to 8.3B parameters have weights of 2.7 to 10.6 GB and fit a 24 GB card, while Qwen3.8-27B takes 30.9 GB in FP8 and Gemma 4 31B 62.6 GB in BF16. By our estimate one 96 GB RTX PRO 6000 holds about 150 pages in flight with Qwen3.8-27B in FP8.
What are the GPU requirements for OCR with an LLM?
A document OCR model such as GLM-OCR, DeepSeek-OCR-2 or chandra-ocr-2 runs on one 24 GB card such as the L4 or RTX PRO 4000, and vLLM’s recipe for Unlimited-OCR states that a single GPU with 8 GB or more is enough for BF16 inference. The card count then follows from the pages per day. A general vision-language model of 27B to 31B parameters needs a 96 GB RTX PRO 6000 or a 141 GB H200 NVL.
How many image tokens does Qwen VL use per page?
Qwen3.8-27B uses the Qwen3-VL processor with a patch size of 16 and a merge size of 2, so one token covers 32 × 32 pixels, with 64 to 16,384 tokens per image. An A4 page rendered at 150 dpi comes to about 2,145 tokens and at 200 dpi to about 3,796, by our arithmetic. The render resolution is therefore the main setting for cost per page.
Can invoice extraction with an LLM run on-premise?
Yes, with open models on your own GPU server: an OCR model converts each page to text, and a text LLM extracts the fields into a JSON schema using vLLM’s structured outputs. GLM-OCR and Unlimited-OCR are MIT-licensed and DeepSeek-OCR-2 and olmOCR-2 Apache 2.0, while Mistral lists its OCR models as Premier with no open licence. The page images and extracted data then stay in your network.
How many pages per day can one GPU process with an OCR model?
Datalab reports 1.44 pages per second for its 5B chandra-ocr-2 on one H100 80GB with vLLM and 96 concurrent sequences, which is about 41,000 pages in 8 hours. Scaled by the lower of memory bandwidth and compute, our estimate is about 0.13 pages per second on an L4, 0.38 on an RTX PRO 5000 and 1.2 on an H200 NVL. Measure your own documents in a pilot, since layout and output length change the rate.
What multimodal LLM hardware handles 100,000 pages a day?
With a 5B OCR model, our estimate is three H200 NVL for an 8-hour window, or two RTX PRO 6000 running over 24 hours, with one copy of the model per card. Reading every page with a 27B vision-language model takes about five times the compute, so it needs five H200 NVL over 24 hours and about 15 for an 8-hour window. A two-stage pipeline, OCR first and a text LLM for the fields, keeps the large model off most pages.

Send us your pages per day, the document types and languages, the processing window and the fields you extract. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna