On-premise machine translation: NMT models, LLMs and the GPUs a self-hosted translation server needs
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A self-hosted translation server runs either dedicated NMT models (OPUS-MT, NLLB-200, MADLAD-400), which translate sentence by sentence and need from under 1 GB to about 21 GB for 16-bit weights, or LLMs such as TranslateGemma, Tower and EuroLLM, which translate with context and need about 10 to 58 GB up to 27B parameters by our estimate
- The NLLB README puts all models under CC-BY-NC 4.0, and the NLLB-200 cards call it “a research model” that “is not released for production deployment”; the Tower models are non-commercial too, OPUS-MT, MADLAD-400, EuroLLM and Mistral Small 3.2 are CC BY 4.0 or Apache 2.0, and TranslateGemma comes under the Gemma Terms of Use
- NLLB-200 3.3B needs about 6.6 GB for 16-bit weights and fits a 16 GB RTX PRO 2000 or a 24 GB L4; in CTranslate2’s published benchmark an OPUS-MT model generated 9,296.7 target tokens per second in float16 on an NVIDIA A10G
- At our planning figure of 2 output tokens per word, 1 million words a day average 69 tokens per second over eight hours, and one 48 GB L40S or RTX PRO 5000 holds a 9B to 12B translation LLM in BF16 for that volume by our estimate
- Context limits set how documents are split: NLLB was trained on inputs of at most 512 tokens and TranslateGemma takes 2K tokens of input, Tower+ 9B 8,192; file formats come from the pipeline, such as Argos Translate’s file library for DOCX, PPTX, ODT, HTML and PDF
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
On-premise machine translation: NMT models or LLMs
On-premise machine translation runs on one of two kinds of model. By our estimate in the sizing section below, one 48 GB card such as the L40S or RTX PRO 5000 handles about 1 million words a day with a 9B to 12B translation LLM, and NMT models need far less. Dedicated neural machine translation (NMT) models, such as OPUS-MT, Meta’s NLLB-200 and Google’s MADLAD-400, translate sentence by sentence. Their dense models range from a 298 MB file for OPUS-MT German to English to MADLAD-400 at 10.7 billion parameters, about 21 GB in 16-bit. LLMs, both general models such as EuroLLM and Mistral Small and translation-tuned ones such as TranslateGemma and Unbabel’s Tower, translate a paragraph or a section with its context. Up to 27B parameters they need about 10 to 58 GB of GPU memory for 16-bit weights, by our estimate at two bytes per parameter, and Tower+ 72B about 146 GB.
The licence on each model card decides which of them a company can deploy. NLLB-200 and the Tower models are published under non-commercial Creative Commons licences. OPUS-MT, MADLAD-400, EuroLLM and Mistral Small 3.2 come under CC BY 4.0 or Apache 2.0, and TranslateGemma under Google’s Gemma Terms of Use. This article covers the self-hosted alternative to a commercial service such as DeepL, for texts that stay on hardware under your control.
Open translation models, languages and licences
| MODEL | TYPE | PARAMETERS | LANGUAGES | LICENCE ON THE CARD |
|---|---|---|---|---|
| OPUS-MT (Marian) | NMT, per pair or multilingual | not stated; 298 MB file for de-en | per pair, or many into one (mul-en) | CC BY 4.0 (project); Apache 2.0 (opus-mt-de-en) |
| NLLB-200 | NMT | 600M, 1.3B distilled; 1.3B, 3.3B; 54.5B MoE | 200+ | CC BY-NC 4.0, research model |
| MADLAD-400 | NMT, T5 | 3B, 7.2B, 10.7B | 400+ | Apache 2.0 |
| Translate | LLM, Gemma 3 | 4B, 12B, 27B | 55 | Gemma Terms of Use |
| TowerInstruct-7B-v0. | LLM, Llama 2 | 7B | 10, Polish not among them | CC BY-NC 4.0; Llama 2 licence for the base |
| Tower+ | LLM, Gemma 2 or Qwen 2.5 | 9B, 72B | 22 in the header, 26 listed | CC BY-NC 4.0 in the text, CC BY-NC-SA 4.0 in the metadata |
| EuroLLM | LLM | 1.7B, 9B, 22B | 35, all 24 EU languages | Apache 2.0 |
| Mistral Small 3.1, 3.2 | LLM | 24B | 25 named | Apache 2.0 |
Hugging Face model cards and file lists, the OPUS-MT and NLLB (fairseq) READMEs and the TranslateGemma technical report, read on 10 October 2026; EuroLLM and Mistral Small as in our guide to open LLMs for European languages.
The NLLB README states: “All models are licensed under CC-BY-NC 4.0 available in Model LICENSE file.” Each NLLB-200 card adds: “NLLB-200 is a research model and is not released for production deployment.” The same cards say the model “is not intended to be used for document translation”. TowerInstruct-7B-v0.2 is CC-BY-NC-4.0 and notes that its Llama 2 base is under the Llama 2 Community License.
MADLAD-400 is Apache 2.0, but its card names the “Research community” as its intended users and says the research models “have not been assessed for production usecases”. The OPUS-MT project licenses its pre-trained models under CC-BY 4.0, while the card of opus-mt-de-en states Apache 2.0, so check the card of each language pair you deploy. The TranslateGemma card says that to access the model “you’re required to review and agree to Google’s usage license”, and our guide to open LLMs for European languages covers how those terms treat derived models. Whether an internal translation service counts as non-commercial use under a CC BY-NC licence is a legal assessment for your legal department.
NLLB GPU requirements and other NMT models
NLLB-200 3.3B is published in 32-bit, as three PyTorch files of 17.6 GB in total. At two bytes per parameter in 16-bit, the 3.3B model needs about 6.6 GB for its weights and the 600M model about 1.2 GB, our estimate from the parameter counts. Either fits a 16 GB RTX PRO 2000, a 24 GB L4 or a 24 GB RTX PRO 4000 SFF. MADLAD-400 at 10.7B needs about 21 GB by the same rule, which fits the 32 GB RTX PRO 4500 or the 48 GB L40S with room to spare; a 24 GB L4 leaves little room for batches, unless the model runs in the 8-bit mode that CTranslate2 offers. The 54.5B mixture-of-experts version of NLLB needs about 109 GB, which one 141 GB H200 NVL holds.
The NLLB card states that the model was trained on inputs of no more than 512 tokens and allows “single sentence translation among 200 languages”. A server in front of it therefore splits each document into sentences, translates them in batches and puts the text back together. The OPUS-MT project offers a web application with an API and a Docker setup for CUDA GPUs.
CTranslate2, the MIT-licensed inference engine of the OpenNMT project, supports NLLB and MADLAD-400 and converts Marian and OPUS-MT models. Its README publishes a benchmark that translates the English to German newstest2014 test set with an OPUS-MT model on an NVIDIA A10G GPU. With Hugging Face Transformers 4.26.1 the model generated 1,022.9 target tokens per second. CTranslate2 3.6.0 in float16 generated 9,296.7 with 909 MB of GPU memory, at the same BLEU score of 27.90. The README states no batch or beam size and notes that the results “are only valid for the configuration used during this benchmark”.
LLM translation on GPU: TranslateGemma, Tower, EuroLLM and Mistral
TranslateGemma, described in Google’s technical report submitted to arXiv on 13 January 2026, is “a suite of open machine translation models based on the Gemma 3 foundation models”, in 4B, 12B and 27B sizes. The cards show 5B, 13B and 29B parameters and a “Total input context of 2K tokens”. In BF16 the three need about 10, 26 and 58 GB for their weights, our estimate at two bytes per parameter. The 4B model fits an L4, the 12B model a 48 GB L40S or RTX PRO 5000, and the 27B model a 96 GB RTX PRO 6000, or a 72 GB RTX PRO 5000 with less room for batches.
Tower+ 9B is built on Gemma 2 9B with a context of 8,192 tokens, and Tower+ 72B on Qwen 2.5 72B with 131,072 tokens. The 72B card shows 73B parameters. By the same rule they need about 18 GB and 146 GB in BF16, so the 72B model takes two RTX PRO 6000 or two H200 NVL at 16-bit.
An LLM that already serves an assistant in the company can translate as well. Our guide to open LLMs for European languages gives the GPU memory of EuroLLM and Mistral Small 3.2, and our Mistral hardware requirements size the current Mistral models for 1, 20 and 100 users. A translation request is a paragraph or a few pages, so its KV cache is small next to a long chat, and throughput sets the limit sooner than memory does. The method for weights and cache per model is in our LLM hardware requirements by model.
EU language coverage of translation models
NLLB-200 and MADLAD-400 cover 200 and over 400 languages, and OPUS-MT publishes models per language pair and multilingual ones such as mul-en. EuroLLM covers all 24 official EU languages, and the Tower+ cards list 26 languages, Polish, Romanian, Czech and Hungarian among them. TowerInstruct-7B-v0.2 supports ten languages, German among them but not Polish, and its card says it is not intended as a document-level translator. The TranslateGemma cards name 55 languages without a list, and the report evaluates 55 language pairs of WMT24++, so test each pair you need.
Sizing a translation server from words per day
Translation volume is planned in words and GPU throughput in tokens. We plan with 2 output tokens per word, an assumption to replace with a count from your own documents, since each tokenizer splits words differently and the ratio differs by language. A preprint on arXiv of May 2026, which counted the tokens of six models on parallel text, reports means of 1.23 tokens per word for English and 1.76 for German. For Polish and Czech it reports about 2.3, and for Hungarian 2.7. For such target languages, plan with 2.5 to 3 tokens per word, a quarter to a half above the rates in the table. An LLM also processes the source text in its prompt, about the same number of input tokens again plus its instructions. Spread over an eight-hour day, the daily output gives the average rate to sustain, and the busy hour runs above it.
| WORDS PER DAY | OUTPUT TOKENS | AVERAGE OVER 8 H | CONFIGURATION |
|---|---|---|---|
| 100,000 | 200,000 | 7 tokens/s | one L4, RTX PRO 2000 or RTX PRO 4000 SFF in an existing server: NMT models or a 4B LLM |
| 1,000,000 | 2 million | 69 tokens/s | one 48 GB L40S or RTX PRO 5000: a 9B to 12B LLM in BF16 next to NMT models |
| 5,000,000 | 10 million | 347 tokens/s | one RTX PRO 6000 Server Edition: a 22B to 27B LLM, or copies of a smaller one |
| 20,000,000 | 40 million | 1,389 tokens/s | two to four RTX PRO 6000 Server Edition or H200 NVL, one model copy per card |
Our estimates, not measurements: 2 output tokens per word as a planning assumption, more for Slavic or Hungarian targets, averaged over 8 hours; configurations by memory fit with room for batching. Measure your chosen model on a test card before you order.
For NMT models the daily total is small against the published rates. At the 9,296.7 tokens per second of the CTranslate2 benchmark, the 10 million tokens of 5 million words would take about 18 minutes of GPU time on that A10G by our arithmetic, with tokens counted in OPUS-MT’s own vocabulary. Response time for interactive users and the peaks of batch jobs decide the card. For NMT models and LLMs alike we found no published translation throughput on the cards we supply. Our L4 vs L40S comparison sets out how those two cards differ, and a day’s sample of your documents on a test card gives the figure for your model.
We supply the L4, the L40S and the RTX PRO cards for servers you already run, and AI servers built to order with configuration and quote within one business day. Send us your words per day, language pairs and preferred model through the form below.
Document translation on-premise: file formats and layout
The pipeline around the model handles the files. It extracts the text from DOCX, PPTX, ODT, HTML or PDF files, sends it to the model in segments sized to the context limit and writes the translation back into the same structure. Argos Translate’s file library, from the LibreTranslate project and licensed under AGPL-3.0, lists “.txt, .odt, .odp, .docx, .pptx, .epub, .html, .srt, .pdf” as supported formats. LibreTranslate, also AGPL-3.0, offers a /translate_file endpoint for whole files. Neither documentation states whether layout is preserved, so test the pipeline with your own templates and with scanned PDFs, which carry no text until OCR has run. The AGPL-3.0 sets conditions when a modified version is offered to users over a network, which is a question for your legal department.
The context limit sets the segment size, from 512 tokens for NLLB and 2K for TranslateGemma to 8,192 for Tower+ 9B. A longer segment lets the model keep terminology consistent across a section, and an LLM prompt can carry a list of approved terms.
Our Private AI/ML service deploys open models on-premise with vLLM, Ollama or NVIDIA AI Enterprise, so data does not leave your network. Describe your document types and language pairs in the form below; the first call is free of charge.
Testing translation quality before you choose
The findings of the WMT24 general translation task, published in November 2024, carry the subtitle “The LLM Era Is Here but MT Is Not Solved Yet”. In WMT25, published in November 2025, the organisers evaluated 60 systems, 36 submitted by participants and 24 collected from LLMs and online translation providers. Their report is subtitled “Time to Stop Evaluating on Easy Test Sets”. Neither report evaluates company documents such as contracts or manuals, which is why the steps below test on your own texts.
- Shortlist models whose licence covers your use and whose card names every language pair you need.
- Take 50 to 100 segments of your own documents per language pair and domain, with tables, product names and legal wording.
- Translate them with each candidate on the same engine and segment size, and have native speakers mark error spans, as the professional annotators of WMT25 did under its Error Span Annotation protocol.
- Measure tokens per second and response time for one day’s sample on a test card, then size the server from the table above.
What we supply
We supply every GPU this article sizes, from the L4 and RTX PRO 2000 to the RTX PRO 6000 Server Edition and the H200 NVL, as cards for servers you already run or in AI servers built to order. Everything comes on one EU contract and invoice with manufacturer warranty, and we check the rack, power and airflow before we quote; our professional GPU range lists every card. Deploying the translation models on-premise is part of our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How do I run machine translation on-premise?
What are the GPU requirements for NLLB?
Can NLLB be used commercially?
Which GPU do I need for LLM translation?
What are the options for a self-hosted alternative to DeepL?
Does an on-premise translation server keep document formatting?
Send us your daily word volume, the language pairs, the file types and the models you are considering. We reply within one business day with a configuration and a quote in writing, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day