Open LLMs for German, Polish and Romanian: European models, licences and GPU memory compared
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- EuroLLM (1.7B, 9B and 22B, Apache 2.0) and Teuken-7B name German, Polish and Romanian among all 24 official EU languages; Teuken’s commercial v0.4 is Apache 2.0, its newer v0.6 CC BY-NC 4.0 for non-commercial use
- For Polish, Bielik v3 (1.5B to 11B, Apache 2.0) and PLLuM (4B and 12B under Apache 2.0, 8B and 70B under Meta’s Llama 3.1 licence) are trained with an emphasis on the language; the Romanian instruct models of OpenLLM-Ro are CC BY-NC 4.0
- Mistral Small 3.1, 3.2 and 4 name all three languages and are Apache 2.0, while Mistral Medium 3.5’s modified MIT licence grants no rights to companies above a monthly revenue threshold
- Qwen3’s launch post lists all three languages among 119, Gemma 4’s card gives 35+ languages out of the box and 140+ in pre-training without naming them, and Llama 3.1, 3.3 and 4 list German but not Polish or Romanian; OpenEuroLLM had published research checkpoints as of 6 October 2026
- In BF16, EuroLLM-22B takes 42.2 GiB, which leaves 40.9 GiB of KV cache on one 96 GB RTX PRO 6000 by our sizing rule, enough for about 24 conversations of 8,192 tokens
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
Open LLMs for German, Polish and Romanian: what the model cards list
The open models whose cards name German, Polish and Romanian include EuroLLM (1.7B, 9B and 22B), Teuken-7B, the 7B and 11B Bielik v3 models and Mistral Small 3.1, 3.2 and 4, and Qwen3’s launch post lists all three among 119 languages. Bielik and PLLuM are trained with an emphasis on Polish, and the OpenLLM-Ro project publishes Romanian fine-tunes for research use. EuroLLM, Bielik v3 and these Mistral Small versions are Apache 2.0, while Teuken’s current version and the OpenLLM-Ro instruct models exclude commercial use and two PLLuM sizes come under Meta’s Llama 3.1 licence. “Open” here means downloadable weights under the licence on the card, which is not always an open source licence: the Open Source Definition does not allow a licence to restrict use “in a business”.
| MODEL, ORIGIN | SIZES | DE, PL, RO | LICENCE ON THE CARD |
|---|---|---|---|
| EuroLLM (EU-funded) | 1.7B, 9B, 22B; EuroMoE 2.6B | all three, among 35 languages | Apache 2.0 |
| Teuken-7B (OpenGPT-X) | 7B | all three, among the 24 EU languages | v0.4 commercial: Apache 2.0; v0.6: CC BY-NC 4.0 |
| Salamandra, ALIA (BSC) | 2B, 7B; ALIA 40B | all three in pre-training; instruct tuning mainly Iberian languages and English | Apache 2.0 |
| Bielik v3 (SpeakLeash) | 1.5B, 4.5B, 7B, 11B | 7B, 11B: all three, among 32 European languages; 1.5B, 4.5B: Polish | Apache 2.0; contact details to download |
| PLLuM (HIVE AI) | 4B, 8B, 12B, 70B | Polish, with English data | 4B, 12B: Apache 2.0; 8B, 70B: Llama 3.1 licence |
| Mistral Small 3.1, 3.2, 4 | 24B; 119B | all three named | Apache 2.0 |
| Mistral Medium 3.5 | 128B | all three named | modified MIT with a revenue threshold |
| Ministral 3, Mistral Large 3 | 3B to 14B; 675B | German named; Polish, Romanian not | Apache 2.0 |
| Apertus (Swiss AI) | 0.5B to 70B | “over 1000 languages”, none named | Apache 2.0 and an acceptable use policy |
| OpenLLM-Ro instruct models | 2B to 9B | Romanian | CC BY-NC 4.0 |
Hugging Face model cards and language metadata, read on 6 October 2026. DE, PL, RO: whether the card or its metadata names German, Polish and Romanian.
EuroLLM, Teuken and OpenEuroLLM: models for all 24 EU languages
EuroLLM is developed by ten universities, research centres and companies with EU funding, and its cards list the 24 official EU languages and 11 more. EuroLLM-22B has 22.6 billion parameters and a context of 32,768 tokens. Its card says it matches or outperforms Gemma-3-27B, Qwen-3-32B and Apertus-70B in translation. The instruct cards of the 9B and 22B models state that the model “has not been aligned to human preferences” and may produce hallucinations or false statements.
Teuken-7B was developed in the research project OpenGPT-X by Fraunhofer, Forschungszentrum Jülich, TU Dresden and DFKI. It covers the 24 official EU languages with a context of 4,096 tokens, which leaves little room for retrieved passages in a RAG prompt. Teuken-7B-instruct-commercial-v0.
OpenEuroLLM started on 1 February 2025 with 20 organisations and funding from the Digital Europe Programme, to build open models for the “EU official languages and beyond”. As of 6 October 2026 its Hugging Face organisation holds research checkpoints, and the card of its instruction-tuned 9B model calls it an “Experimental SFT-stage research checkpoint; not a final production assistant”. In March 2026 the project wrote that it “is looking to release an 8B model by next summer”.
Salamandra, ALIA and Apertus: what pre-training and tuning cover
Salamandra, from BSC, the supercomputing centre in Barcelona, comes in 2B and 7B sizes pre-trained on 35 European languages and code, German, Polish and Romanian among them. ALIA-40b was pre-trained on 9.83 trillion tokens of 35 European languages, and all of these models are Apache 2.0. ALIA-40b-instruct-2606 states that “the post-training process concentrated primarily on Spanish, Catalan, Basque, Galician, and English”, and its metadata, like that of salamandra-7b-instruct-2606, lists only those five languages. Test these instruct models in German, Polish or Romanian before you rely on them.
The Swiss AI Initiative describes Apertus as a “Fully Open Model” with open weights and open training data, and its 2509 card says it “supports over 1000 languages”. Version 1.5, in 8B and 70B sizes, is Apache 2.0, but downloading it means accepting an acceptable use policy and sharing contact details; version 1.1 adds 0.5B, 1.5B and 4B models. The card advises fetching new weights every six months, as the project handles data-protection deletion requests through new releases.
Polish LLMs: Bielik and PLLuM
SpeakLeash and ACK Cyfronet AGH develop Bielik. In its v3 generation the 1.5B and 4.5B models are tagged for Polish only, while the 7B and 11B models are “Multilingual (32 European languages, optimized for Polish)”, with German and Romanian in their metadata. All are Apache 2.0, and the download page asks you to agree to share contact information. The 11B instruct card notes that the model “does not have any moderation mechanisms”. Bielik-PL-11B-v3.
PLLuM is, in its card’s words, “a family of large language models (LLMs) specialized in Polish with additional English data incorporated for broader generalization”, developed by the PLLuM consortium in 2024 and since 2025 under HIVE AI, an initiative for the Polish public sector financed by Poland’s Minister of Digital Affairs. The December 2025 generation, marked 2512, comes in base, instruct and chat variants. PLLuM-4B is built on Google’s gemma-3-4b-pt and PLLuM-12B on Mistral-Nemo-Base-2407, both Apache 2.0 on their cards. Llama-PLLuM-8B and Llama-PLLuM-70B are built on Llama 3.1 and come under Meta’s Llama 3.1 licence and its acceptable use policy; the licence requires anyone who distributes them, or a product or service that contains them, to display “Built with Llama”. Google’s Gemma Terms of Use count “works based on Gemma” as model derivatives and pass their use restrictions on, so how they combine with the Apache 2.0 on the PLLuM-4B card is a question for your legal department.
Romanian LLMs and Romanian in multilingual models
Of the models above, EuroLLM, Teuken, the Salamandra base models, Bielik’s 7B and 11B models, Mistral Small 3.1, 3.2 and 4, Mistral Medium 3.5 and Qwen3 name Romanian, while Llama 3.1, 3.3 and 4 do not. The OpenLLM-Ro project fine-tunes Llama, Mistral, Gemma and Qwen models for Romanian and publishes its instruct models under CC BY-NC 4.0, which excludes commercial use. The RoLlama3.1-8b-Instruct card, version of 23 April 2025, describes the model as intended for research use in Romanian. It reports 6.43 on RoMT-Bench, against 5.69 for the Llama 3.1 8B Instruct it starts from.
Mistral open-weight models: one licence per model
Mistral AI publishes most of its open-weight models under Apache 2.0, among them Mistral NeMo, Mistral Small 3.1 and 3.2 (24B), Mistral Small 4 (119B), the Ministral 3 models (3B, 8B and 14B) and Mistral Large 3 (675B). The Mistral Small 3.1 card names 25 languages, German, Polish and Romanian among them, and the 3.2 card calls that model “a minor update” of 3.1. The newer cards name German among “dozens of languages”; Polish and Romanian appear only in the metadata of Mistral Small 4 and Mistral Medium 3.5, not in that of Ministral 3 or Mistral Large 3.
Mistral Medium 3.5, a dense 128B model, comes under a “Modified MIT License” that grants no rights at all if the “global consolidated monthly revenue” of your company exceeds a threshold set in the licence, and the condition extends to derivatives. Older models such as Mistral Large 2411 and Ministral 8B 2410 use the “Mistral Research License”, which allows non-commercial use; Mistral offers a commercial licence on request.
Qwen3, Gemma 4, Llama and gpt-oss: languages on the cards
| MODEL | WHAT THE CARD SAYS | DE, PL, RO | LICENCE |
|---|---|---|---|
| Qwen3 | “100+ languages and dialects”; 119 in the launch post | all three in the launch post | Apache 2.0 (Qwen3-32B) |
| Gemma 4 | 35+ out of the box, 140+ in pre-training | none named | Apache 2.0 |
| Llama 3.1, 3.3 | eight supported languages | German only | Meta’s community licence and use policy |
| Llama 4 | 12 supported, 200 in pre-training | German only | Llama 4 licence; multimodal rights withheld in the EU |
| gpt-oss | no language list | German tested in MMMLU | Apache 2.0 and a usage policy |
Model cards; Qwen3 launch post (29 April 2025); Meta’s Llama 4 acceptable use policy; OpenAI’s gpt-oss model card (5 August 2025); all read on 6 October 2026.
OpenAI’s gpt-oss model card of 5 August 2025 reports MMMLU, a professionally translated MMLU in 14 languages without Polish or Romanian, on which gpt-oss-120b scores 83.0 in German at high reasoning effort. Our guide to a private ChatGPT alternative compares their licence terms, including Llama 4’s clause on companies with their principal place of business in the EU.
GPU memory per model size
Our sizing rule takes 0.9 of the memory the driver reports and subtracts about 3 GiB of overhead and the weights, two bytes per parameter in BF16; the rest holds the KV cache, as our guide to how much VRAM an LLM needs explains.
| MODEL | BF16 WEIGHTS | RTX PRO 5000, 48 GB | RTX PRO 6000, 96 GB | H200 NVL, 141 GB |
|---|---|---|---|---|
| EuroLLM-9B-Instruct-2512 | 17.0 GiB | 23.2 GiB | 66.0 GiB | 106.3 GiB |
| Bielik-11B-v3. | 20.8 GiB | 19.4 GiB | 62.2 GiB | 102.6 GiB |
| PLLuM-12B-instruct-2512 | 22.8 GiB | 17.4 GiB | 60.2 GiB | 100.5 GiB |
| EuroLLM-22B-Instruct-2512 | 42.2 GiB | does not fit | 40.9 GiB | 81.2 GiB |
| Mistral-Small-3. | 44.7 GiB | does not fit | 38.3 GiB | 78.6 GiB |
| ALIA-40b-instruct-2606 | 75.3 GiB | does not fit | 7.7 GiB | 48.0 GiB |
| Llama-PLLuM-70B-instruct | 131.4 GiB | does not fit | does not fit | does not fit |
Cache room per card by this rule, with 95.6 GiB (RTX PRO 6000) and 140.4 GiB (H200 NVL) as the driver reports them and a nominal 48 GB; parameter counts from Hugging Face, 6 October 2026. “Does not fit” means no room left for the cache.
EuroLLM-22B stores 216 KiB of cache per token in BF16, from the 54 layers and 8 KV heads of 128 dimensions in its config.json; the card’s table gives 56 layers, which does not match the parameter count. The 40.9 GiB left on one RTX PRO 6000 therefore hold about 24 conversations of 8,192 tokens, or 6 at the full 32,768, and an FP8 cache doubles both. Llama-PLLuM-70B needs two RTX PRO 6000 cards in BF16, which leave about 17 GiB of cache per card, or a quantised version; our guide to FP8, NVFP4, INT4 and GGUF sets out what each format saves and what it costs in accuracy.
We build AI servers to order with the RTX PRO 5000, RTX PRO 6000 or H200 NVL and send the configuration and quote within one business day. Send us the models on your shortlist and how many people will use them at once.
Evaluating a model in your language with your documents
Public benchmarks contain none of your contracts or support tickets, so a test of your own decides between two or three candidates.
- Shortlist models whose licence covers your intended use and whose weights fit your GPUs at the precision you plan to run.
- Build the evaluation set described in our article on RAG on company data in every language your users write, with questions in one language about documents in another.
- Serve each candidate with the same engine, prompt, retrieved passages and context limit, for example through vLLM’s OpenAI-compatible API; our comparison of serving engines covers the options.
- Have native speakers grade the answers blind for correctness against the source, the language of the answer, grammar, terminology and the citation.
- Count the tokens each tokenizer produces for a sample of your documents, since the counts differ between models, and measure time to first token on your hardware.
- Run the set again on each new version; EuroLLM, Bielik and PLLuM have each published several generations.
The technical assessment in our Private AI/ML service covers model and GPU selection and a pilot plan with metrics, at a price fixed before work begins. Tell us which languages and document types your assistant has to handle.
What we do
Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise, so data does not leave your network and access is controlled, or on dedicated hardware in a Tier-3 data centre in Lithuania, where data stays in the EU. The first call is free of charge and leaves you with two or three possible solution scenarios. We start with a pilot on one process with clear metrics and scale only what has proved its value. Eurokommerz holds the contract and supplies the GPU servers, with engineering by our partner Vixen.UNO and support under an agreed SLA.
FAQ
What is PLLuM?
Which open LLMs support German?
Which open LLMs are built for Polish?
Is there an open LLM for Romanian?
What is EuroLLM?
Which European open LLMs can be used commercially?
Send us the languages your users write in, the document types the assistant should answer from, the models you are considering and the number of people who will use it. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day