Open source LLM for business: choosing an open-weight model by licence, languages and GPUs
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- An open LLM for a company assistant is chosen by five criteria: licence terms for company use, quality in your languages, a size your GPUs hold at the peak of users, tool calling and context length, and the vendor’s update record
- As read on 10 October 2026, gpt-oss-120b, Qwen3.8-27B, Gemma 4 and Mistral Small 4 are Apache 2.0 and DeepSeek V4-Flash is MIT, while Llama 4, Qwen3.8-Flash-Next, Mistral Medium 3.5, GLM-5.3 and Nemotron 3 Nano and Super come under their vendors’ own licences, and Nemotron 3 Ultra under OpenMDW-1.1
- gpt-oss-120b as released and Qwen3.8-27B in FP8 fit one RTX PRO 6000 or one H200 NVL, Mistral Small 4 in FP8 and DeepSeek-V4-Flash-0731 need two RTX PRO 6000, and DeepSeek-V3.2 and GLM-5.3 need eight H200 NVL, by the estimates in our hardware guides
- Every family on the shortlist has a vLLM tool-call parser in vLLM’s documentation or recipes, and context windows range from 131,072 tokens for gpt-oss to 1M for DeepSeek-V4-Flash, GLM-5.3 and Nemotron 3 Super and 10M for Llama 4 Scout
- LMArena’s text leaderboard of 8 October 2026 lists open-weight models with overlapping rank spreads, such as 55 to 95 for gemma-4-31b, so a test with your own questions decides between two or three candidates
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
Open source LLM for business: how to choose a model
For a company assistant, an open source LLM, meaning an open-weight model run on your own servers, is chosen by five criteria: licence terms that cover your use, answer quality in the languages your staff write, a size that your GPUs hold for the number of users at the peak, tool calling and context length for the tasks you plan, and the vendor’s record of updates. As of October 2026 the families on our shortlist, each with its own hardware guide, are gpt-oss, Qwen3.8, Mistral, Gemma 4, Llama 4, DeepSeek, GLM and Nemotron. A test with your own questions decides which of them suits your assistant; leaderboards and vendor claims help only to narrow the list.
This article uses the term open-weight, because “open source” describes a licence, and names each checkpoint’s licence as its licensor does: Apache 2.0 and MIT are on the Open Source Initiative’s list of approved licences, while the Llama 4 Community License, the Qwen Community License, the GLM-5.3 License and the NVIDIA Nemotron Open Model License are their vendors’ own terms.
This comparison gives facts from model cards, licence files and vLLM documents, read on 9 and 10 October 2026, and no ranking of our own. Sizing per model is in our LLM hardware requirements by model, and the clauses of each licence in our comparison of open LLM licences for company use.
Open-weight model families compared: licence, sizes and languages
| FAMILY | LICENCE | SIZES | LANGUAGES ON CARD | HARDWARE GUIDE |
|---|---|---|---|---|
| gpt-oss (OpenAI) | Apache 2.0 | 20b; 120b, 5.1B active | no list; German in OpenAI’s MMMLU results | gpt-oss-120b |
| Qwen3.8 (Alibaba) | 27B: Apache 2.0; Flash-Next: Qwen Community License 1.0; 2.4T: Qwen3.8-Max License | 27B dense; Flash-Next, 6B active; 2.4T | none on the cards; 119 in the Qwen3 launch post | Qwen |
| Mistral (Mistral AI) | Small 4, Large 3, Ministral 3: Apache 2.0; Medium 3.5: modified MIT | 3B to 14B; 119B; 128B; 675B | “dozens”, German named; Polish and Romanian in Small 4 metadata | Mistral |
| Gemma 4 (Google) | Apache 2.0; Gemma 3 stays under the Gemma Terms of Use | E2B, E4B, 12B, 26B A4B, 31B | 35+ out of the box, 140+ in pre-training, none named | Gemma |
| Llama 4 (Meta) | Llama 4 Community License; multimodal rights withheld from EU-based licensees, end users excepted | Scout 109B, Maverick 400B, 17B active each | 12, German among them, not Polish or Romanian | Llama 4 |
| DeepSeek | MIT | V4-Flash, 13B active; V4.1-Flash; V3.2, 685B | none on the V4-Flash-0731 card | DeepSeek |
| GLM (Z.ai) | GLM-5.3 License; GLM-5. | 4.7-Flash 30B; 5.3-Flash 320B; 5.3 753B | GLM-5.3 tags: English, Chinese | GLM |
| Nemotron 3 (NVIDIA) | Nano, Super: NVIDIA Nemotron Open Model License; Ultra: OpenMDW-1.1 | Nano 4B and 30B; Super 120B; Ultra 550B | Super: English, French, German, Italian, Japanese, Spanish, Chinese | Nemotron |
Model cards, licence files and metadata on Hugging Face, read on 9 and 10 October 2026; Qwen3 launch post of 29 April 2025; OpenAI’s gpt-oss model card of 5 August 2025. “Languages on card” lists what the vendor names, not tested quality.
The licence is set per checkpoint, not per family. Qwen3.8-27B is Apache 2.0, while Qwen3.8-Flash-Next asks for a separate licence from Qwen if the licensee or an affiliate runs a Model as a Service or AI Work Assistant business, with internal use exempt. Gemma 4 is Apache 2.0, while Gemma 3 stays under the Gemma Terms of Use and their Prohibited Use Policy. Mistral Medium 3.5’s modified MIT licence grants no rights if your company’s monthly revenue exceeds a threshold it states, and the condition extends to derivatives.
For Llama 4, Meta’s Acceptable Use Policy withholds the Section 1(a) rights to the multimodal models from companies with their principal place of business in the EU, except for end users of a product that incorporates them, and Meta’s card calls the Llama 4 models “natively multimodal AI models”. The GLM-5.3 License asks Model-as-a-Service operators above a revenue threshold to pass a security review by Z.AI, as the licence spells the name. Whether a clause applies to your company is a legal assessment for your legal department.
Languages: what the model cards state
A model card lists the languages its vendor names, and it says nothing about quality on your contracts or tickets. Nemotron 3 Super names seven languages, German among them, and Llama 4 names twelve, including German but not Polish or Romanian. Gemma 4 gives 35+ languages out of the box and 140+ in pre-training without naming them.
The cards of Qwen3.8-27B and DeepSeek-V4-Flash-0731 name no languages, and GLM-5.3 carries only the tags English and Chinese. OpenAI reports 83.0 for gpt-oss-120b on German MMMLU at high reasoning effort, a developer’s own figure. For German, Polish and Romanian, European models such as EuroLLM, Bielik and PLLuM are compared in our article on open LLMs for European languages.
Model size, GPUs and concurrent users
The GPUs must hold the checkpoint’s weights plus a KV cache for every conversation in flight. By the estimates in our hub, gpt-oss-120b as released (65.3 GB) and Qwen3.8-27B in FP8 (30.9 GB) run on one RTX PRO 6000 or one H200 NVL, and one DGX Spark (128 GB) holds either for a pilot. Mistral Small 4 in FP8 and Llama 4 Scout in FP8 need two RTX PRO 6000, or one to two H200 NVL. DeepSeek-V4-Flash-0731 needs two RTX PRO 6000 for 20 users at 32K. Llama 4 Maverick needs four H200 NVL or eight RTX PRO 6000, and DeepSeek-V3.2 and GLM-5.3 need eight H200 NVL.
For a whole company the peak of requests in flight sets the card count. Our GPU sizing guide by company size works through 2,000 employees with example values: 40 per cent active in the busiest hour, 6 requests an hour each, 30 seconds per request and a peak factor of 2. That gives 80 requests in flight. With gpt-oss-120b at 32K, one RTX PRO 6000 holds about 19 conversations and one H200 NVL about 55, so the peak needs five RTX PRO 6000 copies or two H200 NVL copies.
Two servers with two H200 NVL each still hold 110 conversations if one fails. With DeepSeek-V4-Flash, a copy on two RTX PRO 6000 holds about 43 conversations at 0.12 GiB of FP8 cache each, so two servers with four cards each keep the peak of 80 if one fails. These are memory estimates, and response time under load needs its own test.
We build AI servers to order with H200 NVL or RTX PRO 6000 Server Edition cards and return a configuration and quote within one business day. Send us the models on your shortlist and your expected peak of users.
Tool calling, context length and reasoning modes
An assistant that looks up tickets, queries a database or calls an internal API needs a model that emits tool calls and an engine that parses them. vLLM’s tool-calling documentation marks --enable-auto-tool-choice as mandatory for this, and --tool-call-parser names the format of the model family.
| MODEL | CARD ON TOOLS | VLLM TOOL PARSER | CONTEXT |
|---|---|---|---|
| gpt-oss-120b | “function calling, web browsing, Python code execution, and Structured Outputs” | openai, vLLM docs | 131,072 |
| Qwen3.8-27B | “long-horizon agentic tasks” | qwen3_coder, recipe | 262,144, up to 1M |
| Mistral Small 4 | “native function calling and JSON output” | mistral, vLLM docs and card | 256K |
| Gemma 4 31B | “native function-calling support” | gemma4, recipe | 256K |
| Llama 4 Scout | no statement on the card | llama4_ | 10M |
| Deep | “substantially enhanced agentic capabilities” | deepseek_v4, recipe | 1M |
| GLM-5.3 | “complex coding and long-horizon tasks” | glm47, recipe | 1,048,576 |
| Nemotron 3 Super | “tool use” among its intended workloads | qwen3_coder, recipe | 1M |
Hugging Face model cards; vLLM tool-calling documentation dated 8 October 2026 and vLLM recipes updated between 11 May and 8 October 2026, read on 10 October 2026; context from the cards and our hardware guides.
Plan for the context your requests need rather than the longest a model accepts, because the KV cache grows with the declared context. On one RTX PRO 6000 gpt-oss-120b holds about 19 conversations at 32K, 79 at 8K and 4 at 128K by our estimate, and the context you set must include retrieved passages and tool results.
Reasoning modes change the time each request stays in flight. gpt-oss sets its reasoning level in the system prompt, for example “Reasoning: high”, Mistral Small 4’s card states reasoning effort “configurable per request”, and for high and maximum reasoning effort DeepSeek recommends up to 384K output tokens. Test each use case at the reasoning level it will run.
Release cadence and the vendor’s update record
Open model families publish new versions several times a year, often within the service life of one server. DeepSeek published V4-Flash as a preview, superseded it with the -0731 release and released V4.1-Flash on 10 September 2026, which our DeepSeek guide places on four H200 NVL with its Engram tables in host memory, as vLLM’s recipe runs it. NVIDIA released Nemotron 3 Nano 30B-A3B on 15 December 2025, Super on 11 March 2026, Ultra on 4 June 2026 and 3.5 Lightning on 11 August 2026, and the last two came under OpenMDW-1.1 rather than the Nemotron Open Model License. Google moved from the Gemma Terms of Use for Gemma 3 to Apache 2.0 for Gemma 4.
A new version can therefore change the licence, the memory footprint and the cards it needs. For each vendor on the shortlist, note how often checkpoints appear, whether earlier ones stay downloadable, whether FP8 or NVFP4 checkpoints and vLLM recipes exist for your cards, and whether the licence stayed the same. Before you swap a model in production, run the regression tests described in our guide to LLM evaluation before model upgrades.
Public leaderboards: what LMArena shows
LMArena’s text leaderboard, which lmarena.ai now forwards to arena.ai under the title “Text Arena”, showed 8,734,330 votes across 414 models on 8 October 2026. It describes its rankings as covering “text-to-text tasks across math, coding, creative writing, and other open-ended domains”. Each entry has a score with a confidence interval, a rank spread and a licence column; the licences in this article come from the licensors’ own files.
On that date gemma-4-31b had a score of 1452 ±7, with a rank spread of 55 to 95. qwen3.8-27b had 1438 ±5 and a spread of 81 to 118. gpt-oss-120b had 1352 ±4, and nvidia-nemotron-3-super-120b-a12b had 1360 ±7. deepseek-v4-flash had 1436 ±4 and a spread of 86 to 119, and the page does not say whether this entry is the preview or the -0731 release. A separate entry, deepseek-v4-flash-high-preview, is the preview at a set reasoning level.
A leaderboard helps to build a shortlist, but it holds none of your documents, languages or tools. The page does not define the rank spread; spreads that overlap this widely suggest that neighbouring models are not reliably ordered, so the decision rests on your own test.
Selection procedure for a company assistant
- Write down the use cases, the languages, the peak of users, the context each request needs and the tools the assistant must call.
- Read the licence of each candidate checkpoint, record the date read and pass it to your legal department.
- Drop models whose weights plus cache at the peak do not fit the cards you plan, using our hardware guides.
- Keep two or three candidates and serve each with the same engine, prompts, retrieved passages and context limit.
- Build an evaluation set from your own questions in every language, with tool-call cases and their expected calls, and have native speakers grade the answers blind.
- Load-test the candidate your graders prefer on the production GPU type at the expected peak, measuring time to first token and per-user speed.
- Record the repository, revision and licence of the chosen checkpoint, and run the evaluation set again on every update.
The technical assessment in our Private AI/ML service covers model and GPU selection and a pilot plan with metrics, at a price fixed before work begins. Describe your use cases and languages in the form below.
What we do
Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise, or on dedicated hardware in a Tier-3 data centre in Lithuania, and logs queries and answers so that security and legal see who accesses what. The first call is free of charge and leaves you with two or three possible solution scenarios. We start with a pilot on one process with clear metrics and scale only what has proved its value. Eurokommerz holds the contract and supplies the AI servers with H200 NVL, RTX PRO 6000 Server Edition or L40S cards, with engineering by our partner Vixen.UNO.
FAQ
Which open source LLM should a business use for a company chatbot?
Which open-weight models can a company use commercially?
How do gpt-oss, Qwen, Mistral and Gemma compare for enterprise use?
Which self-hosted LLMs support tool calling?
What hardware does a self-hosted LLM for 2,000 employees need?
Are LLM leaderboards a good way to choose an enterprise model?
Send us your use cases, the languages your staff write in, the expected peak of users and the models you are considering. We reply within one business day, and in the first call we work through 2 to 3 possible solution scenarios with you. The first call is free of charge.
Talk to an expertWe reply within one business day