BLOG · COMPARISON ·

Open source LLM for business: choosing an open-weight model by licence, languages and GPUs

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • An open LLM for a company assistant is chosen by five criteria: licence terms for company use, quality in your languages, a size your GPUs hold at the peak of users, tool calling and context length, and the vendor’s update record
  • As read on 10 October 2026, gpt-oss-120b, Qwen3.8-27B, Gemma 4 and Mistral Small 4 are Apache 2.0 and DeepSeek V4-Flash is MIT, while Llama 4, Qwen3.8-Flash-Next, Mistral Medium 3.5, GLM-5.3 and Nemotron 3 Nano and Super come under their vendors’ own licences, and Nemotron 3 Ultra under OpenMDW-1.1
  • gpt-oss-120b as released and Qwen3.8-27B in FP8 fit one RTX PRO 6000 or one H200 NVL, Mistral Small 4 in FP8 and DeepSeek-V4-Flash-0731 need two RTX PRO 6000, and DeepSeek-V3.2 and GLM-5.3 need eight H200 NVL, by the estimates in our hardware guides
  • Every family on the shortlist has a vLLM tool-call parser in vLLM’s documentation or recipes, and context windows range from 131,072 tokens for gpt-oss to 1M for DeepSeek-V4-Flash, GLM-5.3 and Nemotron 3 Super and 10M for Llama 4 Scout
  • LMArena’s text leaderboard of 8 October 2026 lists open-weight models with overlapping rank spreads, such as 55 to 95 for gemma-4-31b, so a test with your own questions decides between two or three candidates

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

Open source LLM for business: how to choose a model

For a company assistant, an open source LLM, meaning an open-weight model run on your own servers, is chosen by five criteria: licence terms that cover your use, answer quality in the languages your staff write, a size that your GPUs hold for the number of users at the peak, tool calling and context length for the tasks you plan, and the vendor’s record of updates. As of October 2026 the families on our shortlist, each with its own hardware guide, are gpt-oss, Qwen3.8, Mistral, Gemma 4, Llama 4, DeepSeek, GLM and Nemotron. A test with your own questions decides which of them suits your assistant; leaderboards and vendor claims help only to narrow the list.

This article uses the term open-weight, because “open source” describes a licence, and names each checkpoint’s licence as its licensor does: Apache 2.0 and MIT are on the Open Source Initiative’s list of approved licences, while the Llama 4 Community License, the Qwen Community License, the GLM-5.3 License and the NVIDIA Nemotron Open Model License are their vendors’ own terms.

This comparison gives facts from model cards, licence files and vLLM documents, read on 9 and 10 October 2026, and no ranking of our own. Sizing per model is in our LLM hardware requirements by model, and the clauses of each licence in our comparison of open LLM licences for company use.

Open-weight model families compared: licence, sizes and languages

FAMILYLICENCESIZESLANGUAGES ON CARDHARDWARE GUIDE
gpt-oss (OpenAI)Apache 2.020b; 120b, 5.1B activeno list; German in OpenAI’s MMMLU resultsgpt-oss-120b
Qwen3.8 (Alibaba)27B: Apache 2.0; Flash-Next: Qwen Community License 1.0; 2.4T: Qwen3.8-Max License27B dense; Flash-Next, 6B active; 2.4Tnone on the cards; 119 in the Qwen3 launch postQwen
Mistral (Mistral AI)Small 4, Large 3, Ministral 3: Apache 2.0; Medium 3.5: modified MIT3B to 14B; 119B; 128B; 675B“dozens”, German named; Polish and Romanian in Small 4 metadataMistral
Gemma 4 (Google)Apache 2.0; Gemma 3 stays under the Gemma Terms of UseE2B, E4B, 12B, 26B A4B, 31B35+ out of the box, 140+ in pre-training, none namedGemma
Llama 4 (Meta)Llama 4 Community License; multimodal rights withheld from EU-based licensees, end users exceptedScout 109B, Maverick 400B, 17B active each12, German among them, not Polish or RomanianLlama 4
DeepSeekMITV4-Flash, 13B active; V4.1-Flash; V3.2, 685Bnone on the V4-Flash-0731 cardDeepSeek
GLM (Z.ai)GLM-5.3 License; GLM-5.3-Flash and GLM-4.x MIT4.7-Flash 30B; 5.3-Flash 320B; 5.3 753BGLM-5.3 tags: English, ChineseGLM
Nemotron 3 (NVIDIA)Nano, Super: NVIDIA Nemotron Open Model License; Ultra: OpenMDW-1.1Nano 4B and 30B; Super 120B; Ultra 550BSuper: English, French, German, Italian, Japanese, Spanish, ChineseNemotron

Model cards, licence files and metadata on Hugging Face, read on 9 and 10 October 2026; Qwen3 launch post of 29 April 2025; OpenAI’s gpt-oss model card of 5 August 2025. “Languages on card” lists what the vendor names, not tested quality.

The licence is set per checkpoint, not per family. Qwen3.8-27B is Apache 2.0, while Qwen3.8-Flash-Next asks for a separate licence from Qwen if the licensee or an affiliate runs a Model as a Service or AI Work Assistant business, with internal use exempt. Gemma 4 is Apache 2.0, while Gemma 3 stays under the Gemma Terms of Use and their Prohibited Use Policy. Mistral Medium 3.5’s modified MIT licence grants no rights if your company’s monthly revenue exceeds a threshold it states, and the condition extends to derivatives.

For Llama 4, Meta’s Acceptable Use Policy withholds the Section 1(a) rights to the multimodal models from companies with their principal place of business in the EU, except for end users of a product that incorporates them, and Meta’s card calls the Llama 4 models “natively multimodal AI models”. The GLM-5.3 License asks Model-as-a-Service operators above a revenue threshold to pass a security review by Z.AI, as the licence spells the name. Whether a clause applies to your company is a legal assessment for your legal department.

Languages: what the model cards state

A model card lists the languages its vendor names, and it says nothing about quality on your contracts or tickets. Nemotron 3 Super names seven languages, German among them, and Llama 4 names twelve, including German but not Polish or Romanian. Gemma 4 gives 35+ languages out of the box and 140+ in pre-training without naming them.

The cards of Qwen3.8-27B and DeepSeek-V4-Flash-0731 name no languages, and GLM-5.3 carries only the tags English and Chinese. OpenAI reports 83.0 for gpt-oss-120b on German MMMLU at high reasoning effort, a developer’s own figure. For German, Polish and Romanian, European models such as EuroLLM, Bielik and PLLuM are compared in our article on open LLMs for European languages.

Model size, GPUs and concurrent users

The GPUs must hold the checkpoint’s weights plus a KV cache for every conversation in flight. By the estimates in our hub, gpt-oss-120b as released (65.3 GB) and Qwen3.8-27B in FP8 (30.9 GB) run on one RTX PRO 6000 or one H200 NVL, and one DGX Spark (128 GB) holds either for a pilot. Mistral Small 4 in FP8 and Llama 4 Scout in FP8 need two RTX PRO 6000, or one to two H200 NVL. DeepSeek-V4-Flash-0731 needs two RTX PRO 6000 for 20 users at 32K. Llama 4 Maverick needs four H200 NVL or eight RTX PRO 6000, and DeepSeek-V3.2 and GLM-5.3 need eight H200 NVL.

For a whole company the peak of requests in flight sets the card count. Our GPU sizing guide by company size works through 2,000 employees with example values: 40 per cent active in the busiest hour, 6 requests an hour each, 30 seconds per request and a peak factor of 2. That gives 80 requests in flight. With gpt-oss-120b at 32K, one RTX PRO 6000 holds about 19 conversations and one H200 NVL about 55, so the peak needs five RTX PRO 6000 copies or two H200 NVL copies.

Two servers with two H200 NVL each still hold 110 conversations if one fails. With DeepSeek-V4-Flash, a copy on two RTX PRO 6000 holds about 43 conversations at 0.12 GiB of FP8 cache each, so two servers with four cards each keep the peak of 80 if one fails. These are memory estimates, and response time under load needs its own test.

We build AI servers to order with H200 NVL or RTX PRO 6000 Server Edition cards and return a configuration and quote within one business day. Send us the models on your shortlist and your expected peak of users.

Tool calling, context length and reasoning modes

An assistant that looks up tickets, queries a database or calls an internal API needs a model that emits tool calls and an engine that parses them. vLLM’s tool-calling documentation marks --enable-auto-tool-choice as mandatory for this, and --tool-call-parser names the format of the model family.

MODELCARD ON TOOLSVLLM TOOL PARSERCONTEXT
gpt-oss-120b“function calling, web browsing, Python code execution, and Structured Outputs”openai, vLLM docs131,072
Qwen3.8-27B“long-horizon agentic tasks”qwen3_coder, recipe262,144, up to 1M
Mistral Small 4“native function calling and JSON output”mistral, vLLM docs and card256K
Gemma 4 31B“native function-calling support”gemma4, recipe256K
Llama 4 Scoutno statement on the cardllama4_pythonic, vLLM docs10M
DeepSeek-V4-Flash-0731“substantially enhanced agentic capabilities”deepseek_v4, recipe1M
GLM-5.3“complex coding and long-horizon tasks”glm47, recipe1,048,576
Nemotron 3 Super“tool use” among its intended workloadsqwen3_coder, recipe1M

Hugging Face model cards; vLLM tool-calling documentation dated 8 October 2026 and vLLM recipes updated between 11 May and 8 October 2026, read on 10 October 2026; context from the cards and our hardware guides.

Plan for the context your requests need rather than the longest a model accepts, because the KV cache grows with the declared context. On one RTX PRO 6000 gpt-oss-120b holds about 19 conversations at 32K, 79 at 8K and 4 at 128K by our estimate, and the context you set must include retrieved passages and tool results.

Reasoning modes change the time each request stays in flight. gpt-oss sets its reasoning level in the system prompt, for example “Reasoning: high”, Mistral Small 4’s card states reasoning effort “configurable per request”, and for high and maximum reasoning effort DeepSeek recommends up to 384K output tokens. Test each use case at the reasoning level it will run.

Release cadence and the vendor’s update record

Open model families publish new versions several times a year, often within the service life of one server. DeepSeek published V4-Flash as a preview, superseded it with the -0731 release and released V4.1-Flash on 10 September 2026, which our DeepSeek guide places on four H200 NVL with its Engram tables in host memory, as vLLM’s recipe runs it. NVIDIA released Nemotron 3 Nano 30B-A3B on 15 December 2025, Super on 11 March 2026, Ultra on 4 June 2026 and 3.5 Lightning on 11 August 2026, and the last two came under OpenMDW-1.1 rather than the Nemotron Open Model License. Google moved from the Gemma Terms of Use for Gemma 3 to Apache 2.0 for Gemma 4.

A new version can therefore change the licence, the memory footprint and the cards it needs. For each vendor on the shortlist, note how often checkpoints appear, whether earlier ones stay downloadable, whether FP8 or NVFP4 checkpoints and vLLM recipes exist for your cards, and whether the licence stayed the same. Before you swap a model in production, run the regression tests described in our guide to LLM evaluation before model upgrades.

Public leaderboards: what LMArena shows

LMArena’s text leaderboard, which lmarena.ai now forwards to arena.ai under the title “Text Arena”, showed 8,734,330 votes across 414 models on 8 October 2026. It describes its rankings as covering “text-to-text tasks across math, coding, creative writing, and other open-ended domains”. Each entry has a score with a confidence interval, a rank spread and a licence column; the licences in this article come from the licensors’ own files.

On that date gemma-4-31b had a score of 1452 ±7, with a rank spread of 55 to 95. qwen3.8-27b had 1438 ±5 and a spread of 81 to 118. gpt-oss-120b had 1352 ±4, and nvidia-nemotron-3-super-120b-a12b had 1360 ±7. deepseek-v4-flash had 1436 ±4 and a spread of 86 to 119, and the page does not say whether this entry is the preview or the -0731 release. A separate entry, deepseek-v4-flash-high-preview, is the preview at a set reasoning level.

A leaderboard helps to build a shortlist, but it holds none of your documents, languages or tools. The page does not define the rank spread; spreads that overlap this widely suggest that neighbouring models are not reliably ordered, so the decision rests on your own test.

Selection procedure for a company assistant

  1. Write down the use cases, the languages, the peak of users, the context each request needs and the tools the assistant must call.
  2. Read the licence of each candidate checkpoint, record the date read and pass it to your legal department.
  3. Drop models whose weights plus cache at the peak do not fit the cards you plan, using our hardware guides.
  4. Keep two or three candidates and serve each with the same engine, prompts, retrieved passages and context limit.
  5. Build an evaluation set from your own questions in every language, with tool-call cases and their expected calls, and have native speakers grade the answers blind.
  6. Load-test the candidate your graders prefer on the production GPU type at the expected peak, measuring time to first token and per-user speed.
  7. Record the repository, revision and licence of the chosen checkpoint, and run the evaluation set again on every update.

The technical assessment in our Private AI/ML service covers model and GPU selection and a pilot plan with metrics, at a price fixed before work begins. Describe your use cases and languages in the form below.

What we do

Our Private AI/ML service deploys open and commercial models on-premise with vLLM, Ollama or NVIDIA AI Enterprise, or on dedicated hardware in a Tier-3 data centre in Lithuania, and logs queries and answers so that security and legal see who accesses what. The first call is free of charge and leaves you with two or three possible solution scenarios. We start with a pilot on one process with clear metrics and scale only what has proved its value. Eurokommerz holds the contract and supplies the AI servers with H200 NVL, RTX PRO 6000 Server Edition or L40S cards, with engineering by our partner Vixen.UNO.

FAQ

Which open source LLM should a business use for a company chatbot?
Choose an open-weight model by the licence terms for your use, quality in your languages, the GPUs the model needs at your peak of users, tool calling and context length, and the vendor’s update record. As of October 2026, candidates with published hardware figures include gpt-oss-120b, Qwen3.8-27B, Mistral Small 4, Gemma 4, Llama 4, DeepSeek-V4-Flash, GLM-5.3 and Nemotron 3 Super, each under the licence its vendor names. A blind test with your own questions on two or three of them decides.
Which open-weight models can a company use commercially?
As read on 10 October 2026, gpt-oss-120b, Qwen3.8-27B, Gemma 4 and Mistral Small 4 are published under Apache 2.0 and DeepSeek V4-Flash under MIT, both of which grant commercial use. Llama 4, Qwen3.8-Flash-Next, Mistral Medium 3.5, GLM-5.3 and Nemotron 3 Nano and Super come under their vendors’ own licences with conditions such as use policies, revenue thresholds or a separate licence for service businesses, and Nemotron 3 Ultra under OpenMDW-1.1. Whether a clause applies to your company is for your legal department to assess.
How do gpt-oss, Qwen, Mistral and Gemma compare for enterprise use?
gpt-oss-120b, Qwen3.8-27B, Mistral Small 4 and Gemma 4 are all Apache 2.0, with context windows of 131,072, 262,144, 256K and 256K tokens. gpt-oss-120b as released and Qwen3.8-27B in FP8 fit one RTX PRO 6000, Gemma 4 31B fits one but needs a second card for 20 conversations of 32K, and Mistral Small 4 in FP8 needs two RTX PRO 6000. Their cards differ on languages, from no list for gpt-oss to 35+ for Gemma 4, so test them in your own languages.
Which self-hosted LLMs support tool calling?
The gpt-oss, Mistral Small 4 and Gemma 4 cards state function calling, and vLLM’s documentation or recipes give tool-call parsers for gpt-oss, Qwen3.8, Mistral, Gemma 4, Llama 4, DeepSeek-V4-Flash, GLM-5.3 and Nemotron 3 Super. In vLLM the server needs the flags --enable-auto-tool-choice and --tool-call-parser with the parser of the model family.
What hardware does a self-hosted LLM for 2,000 employees need?
With example values of 40 per cent busy-hour users, 6 requests an hour, 30 seconds per request and a peak factor of 2, 2,000 employees produce about 80 requests in flight. For gpt-oss-120b at 32K that needs five RTX PRO 6000 copies or two H200 NVL copies by our estimate, and two servers with two H200 NVL each still hold 110 conversations if one fails. A larger model such as DeepSeek-V4-Flash needs two servers with four RTX PRO 6000 each for the same peak with failover, by our estimate.
Are LLM leaderboards a good way to choose an enterprise model?
LMArena’s text leaderboard of 8 October 2026 listed 414 models with scores, confidence intervals and rank spreads, for example a spread of 55 to 95 for gemma-4-31b. Its rankings cover open-ended text tasks and contain none of your documents, languages or tools, and some entries are previews or run at a set reasoning level, without saying which checkpoint they are. Use it to build a shortlist and decide with a test on your own questions.

Send us your use cases, the languages your staff write in, the expected peak of users and the models you are considering. We reply within one business day, and in the first call we work through 2 to 3 possible solution scenarios with you. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna