BLOG · GUIDE ·

A private ChatGPT alternative for your company: components, sizing and access control

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • A self-hosted alternative to ChatGPT combines a model server with an OpenAI-compatible API such as vLLM, a chat front end such as Open WebUI or LibreChat, sign-in through your identity provider, a document layer that enforces each user’s permissions, a query log and GPU servers sized for peak requests in flight
  • LibreChat is MIT-licensed; Open WebUI’s own licence since v0.6.6 allows internal use without per-user limits but protects its branding above 50 end users in any rolling 30-day period; among models, gpt-oss, Qwen3-32B and Gemma 4 are Apache 2.0, while Llama 4’s use policy withholds its multimodal rights from companies with their principal place of business in the EU
  • Sign-in runs through your identity provider with OpenID Connect in both front ends and SAML natively only in LibreChat; Open WebUI syncs groups from token claims at each login and shares knowledge bases per group, so per-document rights need a filter at retrieval
  • Open WebUI’s audit log is off by default, cuts bodies at 2,048 bytes and excludes the chat paths, and even with them its response bodies hold no answers, which reach the browser over a websocket; a complete record of prompts and answers comes from the saved chats or a gateway, with a retention period decided in advance
  • Size by requests in flight: 0.9 × driver-visible memory minus 3 GiB minus the weights leaves about 22.2 GiB of KV cache for gpt-oss-120b on one RTX PRO 6000, or about 19 conversations at a declared 32K context

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

What a private ChatGPT alternative consists of

A private alternative to ChatGPT runs a ChatGPT-style assistant entirely on servers under your control, on-premise or on dedicated hardware in an EU data centre. Its parts are a model server with an OpenAI-compatible API, such as vLLM, a chat front end with accounts, such as Open WebUI or LibreChat, sign-in through your identity provider, a document layer that applies each user’s access rights, a log of prompts and answers, guardrails, a model whose licence fits your use and GPU servers sized for the peak number of requests in flight.

COMPONENTOPTIONSWHAT TO DECIDE
Model servervLLM, SGLang, NVIDIA NIMmodel, precision, declared context
Chat front endOpen WebUI, LibreChatlicence terms, features per role
Sign-in and rolesOIDC, SAML, LDAP, SCIMlocal login, token lifetime, model access
Documents and RAGuploads, knowledge bases, retrieval servicewhere per-document rights are enforced
Query logchat history, audit log, gatewaywhat is stored, who reads it, how long
Guardrailsfilter functions, guard modelswhich inputs and outputs are checked
ModelApache 2.0 or community licenceslicence terms, languages, memory
GPU serversized by requests in flightcard count, a second host

vLLM, Open WebUI, LibreChat and NeMo Guardrails documentation, read on 6 October 2026.

Model server: vLLM with an OpenAI-compatible API

vLLM, under the Apache 2.0 licence, is described in its documentation as “an HTTP server that implements OpenAI’s Completions API, Chat API, and more”, started with vllm serve and the model name. Its --api-key option checks static keys, one of which the front end presents, on routes such as /v1. vLLM has no user accounts, and its documentation notes that the Chat API’s user parameter is ignored, so only the front end, or a gateway in front of vLLM, knows who asked what. Set --max-model-len to the context conversations need, not the model’s maximum, because it sets how many full-length conversations fit at once. vLLM also “collects anonymous usage data by default”, covering hardware and model configuration, until VLLM_NO_USAGE_STATS=1 is set. Our comparison of LLM serving engines covers SGLang, TensorRT-LLM and Ollama.

Private LLM chat front end: Open WebUI or LibreChat, and their licences

Both front ends connect to any OpenAI-compatible endpoint, Open WebUI under Settings, Admin, Connections with vLLM’s /v1 URL and LibreChat through a custom endpoint with a baseURL in librechat.yaml. Open WebUI searches knowledge bases, optionally with BM25 and reranking; LibreChat uses its RAG API on PostgreSQL with pgvector.

LibreChat is under the MIT licence, its separate admin panel under the AGPL-3.0. Open WebUI has had its own licence since v0.6.6 of 19 April 2025: BSD-3 terms plus a clause that forbids altering, removing, obscuring or replacing its branding unless a deployment has no more than 50 end users “within any rolling thirty (30) day period”, the copyright holder has given written permission or an enterprise licence allows it. Its licence page states “No per-user limits for internal/staff/company-wide use as long as Open WebUI branding is always present and prominent” and that the licence “is not an OSI-approved ‘open source’ license”, which matters where company policy admits only OSI-approved licences.

Single sign-on with OIDC or SAML, roles and groups

Open WebUI and LibreChat both sign users in through OpenID Connect. Open WebUI also documents LDAP and a trusted-header mode for an authenticating reverse proxy, and warns that a wrong configuration “can allow users to authenticate as any user”. We found no native SAML in its documentation; a SAML-only identity provider can be reached through a broker such as Keycloak, which authenticates with external “OpenID Connect or SAML Identity Providers”. LibreChat supports SAML directly, but not together with OpenID Connect, and without Single Logout.

Roles and groups come from the token. With role management on, Open WebUI’s OAUTH_ADMIN_ROLES names the roles that log in as administrators and OAUTH_ALLOWED_ROLES those allowed in at all. ENABLE_OAUTH_GROUP_MANAGEMENT keeps memberships “strictly synchronized” with the token’s claims at each login, which removes users from groups assigned by hand. Permissions are additive, and a group cannot take away what the defaults grant, so the documentation recommends minimal defaults. LibreChat limits sign-in to listed roles with OPENID_REQUIRED_ROLE and manages feature permissions per role, group and user in its separate admin panel, which its documentation describes as “available for testing now”.

New Open WebUI accounts start as pending, and setting ENABLE_LOGIN_FORM to false removes the local password form. Session tokens last 4w by default and, without Redis, stay valid after sign-out; the hardening guide adds that deactivating an account “does not revoke the token already issued to it”, although the role is rechecked on every request. With SCIM 2.0 switched on, the identity provider deactivates users directly, which sets their role to pending.

Document upload, RAG and per-document permissions

Documents arrive as chat uploads or as shared collections the assistant searches. In Open WebUI a knowledge base is shared with groups or users through an access list, and the companion tool oikb mirrors sources such as Git repositories, Confluence spaces and S3 buckets into it. The access list covers the whole knowledge base; we found nothing in the documentation that carries a source system’s per-document permissions into the index.

That suffices where a collection has one audience. Where one repository mixes documents with different readers, permissions have to be enforced at retrieval, by a search that filters on the user’s identity and memberships before ranking, as our article on RAG on company data explains. The front end then has to pass the user’s identity to that retrieval service; in Open WebUI, a filter function receives the full user record.

Assistants built in our Private AI/ML service cite the source and respect each user’s access rights, with data and permissions management and a query log. Write to us with your identity provider and the systems that hold the documents through the form below.

Logging prompts and answers, and how long to keep them

Open WebUI keeps saved conversations in its database, where users can delete them if the Allow Chat Delete permission is granted; temporary chats reduce what is saved, and administrators can read other users’ chats while ENABLE_ADMIN_CHAT_ACCESS keeps its default of true. The separate audit log is off by default; at METADATA it records user, IP address, user agent, method, URL and time, and REQUEST_RESPONSE adds both bodies, cut at 2,048 bytes. Its default exclusions, /chats, /chat and /folders, leave the chat routes out of it. LibreChat deletes temporary chats after 720 hours by default.

A complete query log is a design decision. Even with the chat paths audited, the log records each chat request, which carries the conversation so far, but not the answer: in the browser the HTTP response holds only task IDs, while the answer arrives over a websocket and is saved with the chat. The complete record is either the saved chats, with deletion and temporary chats withheld from users, or a gateway between front end and model server that logs requests and answers; Open WebUI sends it the user’s name, ID, email and role as headers when ENABLE_FORWARD_USER_INFO_HEADERS is true. The European Commission’s guidance is that data “must be stored for the shortest time possible”, with “time limits to erase or review the data stored”. The period itself, and what the GDPR or the AI Act require of an internal assistant, are a legal assessment for your legal department.

The same log documents hybrid use. A public API that some tasks may use is added as one more connection, with its models restricted to the users or groups authorised for them. Keep Open WebUI’s Direct Connections off, because with them “the browser communicates directly with the API provider”, past the server and its log. When a hybrid split pays is the subject of private LLM or cloud API.

Guardrails and open model licences

Guardrails check what goes into the model and what comes out. Open WebUI’s filter functions run an inlet before the model and an outlet before the user, and a filter set as global and active is “force-enabled for all models”. NVIDIA’s NeMo Guardrails, an Apache 2.0 toolkit, adds input, output and retrieval rails, and OpenAI’s gpt-oss-safeguard models, also Apache 2.0, “classify text content based on safety policies that you provide”.

Licences below are taken from each model card and can differ within one family.

MODELLICENCEWHAT TO CHECK
gpt-oss-20b, gpt-oss-120bApache 2.0a usage policy: comply with all applicable law
Qwen3-32B, Qwen3.8-27BApache 2.0Qwen3.8-Flash-Next has the “Qwen Community License 1.0”
Gemma 4, E2B to 31BApache 2.0Gemma 3 stays under the Gemma Terms of Use
Llama 3.3 70BLlama 3.3 community licenceuse policy; “Built with Llama” on products containing it; a licence from Meta above 700 million monthly active users
Llama 4 ScoutLlama 4 community licenceuse policy: multimodal rights withheld in the EU

Model cards, licence files and Meta’s use policies, read on 6 October 2026.

Meta’s Llama 4 acceptable use policy, as read on 6 October 2026, withholds the licence rights for “any multimodal models included in Llama 4” from individuals domiciled in the EU and companies with their principal place of business there, except for “end users of a product or service that incorporates any such multimodal models”. Meta describes the Llama 4 models as “natively multimodal”. Whether such a clause applies to your use is for your legal department.

GPU sizing for concurrent users

Size the GPU by requests in flight at the busiest hour, not by headcount, as our article on concurrent users per RTX PRO 6000 explains. Its rule gives the KV cache 0.9 × the memory the driver reports, minus 3 GiB, minus the weights; vLLM’s gpu_memory_utilization sets that share and now defaults to 0.92. One 96 GB RTX PRO 6000 reports 95.6 GiB, so the 0.9 share is 86.0 GiB. The gpt-oss-120b checkpoint, with its expert weights in MXFP4, takes 65.3 GB, or 60.8 GiB, leaving about 22.2 GiB. Only 18 of its 36 layers keep a cache for the whole context, while the others attend over a 128-token window. A token therefore costs 36 KiB of cache in BF16. At a declared 32K a conversation needs 1.125 GiB, and about 19 fit at once; vLLM prints its own figure at startup.

Each request from the front end is stateless, so the whole conversation is sent with every turn, and a long chat with an uploaded document can fill the declared context; vLLM’s prefix caching reuses the processed history. Open WebUI also calls a model in the background for titles, tags, follow-up suggestions, autocomplete and search queries, and “By default all of that runs on the model the user is chatting with”. Its documentation recommends a small, fast, non-reasoning task model and calls autocomplete “the first thing to turn off”. A task or embedding model runs as a second vLLM instance with its own share of the card’s memory.

We build inference servers for private LLMs with 2 to 8 GPUs per node, sized by model size and concurrent users. Send us the model, the context you plan to declare and your peak requests in flight through the form below.

Rollout in steps: pilot group, policy and training

  1. Choose one process and a pilot group, and agree the metrics the pilot must meet before it grows.
  2. Write the usage policy: data classes allowed in, models and connections per group, what is logged, who reads it and for how long. Our guide to shadow AI policy and controls covers the policy side.
  3. Connect single sign-on with roles and groups, switch off the local login form and vLLM’s usage statistics, leave Direct Connections off and shorten the token lifetime.
  4. Train the pilot group before access: what may go in, how to check a cited source and how to report a wrong answer.
  5. Measure requests in flight, time to first token and answer quality on the pilot group’s questions, and size the production servers from them.
  6. Open the assistant to the next department once the metrics are met, and add a second server before the whole company depends on it.

What we do

Our Private AI/ML service builds this platform under your control, on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania; nothing goes to public APIs unless you enable hybrid mode for a specific task, and what goes there is visible in the query log. Assistants answer from your documents, cite the source and respect each user’s access rights. We start with a pilot on one process with clear metrics, scale only what has proved its value and train your team to run the platform. Eurokommerz holds the contract and supplies the hardware, with engineering by our partner Vixen.UNO; the first call is free of charge, and the price of the technical assessment is fixed before work begins. Data handling during the project is set out on our security and compliance page.

FAQ

What does a self-hosted ChatGPT alternative consist of?
A model server with an OpenAI-compatible API, such as vLLM, and a chat front end with accounts, such as Open WebUI or LibreChat, form the core. Around them sit sign-in through your identity provider by OpenID Connect or SAML, a document layer that enforces each user’s permissions at retrieval, a log of prompts and answers with a retention period, guardrails, a model whose licence fits your use and GPU servers sized for peak requests in flight.
Can a ChatGPT-style assistant run on-premise without sending data out?
Yes, with the model on your own GPU servers and the front end on your network, prompts and documents leave only through connections you add: a public API connection, web search, or Open WebUI’s Direct Connections, with which the browser talks to an API provider directly. Keep Direct Connections switched off and give any public connection only to the users allowed to use it, so its use shows in the query log. vLLM sends anonymous usage statistics by default, which its documentation says contain no sensitive information, until VLLM_NO_USAGE_STATS=1 is set, and Open WebUI downloads its embedding and reranking models and checks for updates unless OFFLINE_MODE is true.
Can a company use Open WebUI without a commercial licence?
Yes for internal use: its licence page states “No per-user limits for internal/staff/company-wide use as long as Open WebUI branding is always present and prominent”. Since v0.6.6 of April 2025, altering, removing or obscuring the Open WebUI branding in a deployment with more than 50 end users in any rolling 30-day period needs the copyright holder’s written permission or an enterprise licence, and the licence page says the licence is not OSI-approved. LibreChat itself is under the MIT licence, without such a clause, and its separate admin panel under the AGPL-3.0.
How do we connect an internal LLM chat to single sign-on?
Open WebUI and LibreChat both sign users in through OpenID Connect; LibreChat also supports SAML, though not at the same time as OpenID Connect, while Open WebUI reaches a SAML-only identity provider through a broker or an authenticating proxy. Map roles and groups from the identity provider into the front end, where Open WebUI synchronises group memberships with the token’s claims at every login, and switch off the local login form. Shorten the session token lifetime too, because Open WebUI’s default is 4w and, without Redis, a token stays valid after sign-out until it expires.
How many GPUs does a private ChatGPT alternative need?
That depends on the peak number of requests in flight and the context you declare, not on headcount. Take 0.9 of the memory the driver reports, subtract 3 GiB and the weights, and divide by the KV cache one conversation needs: on one 96 GB RTX PRO 6000, gpt-oss-120b leaves about 22.2 GiB, enough for about 19 conversations at a declared 32K context. Plan a second card in a second host for availability, whatever the arithmetic says.
Which open models can a company use for an internal chat assistant?
As read on 6 October 2026, gpt-oss-20b and gpt-oss-120b, Qwen3-32B, Qwen3.8-27B and the Gemma 4 models are published under Apache 2.0, while Llama 3.3 and Llama 4 come with Meta’s community licences and acceptable use policies. Llama 4’s policy withholds the rights to its multimodal models from companies with their principal place of business in the EU, and licences can differ within one family, as Qwen3.8-Flash-Next shows. Check the licence on each model card and leave the legal reading to your legal department.

Send us the number of people who would use the assistant, your identity provider, where the documents it should answer from are stored and any model you have in mind. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry – see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna