Open WebUI scaling for 500 to 2,000 users: replicas, Redis, PostgreSQL, SSO and storage
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- One Open WebUI container with SQLite serves a pilot; for multiple instances, Open WebUI’s scaling guide marks PostgreSQL, Redis and an external vector database as required and lists 1,000+ users under the same requirements
- Each replica runs one Uvicorn worker, all replicas share one
WEBUI_SECRET_KEYandOAUTH_, and only one designated pod runs database migrationsSESSION_ TOKEN_ ENCRYPTION_ KEY - Redis carries websockets, locks and token revocation between replicas; the docs call
noevictionthe safe default, accept only volatile policies where memory is capped, recommend append-only persistence and warn of the 1,000-client limit on some distributions - Uploaded files go to S3-compatible storage or a ReadWriteMany volume, vectors to pgvector or another client-server database, and extraction and embeddings to external services such as Tika
- Upgrades stop all replicas, let one instance migrate the schema and then restart the fleet; the docs do not support rolling upgrades across a schema change, and Helm rollback cannot undo migrations
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
Open WebUI scaling: what changes between a pilot and 1,000 users
A pilot of Open WebUI runs as one container with SQLite in its data volume. For a production deployment of 500 to 2,000 users, Open WebUI’s scaling guide moves the deployment to PostgreSQL, adds Redis for websockets and shared state, runs several replicas behind a load balancer, moves the vector index into an external database and takes document extraction and embeddings out of the application process. Its quick reference marks PostgreSQL, Redis and an external vector database as required for multiple instances and lists large deployments of “1000+ users” with the same requirements. The settings below are taken from the Open WebUI documentation as read on 10 October 2026, when v0.11.4 was the latest release on the project’s GitHub page. The Kubernetes guide pins Helm chart 16.5.0 with v0.11.3, and the environment variable reference is up to date with v0.11.1.
| COMPONENT | PILOT | 500 TO 2,000 USERS | SETTINGS |
|---|---|---|---|
| Database | SQLite in the data volume | PostgreSQL with a sized pool | DATABASE_, DATABASE_ |
| Websockets and state | inside the one process | Redis shared by all replicas | REDIS_URL, WEBSOCKET_ |
| Instances | one container | several replicas, one worker each | UVICORN_, one secret key |
| Schema migrations | at every start | one designated pod | ENABLE_ elsewhere |
| Uploaded files | local data directory | S3-compatible bucket or RWX volume | STORAGE_ |
| Vector index | local ChromaDB | pgvector or a client-server database | VECTOR_ |
| Extraction, embeddings | pypdf, Sentence | Tika, an external embedding endpoint | CONTENT_ |
| Sign-in | local accounts | OIDC with role and group claims | ENABLE_ |
Open WebUI documentation: scaling guide, multi-replica troubleshooting, Redis, SSO and Kubernetes pages, read on 10 October 2026.
The choice of front end and the basics of sign-in are covered in our guide to a private ChatGPT alternative.
PostgreSQL and the connection pool
The scaling guide says to switch first: “Before adding replicas or workers, replace the embedded database configuration with PostgreSQL”. DATABASE_URL points to the PostgreSQL database. Several instances on one SQLite file produce “database is locked” errors and data corruption, and SQLite must not sit on a network filesystem, because its file locking is unreliable over NFS. The guide also notes that Open WebUI “does not migrate data between databases”, so decide before the switch whether pilot chats move to PostgreSQL or the production platform starts empty.
Each replica keeps its own connection pool. The guide suggests DATABASE_POOL_SIZE=15 and DATABASE_POOL_MAX_OVERFLOW=20 as a starting point and asks to keep the combined total per instance “well below your PostgreSQL max_connections limit (default is 100)”. By our arithmetic, one replica can then open up to 35 connections, so three replicas reach 105 before migrations, backups or monitoring connect. PostgreSQL 18 also reserves three connections for superusers by default. Either raise max_connections on the database server, which takes effect only at server start, or lower the pool per replica.
The vectors can stay in the same PostgreSQL server: the Kubernetes guide’s production profile sets VECTOR_DB to pgvector with the same database URL, and leaves creating the vector extension to a database administrator. Our vector database comparison covers when pgvector stops being enough.
Redis for websockets, token revocation and shared state
With WEBSOCKET_MANAGER=redis, Open WebUI runs Socket.IO over Redis Pub/Sub, so an answer streamed by one replica reaches a browser connected to another. The Redis page lists REDIS_URL for application state, ENABLE_WEBSOCKET_SUPPORT=true, WEBSOCKET_MANAGER=redis and WEBSOCKET_REDIS_URL, and REDIS_KEY_PREFIX where several Open WebUI instances share one Redis server. Redis also holds locks, the websocket session and usage pools and the list of revoked tokens. The documentation warns that without Redis “signing out does not invalidate a user’s JWT token”, which also applies to a single instance.
| REDIS SETTING | DOCUMENTED VALUE | REASON GIVEN |
|---|---|---|
| Eviction policy | noeviction | allkeys policies can un-revoke tokens |
| With a memory cap | volatile-ttl, volatile-lru | only keys with an expiry are evicted |
| Persistence | appendonly yes, everysec | loss bounded to about one second |
| Client limit | maxclients 10000 | 1,000 clients on some distributions |
| Idle timeout | timeout 1800 | closes idle TCP connections only |
| Connect timeout | REDIS_ | needed with Sentinel |
Open WebUI documentation, Redis page, read on 10 October 2026.
For availability, Open WebUI connects to Redis Sentinel through REDIS_SENTINEL_HOSTS or to a Redis Cluster with REDIS_CLUSTER=true; Sentinel takes precedence when both are set.
Multiple replicas behind a load balancer
The scaling guide describes Open WebUI as stateless once the shared services are in place, so “you can run as many instances as needed behind a load balancer”. Keep UVICORN_WORKERS=1 per container and let the orchestrator add replicas, because every worker is a full Open WebUI process that runs the startup migrations. Set ENABLE_DB_MIGRATIONS=false on all replicas except one designated pod. All replicas need the same WEBUI_SECRET_KEY; with different keys, a token issued by one replica fails on the next and users see login loops and 401 errors. The SSO page adds that OAUTH_ “Must be shared across all instances in a cluster”.
The load balancer or ingress needs websockets, streamed responses without buffering, suitable request and idle timeouts and the upload size you allow. Every public hostname goes into CORS_ALLOW_ORIGIN, or websocket connections fail on an origin mismatch. Session affinity is optional: the Kubernetes guide says it “may be needed for Socket.IO polling” and that “Affinity does not replace Redis”. The scaling guide also raises THREAD_POOL_SIZE to 2000, because the default ceiling “is only 40”.
On Kubernetes, the official chart’s production profile starts three replicas with one worker each, Redis and PostgreSQL provided outside the chart, uploads in an S3 bucket and ingress enabled. The guide states that “User count alone does not determine the number of replicas” and places replicas “on different nodes when you need resilience to a node failure”. Chart 16.5.0 does not create a HorizontalPodAutoscaler or a PodDisruptionBudget through dedicated values, so both are applied as separate resources. ENABLE_OTEL=true with OTEL_EXPORTER_OTLP_ENDPOINT exports traces, metrics and logs. The model servers behind Open WebUI are a separate layer, covered in our article on a private LLM platform on Kubernetes.
Our Private AI/ML service builds a Kubernetes-based platform on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania. Describe your pilot’s layout and the user count you plan for in the form below.
Files, vector index and document extraction
Uploads written by one replica must be readable by all others, or users see “Image unavailable” placeholders and files the model cannot find. The scaling guide’s quick reference marks shared storage “Optional (NFS or S3)” for multiple instances and for “1000+ users”, and its file storage step answers “Not necessarily” to the question whether S3 is needed, since a shared filesystem mount such as NFS or CephFS is enough. The multi-replica page lists shared storage among the “absolute requirements” for a multi-replica setup, and the Kubernetes production profile keeps uploads in an S3-compatible bucket. Read together, the pages leave the choice between a shared filesystem and object storage open and expect every replica to see the same files. Open WebUI names files by UUID, so replicas can share a directory without write conflicts. Mount a ReadWriteMany volume at /app/backend/data on every replica, or set STORAGE_PROVIDER to an object store such as S3. With object storage, the Kubernetes guide treats /app/backend/data as ephemeral: files there outside PostgreSQL and the bucket do not survive a pod replacement.
The default ChromaDB runs inside the application on SQLite. The multi-replica page calls it “not safe for multi-worker or multi-replica deployments”, and the fix is ChromaDB as a separate HTTP server or a client-server database through VECTOR_DB, such as pgvector, Milvus or Qdrant. The scaling guide states that “Only PGVector and ChromaDB will be consistently maintained by the Open WebUI team”.
The scaling guide names the default extraction engine, pypdf, and the default embedding engine, SentenceTransformers, as the two most common causes of memory leaks in production. It moves extraction to Apache Tika with CONTENT_EXTRACTION_ENGINE=tika and TIKA_SERVER_URL, and embeddings to an external endpoint with RAG_EMBEDDING_ENGINE, such as an embedding model served by vLLM.
SSO with OIDC, roles and model access per group
At this scale accounts come from the identity provider. Open WebUI reads the OIDC discovery URL from OPENID_PROVIDER_URL, restricts sign-in to roles in OAUTH_ALLOWED_ROLES and grants administrator rights to OAUTH_ADMIN_ROLES when role management is on. With ENABLE_OAUTH_GROUP_MANAGEMENT, memberships follow the groups claim at each sign-in. The SSO page warns that users are removed from groups “including those manually created or assigned within Open WebUI”, with “no role exemption”, and that a change in the identity provider shows after the user signs out and in again. Groups listed in OAUTH_BLOCKED_GROUPS are never added or removed.
Models, knowledge bases and tools set to private are shared with groups at read or write level. Permissions are additive and deny rules do not exist, so the groups page advises disabling all permissions in sharing groups and setting feature rights in the global defaults. How to map departments and waves to these groups is the subject of our rollout plan for 1,000 employees. Where model access also needs keys per team, token budgets and a log, the connection from Open WebUI points at an LLM gateway instead of the model servers.
Two settings decide whether the environment variables or the admin panel win. ENABLE_PERSISTENT_CONFIG defaults to true, and variables marked as ConfigVar are then read only at the first launch and stored in the database. ENABLE_OAUTH_PERSISTENT_CONFIG defaults to false, which “keeps environment variables authoritative for OAuth”, and both must be true to manage OAuth from the admin panel. Decide which side wins before the first production start.
Licence terms for branding
Since v0.6.6 of 19 April 2025, Open WebUI’s licence forbids altering, removing, obscuring or replacing its branding, described on the licence page as “(name, logo, UI marks, etc.)”, unless a deployment has no more than 50 end users “within any rolling thirty (30) day period”, the copyright holder has given written permission or an enterprise licence allows it. Internal use has “No per-user limits for internal/staff/company-wide use as long as Open WebUI branding is always present and prominent”. WEBUI_NAME sets the name shown and, per the variable reference, appends “(Open WebUI)” when overridden. Whether a planned custom theme stays within these terms is a legal assessment for the company’s legal department.
Backups and upgrades
A backup covers the PostgreSQL database, which holds chats, users, settings and, with pgvector, the vectors, and the bucket or volume with uploads. Open WebUI’s backup page names pg_dump for PostgreSQL, which, per the PostgreSQL documentation, “makes consistent exports even if the database is being used concurrently”. Copy the uploads right after the dump, so that file records and files match. Keep the secret keys and the environment in your secrets store, because tokens and stored OAuth sessions depend on them.
Upgrades follow a fixed order. The multi-replica page states that “Old and new Open WebUI instances must never serve traffic against the same database at the same time”: back up the database, stop all replicas, let one instance migrate the schema, then start the fleet on the new version. Rolling or canary upgrades across a schema change are not supported. The Kubernetes guide scales the deployment to zero, starts one migration-enabled replica with ingress disabled, checks it and then restores the production values, and warns that Helm rollback, including --atomic, “cannot undo database migrations”. Test each upgrade on a staging copy of the database first.
We build and support the platform under an agreed SLA, with changes in agreed maintenance windows and a rollback plan. Tell us which Open WebUI version your pilot runs and how often you plan to upgrade.
Moving from a pilot to the production layout
- Freeze the pilot’s version, record its settings and decide whether its chats and users move to PostgreSQL.
- Provide PostgreSQL with pgvector, set
max_connectionsfor the replicas and pools you plan, and create thevectorextension. - Provide Redis with
noeviction, append-only persistence and Sentinel or Cluster for availability. - Create the bucket or ReadWriteMany volume for uploads, and run Tika and the embedding endpoint as separate services.
- Generate one
WEBUI_SECRET_KEYand one OAuth session key, store them as secrets and decide whether configuration lives in variables or in the database. - Start one replica with migrations enabled and ingress disabled, connect OIDC with role and group claims and test sign-in, sharing and uploads.
- Raise the replica count with migrations disabled, open the ingress with websockets and set
CORS_ALLOW_ORIGIN. - Re-index the knowledge bases into the new vector store, connect OpenTelemetry and write the backup and upgrade runbook.
What we do
Our Private AI/ML service builds an AI platform under your control, with a query log and data and permissions management. We start with a pilot on one process with clear metrics and train your team to run the platform. Eurokommerz holds the contract and supplies the GPU servers, from a single server to a cluster, with engineering by our partner Vixen.UNO and ongoing support under an agreed SLA; the price of the technical assessment is fixed before work begins.
FAQ
How do you scale Open WebUI for many users?
Can Open WebUI run with multiple replicas?
WEBUI_SECRET_KEY and OAuth session encryption key, and only one designated pod should run database migrations.Does Open WebUI need Redis?
noeviction the safe default, accepts only volatile-ttl or volatile-lru where memory must be capped and recommends append-only persistence.Should Open WebUI use PostgreSQL or SQLite in production?
max_connections. Open WebUI does not migrate existing data from SQLite to PostgreSQL itself.How do you set up enterprise SSO for Open WebUI?
OPENID_PROVIDER_URL, the client ID and secret, and turn on role and group management so that roles and memberships follow the token’s claims. Group sync removes users from groups that the claim does not list, including groups assigned by hand and with no exemption for administrators, and a change shows after the next sign-in. Models and knowledge bases are then shared per group.How do you run Open WebUI on Kubernetes with high availability?
Send us the number of users, your identity provider, the Open WebUI version and layout of your pilot and where the platform should run. We reply within one business day, and after the first call you leave with 2 to 3 possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day