BLOG · GUIDE ·

Open WebUI scaling for 500 to 2,000 users: replicas, Redis, PostgreSQL, SSO and storage

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • One Open WebUI container with SQLite serves a pilot; for multiple instances, Open WebUI’s scaling guide marks PostgreSQL, Redis and an external vector database as required and lists 1,000+ users under the same requirements
  • Each replica runs one Uvicorn worker, all replicas share one WEBUI_SECRET_KEY and OAUTH_SESSION_TOKEN_ENCRYPTION_KEY, and only one designated pod runs database migrations
  • Redis carries websockets, locks and token revocation between replicas; the docs call noeviction the safe default, accept only volatile policies where memory is capped, recommend append-only persistence and warn of the 1,000-client limit on some distributions
  • Uploaded files go to S3-compatible storage or a ReadWriteMany volume, vectors to pgvector or another client-server database, and extraction and embeddings to external services such as Tika
  • Upgrades stop all replicas, let one instance migrate the schema and then restart the fleet; the docs do not support rolling upgrades across a schema change, and Helm rollback cannot undo migrations

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

Open WebUI scaling: what changes between a pilot and 1,000 users

A pilot of Open WebUI runs as one container with SQLite in its data volume. For a production deployment of 500 to 2,000 users, Open WebUI’s scaling guide moves the deployment to PostgreSQL, adds Redis for websockets and shared state, runs several replicas behind a load balancer, moves the vector index into an external database and takes document extraction and embeddings out of the application process. Its quick reference marks PostgreSQL, Redis and an external vector database as required for multiple instances and lists large deployments of “1000+ users” with the same requirements. The settings below are taken from the Open WebUI documentation as read on 10 October 2026, when v0.11.4 was the latest release on the project’s GitHub page. The Kubernetes guide pins Helm chart 16.5.0 with v0.11.3, and the environment variable reference is up to date with v0.11.1.

COMPONENTPILOT500 TO 2,000 USERSSETTINGS
DatabaseSQLite in the data volumePostgreSQL with a sized poolDATABASE_URL, DATABASE_POOL_SIZE
Websockets and stateinside the one processRedis shared by all replicasREDIS_URL, WEBSOCKET_MANAGER=redis
Instancesone containerseveral replicas, one worker eachUVICORN_WORKERS=1, one secret key
Schema migrationsat every startone designated podENABLE_DB_MIGRATIONS=false elsewhere
Uploaded fileslocal data directoryS3-compatible bucket or RWX volumeSTORAGE_PROVIDER=s3
Vector indexlocal ChromaDBpgvector or a client-server databaseVECTOR_DB=pgvector
Extraction, embeddingspypdf, SentenceTransformersTika, an external embedding endpointCONTENT_EXTRACTION_ENGINE=tika
Sign-inlocal accountsOIDC with role and group claimsENABLE_OAUTH_GROUP_MANAGEMENT

Open WebUI documentation: scaling guide, multi-replica troubleshooting, Redis, SSO and Kubernetes pages, read on 10 October 2026.

The choice of front end and the basics of sign-in are covered in our guide to a private ChatGPT alternative.

PostgreSQL and the connection pool

The scaling guide says to switch first: “Before adding replicas or workers, replace the embedded database configuration with PostgreSQL”. DATABASE_URL points to the PostgreSQL database. Several instances on one SQLite file produce “database is locked” errors and data corruption, and SQLite must not sit on a network filesystem, because its file locking is unreliable over NFS. The guide also notes that Open WebUI “does not migrate data between databases”, so decide before the switch whether pilot chats move to PostgreSQL or the production platform starts empty.

Each replica keeps its own connection pool. The guide suggests DATABASE_POOL_SIZE=15 and DATABASE_POOL_MAX_OVERFLOW=20 as a starting point and asks to keep the combined total per instance “well below your PostgreSQL max_connections limit (default is 100)”. By our arithmetic, one replica can then open up to 35 connections, so three replicas reach 105 before migrations, backups or monitoring connect. PostgreSQL 18 also reserves three connections for superusers by default. Either raise max_connections on the database server, which takes effect only at server start, or lower the pool per replica.

The vectors can stay in the same PostgreSQL server: the Kubernetes guide’s production profile sets VECTOR_DB to pgvector with the same database URL, and leaves creating the vector extension to a database administrator. Our vector database comparison covers when pgvector stops being enough.

Redis for websockets, token revocation and shared state

With WEBSOCKET_MANAGER=redis, Open WebUI runs Socket.IO over Redis Pub/Sub, so an answer streamed by one replica reaches a browser connected to another. The Redis page lists REDIS_URL for application state, ENABLE_WEBSOCKET_SUPPORT=true, WEBSOCKET_MANAGER=redis and WEBSOCKET_REDIS_URL, and REDIS_KEY_PREFIX where several Open WebUI instances share one Redis server. Redis also holds locks, the websocket session and usage pools and the list of revoked tokens. The documentation warns that without Redis “signing out does not invalidate a user’s JWT token”, which also applies to a single instance.

REDIS SETTINGDOCUMENTED VALUEREASON GIVEN
Eviction policynoevictionallkeys policies can un-revoke tokens
With a memory capvolatile-ttl, volatile-lruonly keys with an expiry are evicted
Persistenceappendonly yes, everysecloss bounded to about one second
Client limitmaxclients 100001,000 clients on some distributions
Idle timeouttimeout 1800closes idle TCP connections only
Connect timeoutREDIS_SOCKET_CONNECT_TIMEOUT=5needed with Sentinel

Open WebUI documentation, Redis page, read on 10 October 2026.

For availability, Open WebUI connects to Redis Sentinel through REDIS_SENTINEL_HOSTS or to a Redis Cluster with REDIS_CLUSTER=true; Sentinel takes precedence when both are set.

Multiple replicas behind a load balancer

The scaling guide describes Open WebUI as stateless once the shared services are in place, so “you can run as many instances as needed behind a load balancer”. Keep UVICORN_WORKERS=1 per container and let the orchestrator add replicas, because every worker is a full Open WebUI process that runs the startup migrations. Set ENABLE_DB_MIGRATIONS=false on all replicas except one designated pod. All replicas need the same WEBUI_SECRET_KEY; with different keys, a token issued by one replica fails on the next and users see login loops and 401 errors. The SSO page adds that OAUTH_SESSION_TOKEN_ENCRYPTION_KEY “Must be shared across all instances in a cluster”.

The load balancer or ingress needs websockets, streamed responses without buffering, suitable request and idle timeouts and the upload size you allow. Every public hostname goes into CORS_ALLOW_ORIGIN, or websocket connections fail on an origin mismatch. Session affinity is optional: the Kubernetes guide says it “may be needed for Socket.IO polling” and that “Affinity does not replace Redis”. The scaling guide also raises THREAD_POOL_SIZE to 2000, because the default ceiling “is only 40”.

On Kubernetes, the official chart’s production profile starts three replicas with one worker each, Redis and PostgreSQL provided outside the chart, uploads in an S3 bucket and ingress enabled. The guide states that “User count alone does not determine the number of replicas” and places replicas “on different nodes when you need resilience to a node failure”. Chart 16.5.0 does not create a HorizontalPodAutoscaler or a PodDisruptionBudget through dedicated values, so both are applied as separate resources. ENABLE_OTEL=true with OTEL_EXPORTER_OTLP_ENDPOINT exports traces, metrics and logs. The model servers behind Open WebUI are a separate layer, covered in our article on a private LLM platform on Kubernetes.

Our Private AI/ML service builds a Kubernetes-based platform on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania. Describe your pilot’s layout and the user count you plan for in the form below.

Files, vector index and document extraction

Uploads written by one replica must be readable by all others, or users see “Image unavailable” placeholders and files the model cannot find. The scaling guide’s quick reference marks shared storage “Optional (NFS or S3)” for multiple instances and for “1000+ users”, and its file storage step answers “Not necessarily” to the question whether S3 is needed, since a shared filesystem mount such as NFS or CephFS is enough. The multi-replica page lists shared storage among the “absolute requirements” for a multi-replica setup, and the Kubernetes production profile keeps uploads in an S3-compatible bucket. Read together, the pages leave the choice between a shared filesystem and object storage open and expect every replica to see the same files. Open WebUI names files by UUID, so replicas can share a directory without write conflicts. Mount a ReadWriteMany volume at /app/backend/data on every replica, or set STORAGE_PROVIDER to an object store such as S3. With object storage, the Kubernetes guide treats /app/backend/data as ephemeral: files there outside PostgreSQL and the bucket do not survive a pod replacement.

The default ChromaDB runs inside the application on SQLite. The multi-replica page calls it “not safe for multi-worker or multi-replica deployments”, and the fix is ChromaDB as a separate HTTP server or a client-server database through VECTOR_DB, such as pgvector, Milvus or Qdrant. The scaling guide states that “Only PGVector and ChromaDB will be consistently maintained by the Open WebUI team”.

The scaling guide names the default extraction engine, pypdf, and the default embedding engine, SentenceTransformers, as the two most common causes of memory leaks in production. It moves extraction to Apache Tika with CONTENT_EXTRACTION_ENGINE=tika and TIKA_SERVER_URL, and embeddings to an external endpoint with RAG_EMBEDDING_ENGINE, such as an embedding model served by vLLM.

SSO with OIDC, roles and model access per group

At this scale accounts come from the identity provider. Open WebUI reads the OIDC discovery URL from OPENID_PROVIDER_URL, restricts sign-in to roles in OAUTH_ALLOWED_ROLES and grants administrator rights to OAUTH_ADMIN_ROLES when role management is on. With ENABLE_OAUTH_GROUP_MANAGEMENT, memberships follow the groups claim at each sign-in. The SSO page warns that users are removed from groups “including those manually created or assigned within Open WebUI”, with “no role exemption”, and that a change in the identity provider shows after the user signs out and in again. Groups listed in OAUTH_BLOCKED_GROUPS are never added or removed.

Models, knowledge bases and tools set to private are shared with groups at read or write level. Permissions are additive and deny rules do not exist, so the groups page advises disabling all permissions in sharing groups and setting feature rights in the global defaults. How to map departments and waves to these groups is the subject of our rollout plan for 1,000 employees. Where model access also needs keys per team, token budgets and a log, the connection from Open WebUI points at an LLM gateway instead of the model servers.

Two settings decide whether the environment variables or the admin panel win. ENABLE_PERSISTENT_CONFIG defaults to true, and variables marked as ConfigVar are then read only at the first launch and stored in the database. ENABLE_OAUTH_PERSISTENT_CONFIG defaults to false, which “keeps environment variables authoritative for OAuth”, and both must be true to manage OAuth from the admin panel. Decide which side wins before the first production start.

Licence terms for branding

Since v0.6.6 of 19 April 2025, Open WebUI’s licence forbids altering, removing, obscuring or replacing its branding, described on the licence page as “(name, logo, UI marks, etc.)”, unless a deployment has no more than 50 end users “within any rolling thirty (30) day period”, the copyright holder has given written permission or an enterprise licence allows it. Internal use has “No per-user limits for internal/staff/company-wide use as long as Open WebUI branding is always present and prominent”. WEBUI_NAME sets the name shown and, per the variable reference, appends “(Open WebUI)” when overridden. Whether a planned custom theme stays within these terms is a legal assessment for the company’s legal department.

Backups and upgrades

A backup covers the PostgreSQL database, which holds chats, users, settings and, with pgvector, the vectors, and the bucket or volume with uploads. Open WebUI’s backup page names pg_dump for PostgreSQL, which, per the PostgreSQL documentation, “makes consistent exports even if the database is being used concurrently”. Copy the uploads right after the dump, so that file records and files match. Keep the secret keys and the environment in your secrets store, because tokens and stored OAuth sessions depend on them.

Upgrades follow a fixed order. The multi-replica page states that “Old and new Open WebUI instances must never serve traffic against the same database at the same time”: back up the database, stop all replicas, let one instance migrate the schema, then start the fleet on the new version. Rolling or canary upgrades across a schema change are not supported. The Kubernetes guide scales the deployment to zero, starts one migration-enabled replica with ingress disabled, checks it and then restores the production values, and warns that Helm rollback, including --atomic, “cannot undo database migrations”. Test each upgrade on a staging copy of the database first.

We build and support the platform under an agreed SLA, with changes in agreed maintenance windows and a rollback plan. Tell us which Open WebUI version your pilot runs and how often you plan to upgrade.

Moving from a pilot to the production layout

  1. Freeze the pilot’s version, record its settings and decide whether its chats and users move to PostgreSQL.
  2. Provide PostgreSQL with pgvector, set max_connections for the replicas and pools you plan, and create the vector extension.
  3. Provide Redis with noeviction, append-only persistence and Sentinel or Cluster for availability.
  4. Create the bucket or ReadWriteMany volume for uploads, and run Tika and the embedding endpoint as separate services.
  5. Generate one WEBUI_SECRET_KEY and one OAuth session key, store them as secrets and decide whether configuration lives in variables or in the database.
  6. Start one replica with migrations enabled and ingress disabled, connect OIDC with role and group claims and test sign-in, sharing and uploads.
  7. Raise the replica count with migrations disabled, open the ingress with websockets and set CORS_ALLOW_ORIGIN.
  8. Re-index the knowledge bases into the new vector store, connect OpenTelemetry and write the backup and upgrade runbook.

What we do

Our Private AI/ML service builds an AI platform under your control, with a query log and data and permissions management. We start with a pilot on one process with clear metrics and train your team to run the platform. Eurokommerz holds the contract and supplies the GPU servers, from a single server to a cluster, with engineering by our partner Vixen.UNO and ongoing support under an agreed SLA; the price of the technical assessment is fixed before work begins.

FAQ

How do you scale Open WebUI for many users?
Move from SQLite to PostgreSQL, add Redis for websockets and shared state, run several replicas with one Uvicorn worker each behind a load balancer and share one secret key between them. Put uploads in S3-compatible storage or a ReadWriteMany volume, vectors in pgvector or another client-server database, and extraction and embeddings in external services. Open WebUI’s scaling guide lists deployments of 1,000+ users with the same requirements as multiple instances.
Can Open WebUI run with multiple replicas?
Yes. Open WebUI’s documentation describes it as stateless once PostgreSQL, Redis and an external vector database are in place, so replicas can run behind a load balancer. All replicas need the same WEBUI_SECRET_KEY and OAuth session encryption key, and only one designated pod should run database migrations.
Does Open WebUI need Redis?
For multiple replicas or workers, yes: the scaling guide marks Redis as required, because websockets, locks and token revocation are shared through it. A single instance runs without Redis, but then signing out does not invalidate a user’s token. The Redis page calls noeviction the safe default, accepts only volatile-ttl or volatile-lru where memory must be capped and recommends append-only persistence.
Should Open WebUI use PostgreSQL or SQLite in production?
SQLite suits a single instance and evaluation; for multiple instances the scaling guide requires PostgreSQL and says to switch before adding replicas or workers. Each replica keeps its own connection pool, so the suggested starting values of 15 plus 20 overflow connections per replica, summed over all replicas, must stay below the server’s max_connections. Open WebUI does not migrate existing data from SQLite to PostgreSQL itself.
How do you set up enterprise SSO for Open WebUI?
Connect the identity provider through OpenID Connect with OPENID_PROVIDER_URL, the client ID and secret, and turn on role and group management so that roles and memberships follow the token’s claims. Group sync removes users from groups that the claim does not list, including groups assigned by hand and with no exemption for administrators, and a change shows after the next sign-in. Models and knowledge bases are then shared per group.
How do you run Open WebUI on Kubernetes with high availability?
The official Helm chart’s production profile runs three replicas with one worker each, external PostgreSQL with pgvector, external Redis and uploads in an S3 bucket. Spread the replicas across nodes and add a PodDisruptionBudget and an autoscaler yourself, because chart 16.5.0 does not create them through dedicated values. Upgrades scale the deployment to zero, let one replica migrate the schema and then restore the replicas.

Send us the number of users, your identity provider, the Open WebUI version and layout of your pilot and where the platform should run. We reply within one business day, and after the first call you leave with 2 to 3 possible solution scenarios. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna