LLM gateway for a company: one API, keys per team, token budgets and a query log
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- An LLM gateway gives every application one OpenAI-compatible endpoint, checks a key per application or team, enforces rate limits and token budgets, routes to private models or, where enabled, to an external API, and logs each request
- vLLM ignores the
userparameter of its Chat Completions API and has no user accounts, so only a gateway that is the only way to the model server knows which team asked what - LiteLLM’s proxy is MIT-licensed outside its enterprise directory; JWT sign-in through the identity provider, key rotation and per-model key budgets need its Enterprise licence
- Envoy AI Gateway has been called Agent Router since 9 September 2026 and stays under Apache 2.0, with token rate limits that need Redis; Kong’s AI Proxy is in its Apache 2.0 gateway code, while token rate limiting and the PII sanitiser need AI Gateway Enterprise
- Fallback lists should name private models only, the per-request opt-out from logging should be disabled, and the log’s contents and retention period are decided before the first request
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
What an LLM gateway does in a company platform
An LLM gateway, also called an AI gateway, is a proxy through which every application reaches the language models. It offers one OpenAI-compatible endpoint, checks the caller’s key, and routes each request to a private model or, only where the company has enabled it, to an external API. It enforces rate limits and token budgets per application and team and writes each request to a query log kept under a retention policy. On a platform with 500 to 2,000 users and several applications, it is also the one component that knows which team consumed how many tokens.
| FUNCTION | WHY IT MATTERS | WHERE DOCUMENTED |
|---|---|---|
| One OpenAI-style endpoint | applications change models without code changes | LiteLLM proxy, Kong AI Proxy, Agent Router |
| Keys per application, team | one key is revoked without affecting other users | LiteLLM virtual keys |
| Identity provider sign-in | access follows the directory account | LiteLLM JWT auth, Envoy Gateway Security |
| Rate limits | a batch job cannot crowd out interactive chat | LiteLLM tpm and rpm limits, Agent Router token limits |
| Token budgets | consumption per department and period | LiteLLM max_budget and budget_ |
| Fallback routing | requests move on when a model fails | LiteLLM fallbacks |
| PII masking | names and IDs masked, or the request blocked | LiteLLM with Presidio, Kong AI Sanitizer |
| Query log, retention | who asked what, kept for a set period | LiteLLM spend logs and retention cleanup |
LiteLLM, Agent Router, Envoy Gateway and Kong documentation, read on 10 October 2026.
Why the model server needs a gateway in front
vLLM checks static API keys but has no user accounts, and its documentation notes that the user parameter of the Chat Completions API is ignored. Several applications that share one vLLM key look identical in its logs, and revoking the key stops all of them. The gateway issues its own keys, knows which application and team each one belongs to and passes the request on with a key that only the gateway holds. This works only if the gateway is the only way in. On Kubernetes, network policies let only the gateway, the endpoint picker and Prometheus reach the serving port, as our guide to a private LLM platform on Kubernetes sets out.
A worked example shows the scale; the figures are an illustration, not a sizing rule. A company of 1,500 staff runs a chat front end for all employees, a RAG assistant for legal and sales, a coding assistant for 60 developers, workflows in n8n and a nightly batch that extracts fields from incoming documents. With one key per application in test and in production, the gateway holds ten service keys, plus personal keys for developers whose IDE plugins call the API directly. The chat front end, whose sign-in our guide to a private ChatGPT alternative covers, has to pass the signed-in user’s identity with each request, or the log records only the application.
LLM gateway options: LiteLLM, Agent Router (Envoy AI Gateway) and Kong
| GATEWAY | LICENCE | STATED LIMITS |
|---|---|---|
| LiteLLM proxy | MIT, except its enterprise directory | JWT auth, key rotation and per-model key budgets need an Enterprise licence |
| Agent Router | Apache 2.0 | formerly Envoy AI Gateway; token rate limits need Redis |
| Kong AI Gateway | AI Proxy: Apache 2.0 code | AI Rate Limiting Advanced and AI Sanitizer are AI Gateway Enterprise only |
| Inference Extension | Apache 2.0 | its endpoint picker chooses a replica; keys and budgets come from the gateway |
LiteLLM, Kong Gateway and Inference Extension LICENSE files, LiteLLM and Agent Router documentation, the Agentic AI Foundation’s post of 9 September 2026, Kong plugin pages and the Gateway API Inference Extension site, read on 10 October 2026.
LiteLLM’s proxy keeps keys, teams and spend in a Postgres database, which its documentation requires for virtual keys. Its licence file says that content outside the enterprise directory “is available under the MIT license”, while that directory has a licence file of its own. The documentation marks several features as requiring “a LiteLLM Enterprise license”, among them JWT-based sign-in, key rotation, budgets per model on a key and logging to some destinations, such as cloud storage buckets and custom callback APIs. Single sign-on for its admin interface works without that licence for up to 5 users, from v1.76.0.
The Agentic AI Foundation announced on 9 September 2026 that “Envoy AI Gateway is now Agent Router”, adding: “the license remains Apache 2.0”. It runs on Envoy Gateway, whose SecurityPolicy checks each request for a valid JWT against a remote JWKS. Kong’s AI Proxy plugin “accepts requests in one of a few defined and standardized OpenAI formats” and translates them for providers that include vLLM and Ollama. Kong’s token-based rate limiting and its PII sanitiser are each “only available as part of our AI Gateway Enterprise offering”.
The Gateway API Inference Extension calls itself “an official Kubernetes project that optimizes self-hosting Generative Models on Kubernetes”. Its endpoint picker tells the gateway which replica of a model should serve each request. Keys, budgets and the log come from the gateway that implements it, so the two layers can be combined: an access gateway for keys and budgets, and replica routing behind it.
API keys per team and sign-in through the identity provider
LiteLLM creates virtual keys through its /key/generate endpoint. Each key carries a team_id or user_id, a list of allowed models and its own limits, and spend is tracked for the team or user attached to it. The model list enforces which team uses which model, for example the coding model on developer keys only. Service keys belong in each application’s secret store, with a rotation date; personal keys end with the person’s access to the platform.
With JWT auth, LiteLLM validates tokens from the identity provider against its public keys and reads team, user and role from claims named in its configuration, so access ends once the directory account is disabled and its last token has expired. Agent Router can use Envoy Gateway’s SecurityPolicy for the same check, and its rate limit example identifies the user from a request header, which the gateway should set from the validated token, not accept from the client.
Rate limits and token budgets per user, team and department
Rate limits protect interactive users from bulk jobs. LiteLLM sets tpm_limit and rpm_limit, tokens and requests per minute, on keys and teams, and limits set on a team apply across all of its keys. In the worked example, the nightly extraction key gets a token limit that leaves the chat model room for its peak if the job runs into working hours.
Budgets in LiteLLM are counted as spend. A key or team has a max_budget and a budget_duration such as 30d, and the documentation says the budget “is reset at the end of specified duration”. Once a key crosses its budget, requests fail, and max_budget_in_team caps one member within a team. Spend comes from a cost per token, so a private model needs an internal accounting rate in input_cost_per_token and output_cost_per_token, which LiteLLM uses only to track cost. With both costs set explicitly to 0, LiteLLM skips all budget checks for that model, so it stays available after a team has used up its budget, but its requests add nothing to spend and its use per department has to be counted in tokens. Decide per model which behaviour you want. The per-department totals then feed the showback or chargeback described in our guide to a shared GPU platform for departments.
Agent Router counts tokens per route in llmRequestCosts, with the types InputToken, CachedInputToken, OutputToken, TotalToken and a CEL expression, and enforces limits through a BackendTrafficPolicy on Envoy Gateway’s global rate limit, backed by Redis. Its documentation states that usage is charged after a response completes and that an admitted stream is not interrupted, so the last answer can overshoot a limit. Kong’s AI Rate Limiting Advanced counts total, prompt or completion tokens, or cost.
Routing to private models, fallbacks and external APIs
LiteLLM tries fallbacks in the listed order when a model fails, with separate lists for requests that exceed a model’s context window and for content policy errors. Retries and timeouts are set per model, and a deployment that fails more often per minute than allowed_fails is taken out of rotation for its cooldown_time. A fallback list that names an external model sends prompts out of the company whenever the private model is down, without anyone having decided it for that data. Keep fallbacks among private models, such as a second replica or a model with a longer context. Add an external API as a separate model name that only the keys of approved tasks may call, so every request to it carries a key, a team and a log entry.
In our Private AI/ML service, public APIs are used only when you enable hybrid mode for a specific task, and what goes there is visible in the query log. Tell us which applications would call the gateway and which tasks might need an external model.
Query log, PII masking and retention
LiteLLM writes a row for each request to its spend log in the database, and disable_spend_logs turns this off. Whether a row also holds the request and response content is set by store_prompts_in_spend_logs or the matching switch in the admin UI, and the documentation states that settings changed in the UI override the config file, so check both before go-live. Retention is set with maximum_, and the documentation states “Retention and cleanup are open source.” The setting turn_off_message_logging keeps messages and responses away from its logging callbacks, while “request metadata will still be logged”. A client can skip logging per request with no-log=True unless the administrator sets global_disable_no_log_param, which a mandatory query log requires.
LiteLLM masks personal data through Presidio, which needs its analysis and anonymisation services running beside the gateway. In pre_call mode it checks the input before the model call, and in post_call mode it checks input and output afterwards. Per entity type, MASK replaces a value with a placeholder such as <PERSON>, and BLOCK rejects the request. Kong’s AI Sanitizer uses an anonymiser service that can run as a container. Masking matters most on routes to external APIs; on a private route it can remove names a contract review needs. It does not detect injected instructions, which our guide to prompt injection and LLM security covers.
Article 5(1)(e) of the GDPR requires personal data to be kept “in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed”, and the log can hold prompts with names and client details. What the log stores, who may read it and for how long are a legal assessment for your legal department.
Policy decisions before the gateway goes live
| DECISION | EXAMPLE, 1,500 STAFF | WHO DECIDES |
|---|---|---|
| Models per team | chat model for all; coding model for developers | CIO, department heads |
| External APIs | off; enabled per task, separate model name | CIO, DPO, legal |
| Fallbacks | private models only | platform team |
| Limits per key | batch below the chat peak; budgets per department | platform team, finance |
| Log contents | metadata for all; text where policy requires it | DPO, security |
| Retention | set before the first entry | legal department |
| Readers of the log | named roles, their access logged | security, DPO |
| Logging opt-out | disabled for all clients | IT security |
An example policy; the decisions are yours. Settings as named in the LiteLLM documentation, read on 10 October 2026.
The same report that shows tokens per department shows whether staff use the approved platform or keep using public tools, which our guide to shadow AI policy and controls addresses. Bring the gateway live in this order.
- Put the gateway in front of every model server and close all other routes to them.
- Create teams and keys per application and environment, with allowed models and limits.
- Connect token validation to the identity provider and remove shared personal keys.
- Set the log contents, the readers and the retention period, and disable the logging opt-out.
- Add external APIs only as separate model names on the keys of approved tasks.
Logging of queries and answers, with data and permissions management, is part of the platform we build. Describe your departments, applications and the log your legal department expects in the form below.
What we do
Our Private AI/ML service builds an AI platform under your control, on your servers or in a Tier-3 data centre in Lithuania, where nothing goes to public services unless you explicitly enable it. A query log, data and permissions management and logging of queries and answers let security and legal see who accesses what, and how. Eurokommerz holds the contract and supplies the hardware, with engineering by our partner Vixen.UNO and support under an agreed SLA; the first call is free of charge, and the price of the technical assessment is fixed before work begins. Data handling during a project is set out on our security and compliance page.
FAQ
What is an LLM gateway?
Is there an open source AI gateway?
What is the LiteLLM proxy used for?
How do you rate limit LLM usage per user or team?
How do you track LLM cost per team for private models?
Should an LLM gateway log prompts and answers?
Send us the applications that will call your models, the departments and users behind them, and the external APIs you are considering. We reply within one business day with the next steps towards a platform with a query log and data and permissions management. The first call is free of charge.
Talk to an expertWe reply within one business day