BLOG · GUIDE ·

LLM gateway for a company: one API, keys per team, token budgets and a query log

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • An LLM gateway gives every application one OpenAI-compatible endpoint, checks a key per application or team, enforces rate limits and token budgets, routes to private models or, where enabled, to an external API, and logs each request
  • vLLM ignores the user parameter of its Chat Completions API and has no user accounts, so only a gateway that is the only way to the model server knows which team asked what
  • LiteLLM’s proxy is MIT-licensed outside its enterprise directory; JWT sign-in through the identity provider, key rotation and per-model key budgets need its Enterprise licence
  • Envoy AI Gateway has been called Agent Router since 9 September 2026 and stays under Apache 2.0, with token rate limits that need Redis; Kong’s AI Proxy is in its Apache 2.0 gateway code, while token rate limiting and the PII sanitiser need AI Gateway Enterprise
  • Fallback lists should name private models only, the per-request opt-out from logging should be disabled, and the log’s contents and retention period are decided before the first request

Eurokommerz × Vixen.UNO: Private AI/ML  Talk to an expert →

What an LLM gateway does in a company platform

An LLM gateway, also called an AI gateway, is a proxy through which every application reaches the language models. It offers one OpenAI-compatible endpoint, checks the caller’s key, and routes each request to a private model or, only where the company has enabled it, to an external API. It enforces rate limits and token budgets per application and team and writes each request to a query log kept under a retention policy. On a platform with 500 to 2,000 users and several applications, it is also the one component that knows which team consumed how many tokens.

FUNCTIONWHY IT MATTERSWHERE DOCUMENTED
One OpenAI-style endpointapplications change models without code changesLiteLLM proxy, Kong AI Proxy, Agent Router
Keys per application, teamone key is revoked without affecting other usersLiteLLM virtual keys
Identity provider sign-inaccess follows the directory accountLiteLLM JWT auth, Envoy Gateway SecurityPolicy
Rate limitsa batch job cannot crowd out interactive chatLiteLLM tpm and rpm limits, Agent Router token limits
Token budgetsconsumption per department and periodLiteLLM max_budget and budget_duration
Fallback routingrequests move on when a model failsLiteLLM fallbacks
PII maskingnames and IDs masked, or the request blockedLiteLLM with Presidio, Kong AI Sanitizer
Query log, retentionwho asked what, kept for a set periodLiteLLM spend logs and retention cleanup

LiteLLM, Agent Router, Envoy Gateway and Kong documentation, read on 10 October 2026.

Why the model server needs a gateway in front

vLLM checks static API keys but has no user accounts, and its documentation notes that the user parameter of the Chat Completions API is ignored. Several applications that share one vLLM key look identical in its logs, and revoking the key stops all of them. The gateway issues its own keys, knows which application and team each one belongs to and passes the request on with a key that only the gateway holds. This works only if the gateway is the only way in. On Kubernetes, network policies let only the gateway, the endpoint picker and Prometheus reach the serving port, as our guide to a private LLM platform on Kubernetes sets out.

A worked example shows the scale; the figures are an illustration, not a sizing rule. A company of 1,500 staff runs a chat front end for all employees, a RAG assistant for legal and sales, a coding assistant for 60 developers, workflows in n8n and a nightly batch that extracts fields from incoming documents. With one key per application in test and in production, the gateway holds ten service keys, plus personal keys for developers whose IDE plugins call the API directly. The chat front end, whose sign-in our guide to a private ChatGPT alternative covers, has to pass the signed-in user’s identity with each request, or the log records only the application.

LLM gateway options: LiteLLM, Agent Router (Envoy AI Gateway) and Kong

GATEWAYLICENCESTATED LIMITS
LiteLLM proxyMIT, except its enterprise directoryJWT auth, key rotation and per-model key budgets need an Enterprise licence
Agent RouterApache 2.0formerly Envoy AI Gateway; token rate limits need Redis
Kong AI GatewayAI Proxy: Apache 2.0 codeAI Rate Limiting Advanced and AI Sanitizer are AI Gateway Enterprise only
Inference ExtensionApache 2.0its endpoint picker chooses a replica; keys and budgets come from the gateway

LiteLLM, Kong Gateway and Inference Extension LICENSE files, LiteLLM and Agent Router documentation, the Agentic AI Foundation’s post of 9 September 2026, Kong plugin pages and the Gateway API Inference Extension site, read on 10 October 2026.

LiteLLM’s proxy keeps keys, teams and spend in a Postgres database, which its documentation requires for virtual keys. Its licence file says that content outside the enterprise directory “is available under the MIT license”, while that directory has a licence file of its own. The documentation marks several features as requiring “a LiteLLM Enterprise license”, among them JWT-based sign-in, key rotation, budgets per model on a key and logging to some destinations, such as cloud storage buckets and custom callback APIs. Single sign-on for its admin interface works without that licence for up to 5 users, from v1.76.0.

The Agentic AI Foundation announced on 9 September 2026 that “Envoy AI Gateway is now Agent Router”, adding: “the license remains Apache 2.0”. It runs on Envoy Gateway, whose SecurityPolicy checks each request for a valid JWT against a remote JWKS. Kong’s AI Proxy plugin “accepts requests in one of a few defined and standardized OpenAI formats” and translates them for providers that include vLLM and Ollama. Kong’s token-based rate limiting and its PII sanitiser are each “only available as part of our AI Gateway Enterprise offering”.

The Gateway API Inference Extension calls itself “an official Kubernetes project that optimizes self-hosting Generative Models on Kubernetes”. Its endpoint picker tells the gateway which replica of a model should serve each request. Keys, budgets and the log come from the gateway that implements it, so the two layers can be combined: an access gateway for keys and budgets, and replica routing behind it.

API keys per team and sign-in through the identity provider

LiteLLM creates virtual keys through its /key/generate endpoint. Each key carries a team_id or user_id, a list of allowed models and its own limits, and spend is tracked for the team or user attached to it. The model list enforces which team uses which model, for example the coding model on developer keys only. Service keys belong in each application’s secret store, with a rotation date; personal keys end with the person’s access to the platform.

With JWT auth, LiteLLM validates tokens from the identity provider against its public keys and reads team, user and role from claims named in its configuration, so access ends once the directory account is disabled and its last token has expired. Agent Router can use Envoy Gateway’s SecurityPolicy for the same check, and its rate limit example identifies the user from a request header, which the gateway should set from the validated token, not accept from the client.

Rate limits and token budgets per user, team and department

Rate limits protect interactive users from bulk jobs. LiteLLM sets tpm_limit and rpm_limit, tokens and requests per minute, on keys and teams, and limits set on a team apply across all of its keys. In the worked example, the nightly extraction key gets a token limit that leaves the chat model room for its peak if the job runs into working hours.

Budgets in LiteLLM are counted as spend. A key or team has a max_budget and a budget_duration such as 30d, and the documentation says the budget “is reset at the end of specified duration”. Once a key crosses its budget, requests fail, and max_budget_in_team caps one member within a team. Spend comes from a cost per token, so a private model needs an internal accounting rate in input_cost_per_token and output_cost_per_token, which LiteLLM uses only to track cost. With both costs set explicitly to 0, LiteLLM skips all budget checks for that model, so it stays available after a team has used up its budget, but its requests add nothing to spend and its use per department has to be counted in tokens. Decide per model which behaviour you want. The per-department totals then feed the showback or chargeback described in our guide to a shared GPU platform for departments.

Agent Router counts tokens per route in llmRequestCosts, with the types InputToken, CachedInputToken, OutputToken, TotalToken and a CEL expression, and enforces limits through a BackendTrafficPolicy on Envoy Gateway’s global rate limit, backed by Redis. Its documentation states that usage is charged after a response completes and that an admitted stream is not interrupted, so the last answer can overshoot a limit. Kong’s AI Rate Limiting Advanced counts total, prompt or completion tokens, or cost.

Routing to private models, fallbacks and external APIs

LiteLLM tries fallbacks in the listed order when a model fails, with separate lists for requests that exceed a model’s context window and for content policy errors. Retries and timeouts are set per model, and a deployment that fails more often per minute than allowed_fails is taken out of rotation for its cooldown_time. A fallback list that names an external model sends prompts out of the company whenever the private model is down, without anyone having decided it for that data. Keep fallbacks among private models, such as a second replica or a model with a longer context. Add an external API as a separate model name that only the keys of approved tasks may call, so every request to it carries a key, a team and a log entry.

In our Private AI/ML service, public APIs are used only when you enable hybrid mode for a specific task, and what goes there is visible in the query log. Tell us which applications would call the gateway and which tasks might need an external model.

Query log, PII masking and retention

LiteLLM writes a row for each request to its spend log in the database, and disable_spend_logs turns this off. Whether a row also holds the request and response content is set by store_prompts_in_spend_logs or the matching switch in the admin UI, and the documentation states that settings changed in the UI override the config file, so check both before go-live. Retention is set with maximum_spend_logs_retention_period, and the documentation states “Retention and cleanup are open source.” The setting turn_off_message_logging keeps messages and responses away from its logging callbacks, while “request metadata will still be logged”. A client can skip logging per request with no-log=True unless the administrator sets global_disable_no_log_param, which a mandatory query log requires.

LiteLLM masks personal data through Presidio, which needs its analysis and anonymisation services running beside the gateway. In pre_call mode it checks the input before the model call, and in post_call mode it checks input and output afterwards. Per entity type, MASK replaces a value with a placeholder such as <PERSON>, and BLOCK rejects the request. Kong’s AI Sanitizer uses an anonymiser service that can run as a container. Masking matters most on routes to external APIs; on a private route it can remove names a contract review needs. It does not detect injected instructions, which our guide to prompt injection and LLM security covers.

Article 5(1)(e) of the GDPR requires personal data to be kept “in a form which permits identification of data subjects for no longer than is necessary for the purposes for which the personal data are processed”, and the log can hold prompts with names and client details. What the log stores, who may read it and for how long are a legal assessment for your legal department.

Policy decisions before the gateway goes live

DECISIONEXAMPLE, 1,500 STAFFWHO DECIDES
Models per teamchat model for all; coding model for developersCIO, department heads
External APIsoff; enabled per task, separate model nameCIO, DPO, legal
Fallbacksprivate models onlyplatform team
Limits per keybatch below the chat peak; budgets per departmentplatform team, finance
Log contentsmetadata for all; text where policy requires itDPO, security
Retentionset before the first entrylegal department
Readers of the lognamed roles, their access loggedsecurity, DPO
Logging opt-outdisabled for all clientsIT security

An example policy; the decisions are yours. Settings as named in the LiteLLM documentation, read on 10 October 2026.

The same report that shows tokens per department shows whether staff use the approved platform or keep using public tools, which our guide to shadow AI policy and controls addresses. Bring the gateway live in this order.

  1. Put the gateway in front of every model server and close all other routes to them.
  2. Create teams and keys per application and environment, with allowed models and limits.
  3. Connect token validation to the identity provider and remove shared personal keys.
  4. Set the log contents, the readers and the retention period, and disable the logging opt-out.
  5. Add external APIs only as separate model names on the keys of approved tasks.

Logging of queries and answers, with data and permissions management, is part of the platform we build. Describe your departments, applications and the log your legal department expects in the form below.

What we do

Our Private AI/ML service builds an AI platform under your control, on your servers or in a Tier-3 data centre in Lithuania, where nothing goes to public services unless you explicitly enable it. A query log, data and permissions management and logging of queries and answers let security and legal see who accesses what, and how. Eurokommerz holds the contract and supplies the hardware, with engineering by our partner Vixen.UNO and support under an agreed SLA; the first call is free of charge, and the price of the technical assessment is fixed before work begins. Data handling during a project is set out on our security and compliance page.

FAQ

What is an LLM gateway?
An LLM gateway is a proxy between applications and language models that offers one OpenAI-compatible endpoint. It checks a key per application or team, enforces rate limits and token budgets, routes requests to private models or, where enabled, to external APIs, and writes each request to a query log. It works only if it is the only way to reach the model servers.
Is there an open source AI gateway?
LiteLLM’s proxy is under the MIT licence outside its enterprise directory, although JWT sign-in, key rotation and per-model key budgets need its Enterprise licence. Agent Router, called Envoy AI Gateway until September 2026, is under Apache 2.0 and runs on Envoy Gateway. Kong’s AI Proxy plugin is in the Apache 2.0 Kong Gateway code, while its token-based rate limiting and PII sanitiser are part of its AI Gateway Enterprise offering.
What is the LiteLLM proxy used for?
The LiteLLM proxy gives applications one OpenAI-compatible endpoint in front of private models and external APIs. It issues virtual keys tied to teams and users, enforces token and request limits and spend budgets, and handles fallbacks between models. It stores keys, teams and its spend log in a Postgres database.
How do you rate limit LLM usage per user or team?
Put the limits in the gateway, since the model server does not know the user. LiteLLM sets tokens and requests per minute on keys and teams, and Agent Router counts input, output or total tokens per route and enforces limits through Envoy Gateway with Redis. Agent Router charges usage after a response completes, so an admitted stream can overshoot a limit.
How do you track LLM cost per team for private models?
Give each team its own keys and let the gateway count tokens per key and team. LiteLLM computes spend from a cost per token, so a private model needs an internal accounting rate, and a model with both costs set explicitly to 0 bypasses its budget checks and adds nothing to spend. The totals per department can then feed showback or chargeback.
Should an LLM gateway log prompts and answers?
A gateway is the one place that knows which user and team sent a request, so the query log belongs there, with its contents, readers and retention period decided before go-live. In LiteLLM, disable the per-request no-log option where the log is mandatory, and set the retention period for spend logs. How long prompts with personal data may be kept is a legal assessment for your legal department.

Send us the applications that will call your models, the departments and users behind them, and the external APIs you are considering. We reply within one business day with the next steps towards a platform with a query log and data and permissions management. The first call is free of charge.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna