Contact centre AI on-premise: GPUs for call transcription and call summaries for 200 to 1,000 agents
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Contact centre AI on-premise runs in two GPU stages, speech-to-text for every call minute and one LLM pass per call, both sized from the concurrent calls in the busy hour and the call minutes per day
- Recording agent and customer on separate channels removes the need for speaker diarisation but doubles the audio streams: 1,000 agents on calls at the same time are 2,000 streams
- NVIDIA states a maximum of 190 live English streams per L40S with Parakeet 0.6B CTC at 160 ms chunks and 900 at 960 ms, so live agent assist for 1,000 agents needs 4 to 16 L40S for speech alone by our estimate
- Summaries are the small stage: in MLPerf Inference v6.0, eight RTX PRO 6000 Server Edition produced 48,613.8 tokens per second on Llama 3.1 8B offline, and by our estimate one card covers the summaries of 1,000 agents
- The AI Act prohibits inferring emotions of workers since 2 February 2025, so agents’ voices stay out of any emotion model; recordings and transcripts need retention periods under the GDPR
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Contact centre AI on-premise: two GPU stages
Contact centre AI on-premise runs in two GPU stages. Speech-to-text turns every call minute into a transcript, and a language model reads each transcript once to write the summary, the reason for the call and the next step. Both stages are sized from the same two numbers: the calls in progress at the same time in the busy hour, and the call minutes per day.
In our estimate below, summaries after the call for 200 English-speaking agents fit on two data-centre cards, while live agent assist for 1,000 agents needs 4 to 16 L40S for speech alone, plus two RTX PRO 6000 Server Edition. Speech models, their VRAM and RTFx are covered in our Whisper and Parakeet sizing guide.
Speech-to-text and LLM: models and GPUs per stage
Live transcription needs a streaming model such as Parakeet CTC or Parakeet RNNT 1.1B multilingual, since NVIDIA’s ASR NIM runs Whisper offline only. After the call, offline models such as Parakeet TDT 0.6B v3, with 25 European languages, can be used.
Record agent and customer on separate channels where the telephone platform allows it. The speaker label then comes from the channel, and no diarisation model is needed, which on a mixed recording can take more GPU time than the transcription. The trade-off is that every call becomes two audio streams.
The LLM pass reads a long transcript and writes a short text. MLPerf’s Llama 3.1 8B benchmark has the same shape, summarising CNN/DailyMail news articles “with an average input length of 778 tokens and output length of 73 tokens”, as MLCommons describes it. We use its results as the reference for an 8B-class summary model; a larger model lowers the rate, as the 70B figures in our guide to batch LLM inference show.
| STAGE | MODEL CLASS | PUBLISHED FIGURE | CARD WE SUPPLY |
|---|---|---|---|
| Live speech-to-text | streaming, Parakeet 0.6B CTC | up to 190 English streams at 160 ms, 900 at 960 ms, one L40S | L40S |
| After the call, English | offline, Parakeet 0.6B CTC | RTFx 3,638 at 32 parallel requests, one L40S | L40S |
| After the call, 25 languages | offline, Parakeet TDT 0.6B v3 | 238.9 times audio duration at batch 8 on English audio, one L4 | L4 |
| Call summary | LLM, 8B class, batch | 48,613.8 tokens per second, eight cards, offline | RTX PRO 6000 Server Edition |
| Live agent assist | LLM, 8B class, latency-bound | 47,581.0 tokens per second, eight cards, server scenario | RTX PRO 6000 Server Edition |
NVIDIA ASR NIM performance page, last updated 7 October 2026 (L40S, English, maximum effective streams as NVIDIA states them); E2E Networks blog of 27 March 2026 (100 English LibriSpeech utterances on one L4); MLPerf® Inference v6.0, data-centre suite, closed division, available systems, Llama 3.1 8B, entry 6.0-0005, eight RTX PRO 6000 Blackwell Server Edition, FP4 weights, offline and server scenarios, retrieved from MLCommons’ v6.0 results repository on 10 October 2026, results verified by MLCommons Association. The server scenario holds time to first token at 2 s or less and time per output token at 100 ms or less, per MLCommons’ benchmark description. The MLPerf name and logo are registered and unregistered trademarks of MLCommons Association in the United States and other countries. All rights reserved. Unauthorized use strictly prohibited. See www.mlcommons.org for more information.
Real-time agent assist or after-call processing
After-call processing transcribes and summarises the recording when the call ends, and the jobs can queue. The hardware keeps pace with the busy hour if the summary must be ready for the next call, or with the day’s total if it may arrive the next morning.
Real-time agent assist transcribes both channels during the call and asks the LLM for a suggested answer each time the customer finishes speaking, so it is sized for the peak. In NVIDIA’s L40S figures, 512 English streams with 960 ms chunks ran at 198.8 ms average latency, and text arrives in steps of about one chunk. With 160 ms chunks, 64 streams ran at 33.9 ms, and NVIDIA states a maximum of 190 streams instead of 900, so short chunks need almost five times the cards.
Each suggestion request sends the conversation so far. vLLM’s automatic prefix caching “caches the KV cache of existing queries, so that a new query can directly reuse the KV cache if it shares the same prefix”, and its documentation names multi-round conversation as a use case. With it enabled, each request computes mainly the new turn instead of the whole call. Set targets for time to first token and time per output token as our guide to LLM latency targets describes, and test them at busy-hour concurrency.
The L40S streaming results we could read on NVIDIA’s page cover English, Spanish, Vietnamese and Mandarin, and we found no stream figure there for German, Polish, Czech or Romanian. For those languages a test with your own recordings sets the count, and the streaming profile’s memory must fit the card.
Sizing from busy-hour calls and minutes per day
- Take the calls in progress at the same time in the busy hour, or, as an upper limit, the number of agents logged in.
- Multiply by two for separate channels; the result is the live streams and the hours of audio per hour at the peak.
- For live assist, divide the streams by the streams per card you plan for; we plan at about 70 per cent of NVIDIA’s stated maximum.
- For after-call processing, the speech stage needs an RTFx at least equal to the channels recorded in the busy hour if summaries must keep pace, or the day’s channel-hours divided by the hours of the processing window if they may wait.
- For the LLM, multiply the calls finished per hour by the output tokens per summary, add the tokens of live suggestions, and divide by the tokens per second you plan per card.
As an example of step 4, if 1,000 agents talk 5 hours a day on two channels, they produce 10,000 channel-hours a day, which needs an RTFx of about 833 in a 12-hour night window and 2,000 to keep pace with a busy hour in which every agent is on a call.
GPU estimate for 200, 500 and 1,000 agents
| AGENTS, BUSY HOUR | LIVE STREAMS | AFTER-CALL SPEECH | LIVE SPEECH, L40S | LLM, RTX PRO 6000 |
|---|---|---|---|---|
| 200 agents, 200 calls | 400 | 1 L40S (English) or 4 L4 | 1 at 960 ms, 4 at 160 ms | 1, with or without assist |
| 500 agents, 500 calls | 1,000 | 1 L40S (English) or 9 L4 | 2 at 960 ms, 8 at 160 ms | 1, with or without assist |
| 1,000 agents, 1,000 calls | 2,000 | 2 L40S (English) or 17 L4 | 4 at 960 ms, 16 at 160 ms | 1 for summaries, 2 with assist |
Our estimates, not measurements, from the figures in the first table: live streams at 70 per cent of NVIDIA’s stated maxima (133 and 630 per L40S, English), after-call speech at half of the published rates (RTFx 1,819 per L40S for English, 119 per L4 for the 25-language model, from an English test, keeping pace with the busy hour), LLM at half of the per-card MLPerf rates (3,038 tokens per second offline, 2,974 in the server scenario). With live speech, summaries use the live transcript and the after-call column is not needed. A tighter latency target for live suggestions than MLPerf’s lowers the rate per card.
The table assumes that every agent is on a call in the busy hour and that calls last 6 minutes on average. It also assumes summaries of 150 output tokens and, with live assist, a suggestion of 80 tokens every 20 seconds per call.
For 1,000 agents, 10,000 calls end in the busy hour, and their summaries come to about 417 output tokens per second, about 14 per cent of what we plan per RTX PRO 6000. Live suggestions add about 4,000 tokens per second, which is why assist doubles the LLM cards. For the 25-language path the only published Parakeet TDT v3 rate we found is the English test on the L4, and none for the L40S or the RTX PRO 6000, so a test run with your own recordings sets the card. If summaries may wait for the night, the speech stage for 1,000 agents fits on 7 L4 by the same arithmetic.
We build AI servers to order with L4, L40S and RTX PRO 6000 Server Edition cards and check the rack, power and airflow before we quote. Send us your agent count, busy-hour calls and languages through the form below.
Recordings and transcripts: storage and retention under the GDPR
ITU-T G.711, the PCM codec for voice, sets “8000 samples per second” with eight bits per sample, so one channel takes 64 kbit/s, or 28.8 MB per hour by our arithmetic. At that rate the 10,000 channel-hours of 1,000 agents come to 288 GB of audio a day. Transcripts and summaries take a small fraction of that.
Recorded calls in which a caller or an agent can be identified are personal data. The GDPR requires that they are “limited to what is necessary” for the purposes (Article 5(1)(c)) and kept in identifiable form “for no longer than is necessary” (Article 5(1)(e)). Audio, transcript and summary can carry separate retention periods, and the deletion job removes the transcript from the search index as well as the file store. Calls can contain data concerning health, a special category under Article 9(1), so give access by role and log searches across transcripts.
Whether calls may be recorded, transcribed and summarised, and for how long each part is kept, is a legal assessment for the company’s legal department. Our guide to a data protection impact assessment for an internal LLM covers the documentation for a system of this kind.
Emotion recognition and the AI Act in the contact centre
Article 5(1)(f) of the AI Act prohibits “the use of AI systems to infer emotions of a natural person in the areas of workplace and education institutions”, except for medical or safety reasons, and the prohibitions have applied since 2 February 2025. Recital 18 explains that an emotion recognition system infers emotions or intentions “on the basis of their biometric data”. The mere detection of “characteristics of a person’s voice, such as a raised voice or whispering” is not covered “unless they are used for identifying or inferring emotions”.
The Commission published guidelines on prohibited practices on 4 February 2025, and CMS summarised their examples on 10 February 2025. Among the prohibited examples, that summary lists “using webcams and voice recognition systems by a call centre to track their employee’s emotions”. Among the examples outside the prohibition, it lists voice recognition that tracks customers’ emotions and “AI systems inferring emotions from written text and not based on biometric data”.
For the platform, the agent’s channel goes to no model that scores emotion from the voice, and separate channels make that a routing rule. Emotion recognition of customers from their voice stays an emotion recognition system: Article 50(3) requires deployers to inform the persons exposed, and Annex III point 1(c) lists it as high-risk. Scoring agents from their transcripts can fall under Annex III point 4(b), which covers systems that “monitor and evaluate the performance and behaviour” of workers. Under the amended AI Act, the high-risk rules for Annex III uses apply from 2 December 2027, as our guide to the EU AI Act for companies using LLMs explains, and summaries, search and suggestions are not purposes listed in Annex III.
Running speech and LLM on one platform with two servers
Live agent assist is part of the agent’s working day, so plan for a server failure. Two servers that each run the speech models and the LLM let one carry on at half capacity while the other is repaired or updated. Recording does not depend on the GPU servers, and after-call jobs wait in a queue until the servers return.
The L40S and the RTX PRO 6000 Server Edition are passive cards that need a server’s airflow, and the 72 W L4 suits dense after-call nodes. NVIDIA’s NIM containers need an NVIDIA AI Enterprise licence for production use, while open-source engines such as vLLM run without one.
Our Private AI/ML service, with engineering by our partner Vixen.UNO, deploys open and commercial models on servers under your control. Tell us which calls you would summarise first and which languages your agents speak.
What we supply
We supply the NVIDIA L4 and L40S and the RTX PRO 6000 Server Edition as GPUs for servers you already run, or in AI servers built to order, assembled and burn-in tested, on one EU contract and invoice with manufacturer warranty. We check the rack, power and airflow before we quote. NVIDIA AI Enterprise licences for NIM come on the same invoice. Deploying the models on-premise is our Private AI/ML service, with engineering by our partner Vixen.UNO and support under an agreed SLA.
FAQ
Can contact centre AI run on-premise?
What server do I need for call transcription?
How much GPU does call summarisation with an LLM need?
Which GPU is right for call centre AI?
What does real-time agent assist on-premise need?
Does the AI Act ban emotion detection in call centres?
Send us the number of agents, the concurrent calls in the busy hour, the call minutes per day, the languages and whether you need live agent assist or summaries after the call. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day