Whisper GPU requirements: sizing an on-premise speech-to-text server for Whisper and Parakeet
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Speech recognition needs little GPU memory: OpenAI gives about 10 GB of VRAM for Whisper large and about 6 GB for large-v3-turbo, Parakeet TDT 0.6B has 600 million parameters, and NVIDIA’s ASR NIM asks for at least 16 GB
- The GPU count follows from the peak number of live streams and from the hours of recorded audio per hour of processing, published as RTFx, seconds of audio transcribed per second of computing
- On one A100 in the Open ASR Leaderboard paper (December 2025), Whisper large-v3 reached an RTFx of 145.5, Canary-1B v2 749 and Parakeet TDT 0.6B v3 3,333; Parakeet v3 and Canary cover 25 European languages, Whisper 99
- NVIDIA’s ASR NIM figures, updated 7 October 2026, give one L40S a maximum of 190 live English streams with Parakeet 0.6B CTC at 160 ms chunks; NIM runs Whisper offline only
- With speaker diarisation the L40S offline rate fell from 3,638 to 101.5 in NVIDIA’s figures; by our estimate a 48 GB L40S or RTX PRO 5000 holds recognition, diarisation and a 20B LLM for summaries
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Whisper GPU requirements in short
The GPU requirements of Whisper and other speech-to-text models are modest, because speech recognition needs far less GPU memory than a language model. OpenAI gives about 10 GB of VRAM for Whisper large and about 6 GB for large-v3-turbo, and NVIDIA’s Parakeet TDT 0.6B v3 has 600 million parameters, so every card we supply, from the 16 GB RTX PRO 2000 upwards, holds one of them. For an on-premise speech-to-text server, what decides the GPU is the workload: how many live streams run at the same moment, and how many hours of recorded audio arrive per hour of processing time.
Model cards and benchmarks publish throughput as RTFx, the inverse of the real-time factor: seconds of audio transcribed per second of computing. An RTFx of 100 turns 100 hours of recordings into one hour of GPU time. The choice of model changes this figure by more than an order of magnitude, so choose the model first and the card second.
Speech recognition models as of October 2026
Whisper large-v3 has 1,550 million parameters and its model card lists 99 languages. Large-v3-turbo reduces the decoding layers “from 32 to 4”, which OpenAI describes as “way faster, at the expense of a minor quality degradation”, and it returns the original language even when translation is requested. OpenAI’s Whisper repository releases code and weights under the MIT licence, while the large-v3 card on Hugging Face is tagged Apache 2.0.
NVIDIA’s Parakeet TDT 0.6B v3, published on Hugging Face on 14 August 2025, transcribes 25 European languages, among them Polish, Czech, Slovak, Romanian, Hungarian and German, with punctuation and word timestamps. Its card states audio “up to 24 minutes long with full attention (on A100 80GB) or up to 3 hours with local attention” in one pass. Canary-1B v2 covers the same 25 languages and also translates between English and the other 24. Both are under CC BY 4.0. Nemotron Speech Streaming 0.6B, in its Hugging Face release of 13 March 2026, is an English streaming model with chunk sizes of 80, 160, 560 and 1,120 ms. Canary-Qwen 2.5B was trained on English only.
| MODEL | PARAMETERS | LANGUAGES | PUBLISHED MEMORY | RTFX ON A100 |
|---|---|---|---|---|
| Whisper large-v3 | 1,550 M | 99 | about 10 GB (OpenAI); 12.5 GB NIM profile | 145.5 |
| Whisper large-v3-turbo | 809 M | 99, no translation | about 6 GB (OpenAI) | not in the paper |
| Parakeet TDT 0.6B v3 | 600 M | 25 European | 14.02 GB NIM profile, batch 1,024 | 3,333 |
| Canary-1B v2 | 978 M | 25 European, translation | at least 6 GB RAM to load (card) | 749 |
| Canary-Qwen 2.5B | 2.5 B | English | not stated | 418 |
| Nemotron ASR Streaming | 600 M (English model) | English; 40 locales in NIM | 4 GB at batch 32, 23 GB at batch 512 (NIM) | not in the paper |
Model cards on Hugging Face and OpenAI’s Whisper repository, read on 9 October 2026; NIM profiles from NVIDIA’s ASR NIM support matrix, last updated 7 October 2026; RTFx from the Open ASR Leaderboard paper (arXiv 2510.06961v3, 10 December 2025), one A100-SXM4-80GB, batch size 64 where memory allowed.
GPU memory for Whisper large-v3 and turbo
OpenAI’s table in the Whisper repository gives about 10 GB of VRAM for the large models and about 6 GB for turbo, which it measured at about eight times the speed of large when transcribing English on an A100. Optimised runtimes can need less. faster-whisper, a reimplementation on CTranslate2, reports in its README 13 minutes of audio transcribed with large-v2 and beam size 5 on an 8 GB GeForce RTX 3070 Ti. The original implementation took 2 minutes 23 seconds at 4,708 MB. faster-whisper in FP16 with batch_size=8 took 17 seconds at 6,090 MB, and in INT8 with the same batch size 16 seconds at 4,500 MB.
NVIDIA’s ASR NIM container, which serves Parakeet, Canary, Whisper and Nemotron models, requires a GPU with compute capability 8.0 or higher and “at least 16 GB of VRAM”. Its profiles reserve memory for large batches: 12.5 GB for Whisper large-v3, 13.75 to 14.02 GB for Parakeet TDT, and up to 50.07 GB for the largest Parakeet RNNT multilingual profile, more than an L40S has. The support matrix adds: “Ensure that the models you deploy do not exceed the available GPU memory.” Memory therefore depends on batch size and on the number of models loaded more than on the model’s parameter count.
Throughput in RTFx: what the published figures show
The Open ASR Leaderboard paper, from Hugging Face, the University of Cambridge and Mistral AI, measured every model on the same A100 at batch size 64 where memory allowed. Parakeet TDT 0.6B v3 transcribed about 23 times as fast as Whisper large-v3 there. On the English test sets its average word error rate was 6.32 per cent, against 7.44 for Whisper large-v3.
For German, the paper’s multilingual table gives Whisper large-v3 4.97 per cent, Parakeet TDT 0.6B v3 4.90 and Canary-1B v2 4.96. Its multilingual table covers German, French, Italian, Spanish and Portuguese only, so for Polish, Czech or Romanian test the models on your own recordings before you size the server.
For the L4, E2E Networks reported on 27 March 2026 a test with 100 LibriSpeech utterances. Whisper large-v3-turbo in BF16 at batch 4 processed audio 41.1 times faster than its duration, in 2,299 MB. Parakeet TDT 0.6B v3 in BF16 at batch 8 did so 238.9 times faster, in 5,151 MB. The test set is small and its utterances are short, so we use these figures as indications.
Real-time transcription: concurrent streams per GPU
Whisper works on 30-second windows, and its card states that “a chunk length of 30-seconds is optimal” for large-v3. NVIDIA’s ASR NIM runs Whisper and Parakeet TDT offline only. For live calls it offers streaming models: Parakeet CTC for English and several other languages, Parakeet RNNT 1.1B multilingual and Nemotron ASR Streaming.
NVIDIA’s ASR NIM performance page, last updated on 7 October 2026, measures streams with the riva_streaming_asr_client and its --simulate_realtime flag on a LibriSpeech file. On the L40S, Parakeet 0.6B CTC in English with 160 ms chunks handled 64 streams at 33.9 ms average and 39.7 ms p90 latency, and NVIDIA states a maximum of 190 effective streams with an n-gram language model. With 960 ms chunks, 512 streams ran at 198.8 ms average latency, with a stated maximum of 900. In streaming, RTFx equals about the number of streams, because each stream arrives at speaking pace, so latency is the figure to watch.
We found no Whisper or Canary result for the L40S on that page, and no figures for the L4 or the RTX PRO cards. The support matrix lists the L4, the L40 and an entry “Blackwell RTX 60xx”, but not the L40S by name, although the performance page measures it. Check the matrix of the NIM release you deploy against the card you order.
Sizing a transcription server by workload
Size live calls by the peak number of concurrent streams. Size recordings by the RTFx they need: 2,000 hours of calls per month, processed in 20 working days of 8 hours, need a sustained RTFx of 12.5.
| WORKLOAD | EXAMPLE VOLUME | MODEL AND ENGINE | SETUP, ESTIMATE |
|---|---|---|---|
| Live calls, English | 100 concurrent calls | Parakeet 0.6B CTC in ASR NIM, 160 ms chunks | 1 L40S, about half of NVIDIA’s stated 190 streams |
| Live calls, English | 300 concurrent calls | Parakeet 0.6B CTC in ASR NIM | 3 L40S at 160 ms chunks, or 1 L40S at 960 ms chunks |
| Meeting recordings | 50 hours a day, speaker labels, summaries | Parakeet TDT v3 or Whisper turbo, diarisation, 20B LLM | 1 L40S or 1 RTX PRO 5000 48 GB |
| Batch archive, EU languages | 10,000 hours, one-off | Parakeet TDT 0.6B v3 | 1 L4, about 42 hours of processing |
| Batch archive, any language | 10,000 hours, one-off | Whisper large-v3-turbo | 4 L4, about 61 hours of processing |
Our estimates, not measurements: live calls from NVIDIA’s ASR NIM figures for the L40S (7 October 2026), planned to about 70 per cent of the stated maximum; archives from E2E Networks’ L4 figures (27 March 2026), 238.9 and 41.1 times the audio duration per card; meetings from NVIDIA’s L40S offline rate with diarisation, which NVIDIA measured with Parakeet 0.6B CTC in English, not with the models in that row.
Your audio, languages and settings change these numbers, so a test run with a few hours of your own recordings comes before the order. Telephone audio, overlapping speakers and background noise lower accuracy, and longer chunks raise latency.
We build transcription servers to order with L4, L40S or RTX PRO cards and check the rack, power and airflow before we quote. Tell us your hours of audio per day and peak concurrent calls in the form below.
Diarisation and LLM summaries on the same server
Speaker labels can cost more than the transcription. In NVIDIA’s L40S figures, Parakeet 0.6B CTC offline with 32 parallel requests reached an RTFx of 3,638 without diarisation and 101.5 with it. For meeting recordings, size the server from the diarised rate. The pyannote community-1 pipeline, under CC BY 4.0, is an open alternative. Its card says pyannote pipelines run on the CPU by default and can be sent to the GPU with a few lines of code, and it publishes no speed figure. Its files download only after you accept its conditions on Hugging Face.
Summaries and minutes need a language model next to the recogniser. Canary-Qwen 2.5B has an LLM mode that can “summarize it or answer questions about it”, working from the transcript, in English only. For other languages, run a separate LLM. By our estimate, gpt-oss-20b, at 13.8 GB, fits next to Whisper turbo and a diariser on one 48 GB L40S or RTX PRO 5000. gpt-oss-120b, at 65.3 GB, needs a 96 GB RTX PRO 6000 to run next to the recogniser. Our LLM hardware requirements by model cover the larger models.
We supply the cards for recognition, diarisation and summaries in one server. Describe the languages, the recordings and the summaries you need through the form below.
Which GPUs we supply fit a speech-to-text server
| CARD | MEMORY, POWER | COOLING, FORMAT | FIT, OUR READING |
|---|---|---|---|
| NVIDIA L4 | 24 GB, 72 W | passive, low profile, single slot | batch transcription in servers; listed in the ASR NIM matrix |
| RTX PRO 2000 | 16 GB, 70 W | active, 2.7 by 6.6 in, dual slot | one model, at the 16 GB NIM minimum, not named in its matrix |
| RTX PRO 4000 SFF | 24 GB, 70 W | active, low profile, dual slot | small sites without server airflow |
| RTX PRO 4000 | 24 GB, 145 W | active, single slot | workstations and compact servers |
| RTX PRO 4500 | 32 GB, 200 W | active, dual slot | recognition plus a small LLM |
| RTX PRO 5000 | 48 or 72 GB, 300 W | active, dual slot | recognition, diarisation and a 20B LLM |
| NVIDIA L40S | 48 GB, 350 W | passive, dual slot | live calls, with NVIDIA’s published stream figures |
| RTX PRO 6000 Server Edition | 96 GB, up to 600 W | passive, dual slot | recognition plus a 120B-class LLM |
Memory and power from our GPU page and NVIDIA product pages, read on 9 October 2026; fit is our reading of the figures in this article.
The passive cards need a server’s airflow, and the actively cooled RTX PRO cards suit workstations and edge systems. Our L4 vs L40S comparison and our guide to the 70 W cards cover formats and server counts.
faster-whisper is under the MIT licence and runs without an NVIDIA AI Enterprise licence. NVIDIA’s NIM FAQ, updated on 6 August 2026, states that production use of NIM requires an NVIDIA AI Enterprise licence, which is licensed per GPU.
Data protection and security of a transcription server
Recorded calls and meetings in which people can be identified are personal data under the GDPR, and whether and how they may be transcribed is an assessment for the company’s legal department. On the technical side, check that no component sends audio to a hosted service. The pyannote card, for example, offers its Precision-2 model through an API on pyannoteAI’s servers. Keep recordings and transcripts on the server under your control, restrict access to its API and harden its management controller, drivers and container runtime as in our guide to securing a GPU server. Camera analytics, a related workload, is covered in our GPU server for video analytics.
What we supply
We supply the NVIDIA L4 and L40S, the RTX PRO 2000, 4000, 4000 SFF, 4500 and 5000, and the RTX PRO 6000 Server Edition as GPUs for servers you already run, or in AI servers built to order, assembled and burn-in tested, on one EU contract and invoice with manufacturer warranty. Before you order, we confirm that the card fits the chassis and slot and that the rack can power and cool it. NVIDIA AI Enterprise licences for NIM come on the same invoice. Operating system, drivers, CUDA and a container runtime are installed on request, and speech models and LLMs on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
What are the GPU requirements for Whisper?
How much VRAM does Whisper large-v3 need?
Which GPU for real-time transcription?
Can speech to text run on-premise?
What is NVIDIA Parakeet?
How many GPUs does a transcription server need?
Send us the hours of audio per day, the peak number of concurrent calls, the languages and whether you need speaker labels and summaries. We reply within one business day with a configuration and a quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day