Embedding and reranker servers for RAG: which GPU, and how many for millions of documents
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- Embedding models and rerankers for RAG are small: Qwen3-Embedding-0.6B takes about 1.2 GB in BF16 and the 8B model about 16 GB, so throughput, not memory, decides their GPU
- NVIDIA’s NIM documentation lists 49.2 texts of 512 tokens per second on an L4 and 144.3 on an L40S for its 1B embedding model in FP8 (version 1.14.0, 30 September 2026), and 52.1 and 171.1 reranked passages per second for a 1B reranker (17 August 2026)
- By our arithmetic from those figures, 3 million chunks of 512 tokens take about 17 hours to embed on one L4 and about 5.8 hours on one L40S; at 40 passages per query, an L4 reranks about 1.3 queries per second and an L40S about 4.3, or 4.5 with three requests in parallel
- Text Embeddings Inference 1.9 has images for Ada cards such as the L4 and L40S and an image marked experimental for Blackwell 12.0, the compute capability of every RTX PRO Blackwell card; NIM in production requires NVIDIA AI Enterprise
- Retrieval models can run on a small card of their own or share an RTX PRO 6000 as 24 GB MIG instances; the L4, L40S, RTX PRO 4000 and RTX PRO 2000 are not on NVIDIA’s MIG list
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
Embedding and reranker GPU requirements for RAG
The GPU requirements of an embedding model and a reranker are small next to those of the language model in a RAG system. Qwen3-Embedding-0.6B needs about 1.2 GB for its weights in BF16 and the 8B model about 16 GB, so one 24 GB card such as the NVIDIA L4 holds the small models with room to spare. What sizes the GPU is throughput: tokens per second when the corpus is embedded, and queries per second times passages per query when the reranker scores search results.
By our arithmetic from NVIDIA’s figures, 3 million chunks of 512 tokens take about 17 hours to embed on one L4 and about 5.8 hours on one L40S with a 1B embedding model. At 40 passages per query and three requests in parallel, one L4 reranks about 1.3 queries per second and one L40S about 4.5. A corpus of a few million chunks with a few queries per second can therefore run both models on one small card or a MIG instance next to the LLM. A separate retrieval server makes sense for tens of millions of chunks or a full re-index that must finish within a maintenance window.
Our guide to RAG on company data covers chunking, permissions and evaluation; this article covers the hardware.
Embedding and reranker models: parameters, dimensions, languages
The models below are documented for multilingual corpora as of October 2026. NVIDIA’s embedding model lists 26 evaluated languages, among them German, French, Polish, Czech and Hungarian.
| MODEL | PARAMETERS | VECTOR DIMENSIONS | MAX INPUT | LANGUAGES, LICENCE |
|---|---|---|---|---|
| Qwen3-Embedding-0. | 0.6B | 1,024, MRL | 32K tokens | 100+; Apache 2.0 |
| Qwen3-Embedding-4B | 4B | 2,560, MRL | 32K tokens | 100+; Apache 2.0 |
| Qwen3-Embedding-8B | 8B | 4,096, MRL | 32K tokens | 100+; Apache 2.0 |
| BAAI bge-m3 | 568M | 1,024 dense, plus sparse and multi-vector | 8,192 tokens | 100+; MIT |
| Embedding | 300M | 768, or 512, 256, 128 | 2,048 tokens | 100+; Gemma Terms of Use |
| llama-nemotron-embed-1b-v2 | 1B | 384 to 2,048 | 8,192 tokens | 26 evaluated; NVIDIA Open Model License |
| Qwen3-Reranker 0.6B, 4B, 8B | 0.6B, 4B, 8B | none, scores pairs | 32K tokens | 100+; Apache 2.0 |
| bge-reranker-v2-m3 | 0.6B | none, scores pairs | 512 in the card’s example | multilingual; Apache 2.0 |
Hugging Face model cards of Qwen, BAAI, Google and NVIDIA, read 10 October 2026; bge-m3’s 568M total parameters from NVIDIA’s embedding NIM support matrix, updated 3 October 2026. MRL: the vector can be shortened (Matryoshka representation learning).
In BF16 the weights take two bytes per parameter, about 1.2 GB for a 0.6B model, 8 GB for a 4B model and 16 GB for an 8B model. A serving engine reserves more for its batches: NVIDIA’s support matrix gives 3.6 GB of GPU memory for the non-optimised profile of llama-nemotron-embed-1b-v2 and 33 GB for bge-m3, whose weights take about 1.1 GB in FP16 by our arithmetic.
A million 1,024-dimensional float32 vectors take 4.10 GB before any index, as our vector database comparison shows, and 4,096 dimensions take four times as much. Qwen3 Embedding, EmbeddingGemma and NVIDIA’s model return shorter vectors on request; test one on your evaluation set before sizing the index.
Serving engines: Text Embeddings Inference, vLLM and NVIDIA NIM
Hugging Face’s Text Embeddings Inference (TEI) serves embedding models with Qwen3, Gemma3, XLM-RoBERTa, BERT and other architectures, and as re-rankers CamemBERT and XLM-RoBERTa sequence classification models. Each GPU architecture has its own image, version 1.9 as of October 2026: the 89- image for Ada cards such as the L4 and L40S, hopper- for Hopper, and 120- for Blackwell 12.0, marked experimental and listed with GeForce examples. NVIDIA’s CUDA GPU list gives compute capability 12.0 for every RTX PRO Blackwell card we supply, so on these cards TEI runs from that experimental image.
For sizing, --max-batch-tokens defaults to 16,384, and the documentation calls it “one critical control to allow maximum usage of the available hardware” that TEI cannot set automatically. --auto-truncate defaults to true, so a chunk longer than the model’s maximum input is cut rather than rejected, which matters for EmbeddingGemma at 2,048 tokens.
vLLM runs embedding and scoring models in its pooling runner, with /v1/embeddings, /score and /rerank endpoints. Qwen’s model cards require vLLM 0.8.5 or later for both Qwen3 series. The Qwen3 reranker is a language model that scores each query and passage by the probability of the tokens “yes” and “no”; it is not among TEI’s listed re-ranker architectures, so it runs on vLLM.
NVIDIA’s NeMo Retriever NIM containers expose an OpenAI-compatible API. For llama-nemotron-embed-1b-v2 and the rerank-1b-v2 and 500m-v2 models, the support matrices list optimised FP16 and FP8 profiles by compute capability 12.0, 10.0, 9.0 and 8.9, which covers the L4, the L40S and the RTX PRO Blackwell cards. NVIDIA’s NIM FAQ states: “Using NIM in production requires an NVIDIA AI Enterprise license.”
Published throughput on L4, L40S and RTX PRO 6000
The figures we found for embedding and reranking on the cards we supply come from NVIDIA’s NIM performance pages. The current embedding page, updated on 3 October 2026, uses “a separate synthetic AIPerf benchmark matrix” and no longer lists the L4, so the L4 and L40S rows for the 1B text model, which that page calls Llama Nemotron Embed 1B, come from the documentation of NIM version 1.14.0, updated on 30 September 2026. We found no comparable published figures for the Qwen3 or BGE models on these cards.
| MODEL, PRECISION | GPU | REQUEST | THROUGHPUT | AVG LATENCY |
|---|---|---|---|---|
| Embed 1B, FP8 | L4 | 64 texts of 512 tokens | 49.2 texts/s | 1,300 ms |
| Embed 1B, FP8 | L40S | 64 texts of 512 tokens | 144.3 texts/s | 443 ms |
| Embed VL 1B v2, FP8 | L40S | 64 texts of 512 tokens | 159.3 texts/s | 395 ms |
| Embed VL 1B v2, mixed | RTX PRO 6000 Server | 64 texts of 512 tokens | 156.6 texts/s | 402 ms |
| Nemotron 3 Embed 1B, NVFP4 | RTX PRO 6000 Server | 64 texts of 512 tokens | 511.4 texts/s | 118 ms |
| Rerank 1B, FP8 | L4 | 40 passages of 512 tokens | 52.1 passages/s | 768 ms |
| Rerank 1B, FP8 | L40S | 40 passages of 512 tokens | 171.1 passages/s | 234 ms |
| Rerank 500m, FP8 | L40S | 40 passages of 512 tokens | 297.0 passages/s | 134 ms |
| Rerank VL 1B v2, FP8 | RTX PRO 6000 Server | 40 passages of 514 tokens | 246.1 passages/s | 161 ms |
NVIDIA NeMo Retriever embedding NIM performance pages for version 1.14.0 (updated 30 September 2026, Embed 1B rows) and latest (updated 3 October 2026, other embedding rows), and reranking NIM performance page (updated 17 August 2026), read 10 October 2026; one request at a time.
For embedding at 512 tokens the L40S processes 2.9 times as many texts as the L4, and for reranking 3.3 times as many passages. The RTX PRO 6000 Server Edition at mixed precision is level with the L40S on the VL model. With Nemotron 3 Embed 1B in NVFP4 it processes 511.4 texts per second, about 3.5 times the L40S row for Embed 1B, though the two rows use different models. Embedding a search query is a light load, and the L4 table lists 13 ms for one query of 20 tokens.
How long indexing takes: documents, chunks and tokens
Indexing is a one-off peak, repeated in full whenever the embedding model changes, and it is estimated in four steps.
- Run your chunker on a sample and take the average number of chunks per document.
- Count the average tokens per chunk with the embedding model’s tokenizer.
- Divide the total chunks by the throughput of your model on the candidate card, from a published figure or a test with your own chunks.
- Add document parsing and OCR, which run separately and often take longer, and the writes to the vector database.
As in our vector database comparison, 300,000 documents at ten chunks each give 3 million chunks. At 512 tokens each, that is 1.54 billion tokens. By our estimate from the throughput above, the full run with Embed 1B in FP8 takes about 17 hours on one L4 and 5.8 hours on one L40S. One RTX PRO 6000 Server Edition takes about 5.3 hours with the VL model at mixed precision and about 1.6 hours with Nemotron 3 Embed 1B in NVFP4.
An 8B embedding model does several times the arithmetic per token of these 1B models, so plan its full run in days rather than hours and test it first. After the first run, only new and changed documents are embedded, a small load for any of these cards.
Sizing the reranker: queries per second and passages
A reranker scores every retrieved passage against the query, so its load is queries per second times passages per query, at the token length of query plus passage. In NVIDIA’s figures at 40 passages of 512 tokens, the 1B reranker handles 52.7 passages per second on an L4 at concurrency 3, about 1.3 queries per second. On an L40S it handles 181.6, about 4.5 queries per second, and the 500m model on the L40S 389, about 9.7.
Our article on a private ChatGPT server by company size uses example values that give 1,200 requests in the busiest hour for 500 employees and a peak factor of 2. With one retrieval per request, that is about 0.67 queries per second at the peak, which one L4 covers. For 2,000 employees the same values give about 2.7 queries per second, beyond the L4 and within the L40S. An agent that searches several times per answer multiplies the rate.
Each answer also waits for the reranker before the LLM starts, 768 ms for 40 passages on the L4 and 234 ms on the L40S in NVIDIA’s tables with one request at a time. With three requests in parallel, the setting behind the queries per second above, the tables list 2,254 ms and 652 ms. Fewer passages per query reduce both load and wait.
We supply the L4, the L40S and the RTX PRO cards for retrieval models, as cards for your servers or in AI servers built to order. Tell us your document count, chunk length and peak queries per second, and we size the cards from them.
A separate card or a MIG instance on the LLM GPU
The first of three placements gives the retrieval models a small card of their own in the LLM server: the passive, low-profile L4 at 72 W, the RTX PRO 4000 SFF at 70 W, the 16 GB RTX PRO 2000 at 70 W, or the 32 GB RTX PRO 4500. The second is MIG on the LLM card. NVIDIA’s MIG guide lists up to four 24 GB instances on every RTX PRO 6000 edition and up to seven on the H200 NVL, each with its own memory and fault isolation, as our guide to GPU sharing with MIG, time-slicing and MPS explains.
The third is running the engines side by side on one card without partitioning. Each engine then reserves its own share of memory and they compete for compute, so run a full re-index at night or on a separate card. NVIDIA publishes no throughput for a MIG instance or for the RTX PRO 2000, 4000 and 4500 with these models, so test your model on the configuration you plan.
Example configurations by corpus size and query load
| CORPUS AND LOAD | FULL INDEX, ESTIMATE | RERANKING CAPACITY | CONFIGURATION |
|---|---|---|---|
| 3M chunks, 1 query/s | 17 h on L4, 5.8 h on L40S | L4: 1.3 queries/s | one L4 in the LLM server, or a 24 GB MIG instance on an RTX PRO 6000 (test it) |
| 3M chunks, 4 queries/s | 5.8 h on L40S | L40S: 4.5 with 1B, 9.7 with 500m | one L40S for both models, the 500m reranker for headroom |
| 30M chunks, 4 queries/s | 58 h on one L40S, 16 h on an RTX PRO 6000 in NVFP4 | as above | two L40S in a retrieval server, or an RTX PRO 6000 Server Edition |
| Qwen3 8B pair, any corpus | several times the 1B time; test | no published figure | one 48 GB L40S, or two 24 GB MIG instances |
Our estimates: chunks of 512 tokens divided by the published texts per second above; reranking at 40 passages per query and three requests in parallel. Parsing, OCR and database writes excluded. The 8B pair’s 32 GB of BF16 weights as in our company-size article.
We check the rack, power and airflow before we quote, and return a configuration and quote within one business day. Send us the corpus size and the query load through the form below, rough estimates are enough.
What we supply
We supply the cards for retrieval models, the L4, L40S, RTX PRO 2000, 4000, 4000 SFF and 4500, and the RTX PRO 6000 and H200 NVL for MIG next to the LLM, as professional NVIDIA GPUs or in AI servers built to order. Servers are burn-in tested, with manufacturer warranty on every component, on one EU contract and invoice, and NVIDIA AI Enterprise licences for NIM come on the same invoice. The RAG platform itself, with assistants that cite the source and respect each user’s access rights, is our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How much GPU memory does an embedding model need?
How long does it take to embed a million documents?
Which GPU do I need for a reranker?
Can the embedding model run on the same GPU as the LLM?
Does Text Embeddings Inference run on RTX PRO Blackwell GPUs?
Which embedding model works for multilingual company documents?
Send us the number of documents, the average chunk count and length, the embedding and reranker models you are considering, the peak queries per second and the servers you run. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.
Talk to an expertWe reply within one business day