BLOG · GUIDE ·

Embedding and reranker servers for RAG: which GPU, and how many for millions of documents

Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software

IN BRIEF
  • Embedding models and rerankers for RAG are small: Qwen3-Embedding-0.6B takes about 1.2 GB in BF16 and the 8B model about 16 GB, so throughput, not memory, decides their GPU
  • NVIDIA’s NIM documentation lists 49.2 texts of 512 tokens per second on an L4 and 144.3 on an L40S for its 1B embedding model in FP8 (version 1.14.0, 30 September 2026), and 52.1 and 171.1 reranked passages per second for a 1B reranker (17 August 2026)
  • By our arithmetic from those figures, 3 million chunks of 512 tokens take about 17 hours to embed on one L4 and about 5.8 hours on one L40S; at 40 passages per query, an L4 reranks about 1.3 queries per second and an L40S about 4.3, or 4.5 with three requests in parallel
  • Text Embeddings Inference 1.9 has images for Ada cards such as the L4 and L40S and an image marked experimental for Blackwell 12.0, the compute capability of every RTX PRO Blackwell card; NIM in production requires NVIDIA AI Enterprise
  • Retrieval models can run on a small card of their own or share an RTX PRO 6000 as 24 GB MIG instances; the L4, L40S, RTX PRO 4000 and RTX PRO 2000 are not on NVIDIA’s MIG list

Supplied by Eurokommerz: AI servers, built to order  Request a configuration →

Embedding and reranker GPU requirements for RAG

The GPU requirements of an embedding model and a reranker are small next to those of the language model in a RAG system. Qwen3-Embedding-0.6B needs about 1.2 GB for its weights in BF16 and the 8B model about 16 GB, so one 24 GB card such as the NVIDIA L4 holds the small models with room to spare. What sizes the GPU is throughput: tokens per second when the corpus is embedded, and queries per second times passages per query when the reranker scores search results.

By our arithmetic from NVIDIA’s figures, 3 million chunks of 512 tokens take about 17 hours to embed on one L4 and about 5.8 hours on one L40S with a 1B embedding model. At 40 passages per query and three requests in parallel, one L4 reranks about 1.3 queries per second and one L40S about 4.5. A corpus of a few million chunks with a few queries per second can therefore run both models on one small card or a MIG instance next to the LLM. A separate retrieval server makes sense for tens of millions of chunks or a full re-index that must finish within a maintenance window.

Our guide to RAG on company data covers chunking, permissions and evaluation; this article covers the hardware.

Embedding and reranker models: parameters, dimensions, languages

The models below are documented for multilingual corpora as of October 2026. NVIDIA’s embedding model lists 26 evaluated languages, among them German, French, Polish, Czech and Hungarian.

MODELPARAMETERSVECTOR DIMENSIONSMAX INPUTLANGUAGES, LICENCE
Qwen3-Embedding-0.6B0.6B1,024, MRL32K tokens100+; Apache 2.0
Qwen3-Embedding-4B4B2,560, MRL32K tokens100+; Apache 2.0
Qwen3-Embedding-8B8B4,096, MRL32K tokens100+; Apache 2.0
BAAI bge-m3568M1,024 dense, plus sparse and multi-vector8,192 tokens100+; MIT
EmbeddingGemma-300m300M768, or 512, 256, 1282,048 tokens100+; Gemma Terms of Use
llama-nemotron-embed-1b-v21B384 to 2,0488,192 tokens26 evaluated; NVIDIA Open Model License
Qwen3-Reranker 0.6B, 4B, 8B0.6B, 4B, 8Bnone, scores pairs32K tokens100+; Apache 2.0
bge-reranker-v2-m30.6Bnone, scores pairs512 in the card’s examplemultilingual; Apache 2.0

Hugging Face model cards of Qwen, BAAI, Google and NVIDIA, read 10 October 2026; bge-m3’s 568M total parameters from NVIDIA’s embedding NIM support matrix, updated 3 October 2026. MRL: the vector can be shortened (Matryoshka representation learning).

In BF16 the weights take two bytes per parameter, about 1.2 GB for a 0.6B model, 8 GB for a 4B model and 16 GB for an 8B model. A serving engine reserves more for its batches: NVIDIA’s support matrix gives 3.6 GB of GPU memory for the non-optimised profile of llama-nemotron-embed-1b-v2 and 33 GB for bge-m3, whose weights take about 1.1 GB in FP16 by our arithmetic.

A million 1,024-dimensional float32 vectors take 4.10 GB before any index, as our vector database comparison shows, and 4,096 dimensions take four times as much. Qwen3 Embedding, EmbeddingGemma and NVIDIA’s model return shorter vectors on request; test one on your evaluation set before sizing the index.

Serving engines: Text Embeddings Inference, vLLM and NVIDIA NIM

Hugging Face’s Text Embeddings Inference (TEI) serves embedding models with Qwen3, Gemma3, XLM-RoBERTa, BERT and other architectures, and as re-rankers CamemBERT and XLM-RoBERTa sequence classification models. Each GPU architecture has its own image, version 1.9 as of October 2026: the 89- image for Ada cards such as the L4 and L40S, hopper- for Hopper, and 120- for Blackwell 12.0, marked experimental and listed with GeForce examples. NVIDIA’s CUDA GPU list gives compute capability 12.0 for every RTX PRO Blackwell card we supply, so on these cards TEI runs from that experimental image.

For sizing, --max-batch-tokens defaults to 16,384, and the documentation calls it “one critical control to allow maximum usage of the available hardware” that TEI cannot set automatically. --auto-truncate defaults to true, so a chunk longer than the model’s maximum input is cut rather than rejected, which matters for EmbeddingGemma at 2,048 tokens.

vLLM runs embedding and scoring models in its pooling runner, with /v1/embeddings, /score and /rerank endpoints. Qwen’s model cards require vLLM 0.8.5 or later for both Qwen3 series. The Qwen3 reranker is a language model that scores each query and passage by the probability of the tokens “yes” and “no”; it is not among TEI’s listed re-ranker architectures, so it runs on vLLM.

NVIDIA’s NeMo Retriever NIM containers expose an OpenAI-compatible API. For llama-nemotron-embed-1b-v2 and the rerank-1b-v2 and 500m-v2 models, the support matrices list optimised FP16 and FP8 profiles by compute capability 12.0, 10.0, 9.0 and 8.9, which covers the L4, the L40S and the RTX PRO Blackwell cards. NVIDIA’s NIM FAQ states: “Using NIM in production requires an NVIDIA AI Enterprise license.”

Published throughput on L4, L40S and RTX PRO 6000

The figures we found for embedding and reranking on the cards we supply come from NVIDIA’s NIM performance pages. The current embedding page, updated on 3 October 2026, uses “a separate synthetic AIPerf benchmark matrix” and no longer lists the L4, so the L4 and L40S rows for the 1B text model, which that page calls Llama Nemotron Embed 1B, come from the documentation of NIM version 1.14.0, updated on 30 September 2026. We found no comparable published figures for the Qwen3 or BGE models on these cards.

MODEL, PRECISIONGPUREQUESTTHROUGHPUTAVG LATENCY
Embed 1B, FP8L464 texts of 512 tokens49.2 texts/s1,300 ms
Embed 1B, FP8L40S64 texts of 512 tokens144.3 texts/s443 ms
Embed VL 1B v2, FP8L40S64 texts of 512 tokens159.3 texts/s395 ms
Embed VL 1B v2, mixedRTX PRO 6000 Server64 texts of 512 tokens156.6 texts/s402 ms
Nemotron 3 Embed 1B, NVFP4RTX PRO 6000 Server64 texts of 512 tokens511.4 texts/s118 ms
Rerank 1B, FP8L440 passages of 512 tokens52.1 passages/s768 ms
Rerank 1B, FP8L40S40 passages of 512 tokens171.1 passages/s234 ms
Rerank 500m, FP8L40S40 passages of 512 tokens297.0 passages/s134 ms
Rerank VL 1B v2, FP8RTX PRO 6000 Server40 passages of 514 tokens246.1 passages/s161 ms

NVIDIA NeMo Retriever embedding NIM performance pages for version 1.14.0 (updated 30 September 2026, Embed 1B rows) and latest (updated 3 October 2026, other embedding rows), and reranking NIM performance page (updated 17 August 2026), read 10 October 2026; one request at a time.

For embedding at 512 tokens the L40S processes 2.9 times as many texts as the L4, and for reranking 3.3 times as many passages. The RTX PRO 6000 Server Edition at mixed precision is level with the L40S on the VL model. With Nemotron 3 Embed 1B in NVFP4 it processes 511.4 texts per second, about 3.5 times the L40S row for Embed 1B, though the two rows use different models. Embedding a search query is a light load, and the L4 table lists 13 ms for one query of 20 tokens.

How long indexing takes: documents, chunks and tokens

Indexing is a one-off peak, repeated in full whenever the embedding model changes, and it is estimated in four steps.

  1. Run your chunker on a sample and take the average number of chunks per document.
  2. Count the average tokens per chunk with the embedding model’s tokenizer.
  3. Divide the total chunks by the throughput of your model on the candidate card, from a published figure or a test with your own chunks.
  4. Add document parsing and OCR, which run separately and often take longer, and the writes to the vector database.

As in our vector database comparison, 300,000 documents at ten chunks each give 3 million chunks. At 512 tokens each, that is 1.54 billion tokens. By our estimate from the throughput above, the full run with Embed 1B in FP8 takes about 17 hours on one L4 and 5.8 hours on one L40S. One RTX PRO 6000 Server Edition takes about 5.3 hours with the VL model at mixed precision and about 1.6 hours with Nemotron 3 Embed 1B in NVFP4.

An 8B embedding model does several times the arithmetic per token of these 1B models, so plan its full run in days rather than hours and test it first. After the first run, only new and changed documents are embedded, a small load for any of these cards.

Sizing the reranker: queries per second and passages

A reranker scores every retrieved passage against the query, so its load is queries per second times passages per query, at the token length of query plus passage. In NVIDIA’s figures at 40 passages of 512 tokens, the 1B reranker handles 52.7 passages per second on an L4 at concurrency 3, about 1.3 queries per second. On an L40S it handles 181.6, about 4.5 queries per second, and the 500m model on the L40S 389, about 9.7.

Our article on a private ChatGPT server by company size uses example values that give 1,200 requests in the busiest hour for 500 employees and a peak factor of 2. With one retrieval per request, that is about 0.67 queries per second at the peak, which one L4 covers. For 2,000 employees the same values give about 2.7 queries per second, beyond the L4 and within the L40S. An agent that searches several times per answer multiplies the rate.

Each answer also waits for the reranker before the LLM starts, 768 ms for 40 passages on the L4 and 234 ms on the L40S in NVIDIA’s tables with one request at a time. With three requests in parallel, the setting behind the queries per second above, the tables list 2,254 ms and 652 ms. Fewer passages per query reduce both load and wait.

We supply the L4, the L40S and the RTX PRO cards for retrieval models, as cards for your servers or in AI servers built to order. Tell us your document count, chunk length and peak queries per second, and we size the cards from them.

A separate card or a MIG instance on the LLM GPU

The first of three placements gives the retrieval models a small card of their own in the LLM server: the passive, low-profile L4 at 72 W, the RTX PRO 4000 SFF at 70 W, the 16 GB RTX PRO 2000 at 70 W, or the 32 GB RTX PRO 4500. The second is MIG on the LLM card. NVIDIA’s MIG guide lists up to four 24 GB instances on every RTX PRO 6000 edition and up to seven on the H200 NVL, each with its own memory and fault isolation, as our guide to GPU sharing with MIG, time-slicing and MPS explains.

The third is running the engines side by side on one card without partitioning. Each engine then reserves its own share of memory and they compete for compute, so run a full re-index at night or on a separate card. NVIDIA publishes no throughput for a MIG instance or for the RTX PRO 2000, 4000 and 4500 with these models, so test your model on the configuration you plan.

Example configurations by corpus size and query load

CORPUS AND LOADFULL INDEX, ESTIMATERERANKING CAPACITYCONFIGURATION
3M chunks, 1 query/s17 h on L4, 5.8 h on L40SL4: 1.3 queries/sone L4 in the LLM server, or a 24 GB MIG instance on an RTX PRO 6000 (test it)
3M chunks, 4 queries/s5.8 h on L40SL40S: 4.5 with 1B, 9.7 with 500mone L40S for both models, the 500m reranker for headroom
30M chunks, 4 queries/s58 h on one L40S, 16 h on an RTX PRO 6000 in NVFP4as abovetwo L40S in a retrieval server, or an RTX PRO 6000 Server Edition
Qwen3 8B pair, any corpusseveral times the 1B time; testno published figureone 48 GB L40S, or two 24 GB MIG instances

Our estimates: chunks of 512 tokens divided by the published texts per second above; reranking at 40 passages per query and three requests in parallel. Parsing, OCR and database writes excluded. The 8B pair’s 32 GB of BF16 weights as in our company-size article.

We check the rack, power and airflow before we quote, and return a configuration and quote within one business day. Send us the corpus size and the query load through the form below, rough estimates are enough.

What we supply

We supply the cards for retrieval models, the L4, L40S, RTX PRO 2000, 4000, 4000 SFF and 4500, and the RTX PRO 6000 and H200 NVL for MIG next to the LLM, as professional NVIDIA GPUs or in AI servers built to order. Servers are burn-in tested, with manufacturer warranty on every component, on one EU contract and invoice, and NVIDIA AI Enterprise licences for NIM come on the same invoice. The RAG platform itself, with assistants that cite the source and respect each user’s access rights, is our Private AI/ML service, with engineering by our partner Vixen.UNO.

FAQ

How much GPU memory does an embedding model need?
The weights take two bytes per parameter in BF16, about 1.2 GB for Qwen3-Embedding-0.6B and about 16 GB for the 8B model. The serving engine reserves more for its batches, and NVIDIA’s support matrix gives 3.6 GB for the non-optimised profile of its 1B embedding model and 33 GB for bge-m3. Measure the memory use with your engine’s batch settings on a test card.
How long does it take to embed a million documents?
Multiply documents by chunks per document and divide by the published or tested throughput of your model on your card. With NVIDIA’s figures for a 1B model at 512 tokens per chunk, 3 million chunks take about 17 hours on one L4 and about 5.8 hours on one L40S by our arithmetic. Parsing, OCR and database writes come on top, and an 8B embedding model takes several times as long.
Which GPU do I need for a reranker?
Size it by queries per second times passages per query. In NVIDIA’s NIM figures for 40 passages of 512 tokens with three requests in parallel, a 1B reranker handles about 1.3 queries per second on an L4 and about 4.5 on an L40S, and a 500m reranker about 9.7 on the L40S. One request at a time waits 768 ms on the L4 and 234 ms on the L40S, and about three times as long at that load.
Can the embedding model run on the same GPU as the LLM?
Yes, preferably in its own MIG instance: NVIDIA’s MIG guide lists up to four 24 GB instances on an RTX PRO 6000 and up to seven on an H200 NVL, each with its own memory and fault isolation. Without MIG, both engines share memory and compute, and a full re-index slows the chat users. The L4, L40S, RTX PRO 4000 and RTX PRO 2000 do not appear on NVIDIA’s MIG list.
Does Text Embeddings Inference run on RTX PRO Blackwell GPUs?
TEI 1.9 has an image for Blackwell compute capability 12.0, which its documentation marks as experimental and lists with GeForce examples. NVIDIA gives compute capability 12.0 for the RTX PRO 6000, 4500, 4000 and 2000 Blackwell cards, so they run from that image. The L4 and L40S use TEI’s regular Ada Lovelace image.
Which embedding model works for multilingual company documents?
Qwen3 Embedding, BGE-M3 and EmbeddingGemma each state support for more than 100 languages, and NVIDIA’s llama-nemotron-embed-1b-v2 lists 26 evaluated languages including German, French and Polish. Leaderboard positions change, so compare two or three candidates on an evaluation set built from your users’ own questions. Licences differ: Apache 2.0 for Qwen3, MIT for BGE-M3, Gemma Terms of Use and the NVIDIA Open Model License for the others.

Send us the number of documents, the average chunk count and length, the embedding and reranker models you are considering, the peak queries per second and the servers you run. We reply within one business day with a configuration and a quote, with the rack, power and airflow checked before we quote.

Talk to an expert
Talk to an expert

We reply within one business day

By sending this form you agree that we process your details to answer your enquiry; see our privacy policy.

request@eurokommerz.at
Jordangasse 7, 1010 Vienna