Vector database comparison for RAG: pgvector, Qdrant, Milvus and OpenSearch
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- PostgreSQL with pgvector is enough when PostgreSQL already runs in-house and the vectors fit one server’s RAM; its SQL filter runs after the HNSW scan, so permission filters need the iterative index scans added in version 0.8.0 or an exact search through an index on the ACL column
- Qdrant filters during the HNSW traversal, using extra graph edges built from its payload indexes, and reads small filtered sets through the payload index; Milvus filters the entities first and runs the ANN search within them
- OpenSearch’s efficient filtering puts the filter inside the knn clause with the Lucene and Faiss engines, while a post-filter may return significantly fewer than k results; in a hybrid query, the top-level filter covers all subqueries
- A million 1,024-dimensional float32 vectors take 4.10 GB before any index; by each project’s own formula, HNSW brings Milvus to 4.22 GB at graph degree 32 and OpenSearch with Faiss to 4.65 GB at m = 16
- pgvector is under the PostgreSQL License, Qdrant, Milvus and OpenSearch under Apache 2.0; GPU acceleration is documented for Milvus indexes, Qdrant index builds and OpenSearch’s remote index build, not for pgvector
Eurokommerz × Vixen.UNO: Private AI/ML Talk to an expert →
pgvector, Qdrant, Milvus and OpenSearch compared
For a company RAG system with a few million chunks, PostgreSQL with pgvector is enough when PostgreSQL already runs in-house and iterative index scans are switched on for filtered queries. Qdrant and Milvus are databases built for vector search and apply metadata filters inside or before it; Milvus documents its Distributed mode for “100 million up to tens of billions of vectors”. OpenSearch suits a company that already runs an OpenSearch cluster and wants keyword and vector search in one query.
Our guide to a private ChatGPT alternative for a company places the database in the full stack, and our comparison of fine-tuning and RAG covers whether retrieval is the right approach.
| DATABASE | INDEX TYPES | FILTERING | LICENCE | FITS |
|---|---|---|---|---|
| PostgreSQL + pgvector 0.8.7 | HNSW, IVFFlat; halfvec and bit types; binary quantisation through an expression index | SQL condition after the index scan; iterative scans since 0.8.0; row security | PostgreSQL License | PostgreSQL already in production; vectors that fit one server’s RAM |
| Qdrant 1.19.2 | HNSW; scalar, binary, product and TurboQuant quantisation | payload filter during the HNSW traversal; optional ACORN since 1.16.0 | Apache 2.0 | a dedicated vector service with restrictive filters |
| Milvus 3.0.2 | IVF and HNSW families, SCANN, DISKANN and AISAQ on disk, four GPU indexes | filter first, then the ANN search within the matches | Apache 2.0 | large corpora, Distributed mode on Kubernetes, GPU indexes |
| OpenSearch 3.9.0 (k-NN) | HNSW (Lucene, Faiss), IVF and product quantisation (Faiss) | efficient filtering inside the knn clause; post-filters may return fewer than k | Apache 2.0 | an existing OpenSearch cluster; keyword and vector search together |
Each project’s documentation, release pages and LICENSE file, read 6 October 2026: pgvector 0.8.7 (PGXN, 1 October 2026), Qdrant v1.19.2 (5 October 2026), Milvus 3.0.2 (20 September 2026), OpenSearch 3.9.0 (29 September 2026).
Index types: HNSW, IVF, DiskANN and quantisation
All four offer HNSW, a graph index that trades exact results for speed. pgvector offers it next to IVFFlat; its README says HNSW has “better query performance than IVFFlat (in terms of speed-recall tradeoff), but has slower build times and uses more memory”. The vector type can be indexed up to 2,000 dimensions and halfvec up to 4,000; for larger embeddings the README lists binary quantisation, indexing subvectors and dimensionality reduction.
Qdrant “currently only uses HNSW as a dense vector index” and saves memory through quantisation: scalar quantisation stores each value as an 8-bit integer, a fourfold reduction, while binary quantisation and TurboQuant reach up to 32 times. Milvus lists FLAT, the IVF and HNSW families with quantised variants, SCANN, four GPU indexes, AISAQ and DISKANN, which keeps its graph on disk for datasets “too big to fit in memory”. OpenSearch builds HNSW with the Lucene or Faiss engine and IVF with Faiss; its on_disk mode quantises float vectors to 1 bit by default and rescores the top results with full-precision vectors read from disk.
Filtering by permissions during the vector search
Our guide to RAG on company data explains why access rights are enforced at retrieval, with each chunk carrying the users and directory groups allowed to read its source. An approximate index collects a limited number of candidates from its graph. If the permission filter runs only on those candidates afterwards, a user who may read a small share of the corpus receives fewer results than requested.
pgvector’s README says that with approximate indexes “filtering is applied after the index is scanned”. With HNSW at the default hnsw.ef_search of 40, a condition that matches 10% of rows leaves “only 4 rows” on average. Iterative index scans, added in version 0.8.0 of October 2024, keep scanning until enough rows pass, up to 20,000 visited tuples by default; SET hnsw.iterative_ switches them on. For filters that match few rows, the README suggests an ordinary index on the filter column, which gives exact search; for an array of allowed groups, PostgreSQL’s GIN index supports the overlap operator &&.
Row security policies can enforce the rule in the database. Superusers and BYPASSRLS roles always bypass them, and the table owner does too unless the table is set to FORCE ROW LEVEL SECURITY. PostgreSQL checks a policy on each row the scan returns, so for an HNSW index it acts like any other filter and needs iterative scans as well.
Qdrant filters during the graph search. It extends HNSW “with additional edges based on indexed payload values”, and its documentation asks for payload indexes before data is ingested. A query planner estimates how many points a filter matches and, below a threshold, reads them through the payload index and scores them exactly. Since v1.16.0, a query can switch on ACORN (off by default), which explores neighbours of neighbours when direct neighbours are filtered out and “improves search accuracy at the cost of performance”. Qdrant’s JWT tokens grant read or read-write rights per collection, so your application adds the per-user filter to every query.
Milvus applies the filter before the search, and its documentation lists the steps “Filter entities that match the filtering conditions” and “Conduct the ANN search within the filtered entities.” For complex expressions, "hints": "iterative_filter" in the search parameters searches in iterations and filters what each returns until enough results are found. Allowed groups are checked with ARRAY_CONTAINS_ANY, and RBAC ends at collection level, so here too the per-user filter comes from the application.
In OpenSearch, efficient filtering puts the filter inside the knn clause and applies it during the vector search, which “ensures that k results are returned (if there are at least k results in total)”, with Lucene HNSW from 2.4 and Faiss HNSW from 2.9. When its approximate search returns fewer than k, Faiss falls back to an exact search over the filtered documents. A post-filter “may return significantly fewer than k results for a restrictive filter”. Document-level security limits what a role can retrieve through a role query with variables such as ${user.name}; its documentation does not mention k-NN queries, so test that a restricted user still receives k results.
Assistants built in our Private AI/ML service cite the source and respect each user’s access rights. Describe where your documents are stored and how access to them is granted in the form below.
Hybrid search: keywords and vectors under one filter
Part numbers and error codes need keyword matching next to vector similarity, and the permission filter has to cover both parts, or restricted chunks return through the keyword search. pgvector’s README pairs it with “Postgres full-text search for hybrid search”, merged by Reciprocal Rank Fusion (RRF) or a cross-encoder; a row security policy covers both SQL queries, while a WHERE condition has to be written into each. Qdrant fuses candidates from dense and sparse prefetches with RRF or distribution-based score fusion and an IDF modifier for BM25-style weights; each prefetch takes its own filter, so repeat the permission filter in each.
Milvus turns a text field into sparse vectors with a built-in BM25 function. In its hybrid search each AnnSearchRequest has its own expr, so the permission expression goes into every request. In OpenSearch, hybrid search (introduced in 2.11) runs a hybrid query of up to five query clauses, and its top-level filter applies to “all the subqueries of the hybrid query”.
Memory per million vectors: the documented formulas
The raw size is number of vectors × dimensions × bytes per value, so a million 1,024-dimensional float32 embeddings need 1,000,000 × 1,024 × 4 = 4,096,000,000 bytes, or 4.10 GB, and 3.07 GB at 768 dimensions. Each project documents its own index estimate.
| DATABASE | FORMULA IN THE DOCS | 1M × 1,024 DIMS | TO REDUCE IT |
|---|---|---|---|
| PostgreSQL + pgvector | 4 × d + 8 bytes per vector; no HNSW size formula | 4.10 GB, plus an index that stores each vector again | halfvec: 2 × d + 8 bytes, 2.06 GB; a halfvec index |
| Qdrant | d × 4; graph m × 2 × 4 × 1.2; IDs 52 bytes; all × 1.2 | 4.30 GB; 5.16 GB with headroom | int8 scalar: 1 byte per value, 1.02 GB; originals on disk |
| Milvus (HNSW) | d × 4 + graph degree × 4 bytes | 4.22 GB at degree 32 | IVF_SQ8 or HNSW_SQ: 1 byte per value |
| OpenSearch (Faiss HNSW) | 1.1 × (4 × d + 8 × m) bytes | 4.65 GB at m = 16 | half_float type (3.9): 2.39 GB |
Our arithmetic from the pgvector README, Qdrant’s capacity planning page (m = 16), Milvus’s “Index explained” page (its example degree of 32) and OpenSearch’s k-NN pages (m = 16); d = dimensions, m = HNSW links per node; GB = 10⁹ bytes; text, metadata, row overhead and replicas excluded.
Qdrant’s headroom covers “the OS page cache, Qdrant’s runtime overhead, and temporary work during optimization”, and replicas multiply its point count. OpenSearch gives Faiss indexes a share of the memory left after the Java heap, 50% by default. In its documentation’s example, a 100 GB machine with a 32 GB heap leaves 68 GB, of which the k-NN plugin uses 34 GB; by our arithmetic, that holds about 7.3 million 1,024-dimensional vectors at m = 16. pgvector’s HNSW index stores each vector again next to its graph links. Build a test index and read its size with pg_relation_size; indexes build “significantly faster when the graph fits into maintenance_work_mem”.
We select and supply the GPU servers, fast storage and low-latency networking for such a platform, from a single server to a cluster. Send us the chunk count, the embedding dimensions and the growth you expect through the form below.
GPU acceleration: index builds and GPU indexes
Milvus has four GPU indexes, GPU_CAGRA, GPU_IVF_FLAT, GPU_IVF_PQ and GPU_BRUTE_FORCE. GPU_CAGRA uses about 1.8 times the memory of the original vector data, and with adapt_for_cpu it is built on the GPU and searched on the CPU. Since v1.13.0, Qdrant can build its indexes on GPUs, through Vulkan 1.3 in separate gpu-nvidia images, with up to 16 GB of vector data per GPU and indexing iteration. OpenSearch 3.0 added a GPU-accelerated remote build service for Faiss HNSW indexes, with an Amazon S3 repository as the intermediate store. pgvector’s README does not mention GPUs. For a few million chunks, size the GPUs for the language and embedding models first; GPU index builds matter once a rebuild no longer fits the maintenance window.
Operations: high availability, upgrades and backups
pgvector “uses the write-ahead log (WAL), which allows for replication and point-in-time recovery”, so your PostgreSQL standbys and archiving cover the vectors. Version 0.8.7 of 1 October 2026 fixes a buffer overflow in IVFFlat index builds (CVE-2026-103484), and the project advises all users of earlier versions to upgrade; an upgrade ends with ALTER EXTENSION vector UPDATE; in each database. Qdrant’s replication factor defaults to 1, “meaning no additional copy is maintained automatically”. With a factor of at least 2 on every collection, restarting the nodes one at a time upgrades a cluster without downtime, and each upgrade passes through the latest patch of every intermediate minor version. Connections stay unencrypted until TLS is enabled.
Milvus Standalone runs from one Docker image and Distributed on Kubernetes, with metadata in etcd, data in object storage such as MinIO and a write-ahead log on Woodpecker, Kafka or Pulsar. Consistent backups of pgvector, Qdrant and Milvus are covered in our article on backing up an AI server.
When PostgreSQL with pgvector is enough
pgvector is enough when PostgreSQL already runs in production with replication and backups, the vectors fit one server’s RAM, the embedding model has at most 2,000 dimensions (4,000 as halfvec) and permissions live in tables that a query can join or a row security policy can enforce. By our arithmetic, 300,000 documents at an average of ten chunks give 3 million chunks, or 12.3 GB of vector data at 1,024 dimensions, and an HNSW index stores the vectors a second time. VMware’s Private AI Services also indexes its knowledge bases in pgvector, as our article on VMware Private AI Foundation describes.
A dedicated engine makes sense when the vectors outgrow one server’s RAM, when quantisation with built-in rescoring has to keep memory within budget, when filters should act during the graph search rather than after it, or when vectors should run apart from the transactional databases. OpenSearch fits where its cluster and security model already exist and keyword search matters as much as similarity.
What we do
Under our Private AI/ML service, our engineering partner Vixen.UNO builds RAG platforms whose assistants cite the source and respect each user’s access rights, with a query log and data and permissions management. The first call is free of charge; the paid technical assessment covers the solution architecture, the model and GPU selection and a pilot plan with metrics, at a price fixed before work begins. Eurokommerz holds the contract and supplies the AI servers, storage and networking, with EU invoicing and warranty under European law. The platform runs on your servers or on dedicated hardware in a Tier-3 data centre in Lithuania.
FAQ
Which vector database should I use for RAG?
What is the difference between pgvector and Qdrant?
Milvus or Qdrant: which one fits a company RAG system?
How does OpenSearch vector search handle filters?
How much memory does a vector database need per million vectors?
How do I filter vector search results by user permissions?
Send us the number of documents and chunks, the embedding model and its dimensions, how access rights are managed today and which databases you already run. We reply within one business day with next steps, starting with a first call that leaves you with two or three possible solution scenarios. The first call is free of charge.
Talk to an expertWe reply within one business day