Exact search versus approximate IVFFlat, HNSW and DiskANN indexes, and how to tune each for recall and latency.
From Ultra Transcenders AI-200 by Tony Rough (publishing soon)
Without a vector index, pgvector performs an exact search with perfect recall but scans every row. Approximate nearest neighbour (ANN) indexes trade a little recall for much lower latency and compute, and each has build-time and query-time knobs.
| Aspect | IVFFlat | HNSW | DiskANN (pg_diskann) |
|---|---|---|---|
| Algorithm | Inverted file: k-means clusters into lists | Multi-layer navigable graph | Graph-based ANN from Microsoft Research |
| Build speed and memory | Fastest build, least memory | Slower build, more memory than IVFFlat | Slower and more memory than IVFFlat |
| Speed/recall trade-off | Lowest of the three | Better than IVFFlat | High recall, high QPS, low latency at large scale |
| Training step | Yes; build after data is loaded | No; can be created on an empty table | No; can be created on an empty table |
| Build options | lists |
m (default 16), ef_construction (default 64) |
max_neighbors (default 32), l_value_ib, product_quantized |
| Query-time setting | ivfflat.probes (default 1) |
hnsw.ef_search (default 40) |
diskann.l_value_is (default 100), diskann.iterative_search |
| Maximum indexed dimensions | 2,000 | 2,000 | 16,000 with product quantization (pg_diskann 0.6 or later) |
| Availability | pgvector | pgvector | Azure Database for PostgreSQL flexible server only |
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
CREATE INDEX ON documents USING diskann (embedding vector_cosine_ops);
SET hnsw.ef_search = 100;lists = rows / 1000 for up to 1 million rows, or sqrt(rows) for larger tables, and probes = lists / 10 (up to 1 million rows) or sqrt(lists). More probes raise recall and cost; setting probes equal to lists makes the search exact, at which point the planner doesn’t use the index. Load data before building, because the clusters are trained on existing rows.m or ef_construction for a better graph at higher build cost; raise hnsw.ef_search at query time for better recall and slower queries.max_neighbors 32 for under 1 million rows, 64 for 1-50 million and 96 above 50 million, with l_value_ib 100 and diskann.l_value_is 100; turn on product_quantized (preview) above 1 million rows to shrink the index so more of it stays in memory. With product quantization, Learn suggests pq_param_num_chunks of one third of the embedding dimensions and a two-step rerank (fetch, say, 50 candidates by ANN, then reorder by exact distance and return 10).Query-time parameters can be set for the session with SET or for one transaction with SET LOCAL inside BEGIN ... COMMIT, which is the safer choice with connection poolers.
maintenance_work_mem for the build session (for example SET maintenance_work_mem = '8GB';, depending on server memory), and consider scaling compute up for the build and back down afterwards.parallel_workers storage parameter and the max_parallel_maintenance_workers, max_parallel_workers and max_worker_processes server parameters. max_worker_processes needs a server restart; if resources don’t allow parallel workers, PostgreSQL falls back to a serial build. The leader process doesn’t take part in a parallel index build.pg_stat_progress_create_index.EXPLAIN (ANALYZE, VERBOSE, BUFFERS) to confirm an Index Scan using ... node; without an index, raising max_parallel_workers_per_gather can speed up exact searches.Common trap: Building an HNSW or IVFFlat index on
vector(3072)embeddings - both are limited to 2,000 dimensions and the build fails. Reduce dimensions (for example with the model’sdimensionsoption), or use DiskANN with product quantization, which supports up to 16,000.
Common trap: Creating an index with
vector_l2_opsand querying with<=>- the operator must match the operator class (<=>needsvector_cosine_ops), otherwise the planner falls back to a sequential scan.
This note is one section of Ultra Transcenders AI-200: Developing AI Cloud Solutions on Azure, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.
Publishing soon on Amazon in Kindle and paperback editions.
About the book · AI-200 terms in the glossary · All AI-200 study notes
The five Cosmos DB consistency levels, their RU and latency trade-offs, and when to choose each.
How to tell commands, discrete events and telemetry streams apart and pick the right Azure messaging service.
How to define KEDA scalers for queues, topics and other event sources in Container Apps.
What creates a new revision, and how single and multiple revision modes change deployments.
The main Redis caching patterns, how to expire and invalidate entries, and the trade-offs of each.
How the Functions hosting plans differ in scaling, networking and cold start, and which to choose.
Control plane versus data plane, Azure RBAC versus vault access policies, and the roles apps need.