FREE STUDY NOTES · AI-200

pgvector indexes on Azure PostgreSQL: IVFFlat vs HNSW vs DiskANN

Exact search versus approximate IVFFlat, HNSW and DiskANN indexes, and how to tune each for recall and latency.

From Ultra Transcenders AI-200 by Tony Rough (publishing soon)

Without a vector index, pgvector performs an exact search with perfect recall but scans every row. Approximate nearest neighbour (ANN) indexes trade a little recall for much lower latency and compute, and each has build-time and query-time knobs.

Comparing the index types

Aspect IVFFlat HNSW DiskANN (pg_diskann)
Algorithm Inverted file: k-means clusters into lists Multi-layer navigable graph Graph-based ANN from Microsoft Research
Build speed and memory Fastest build, least memory Slower build, more memory than IVFFlat Slower and more memory than IVFFlat
Speed/recall trade-off Lowest of the three Better than IVFFlat High recall, high QPS, low latency at large scale
Training step Yes; build after data is loaded No; can be created on an empty table No; can be created on an empty table
Build options lists m (default 16), ef_construction (default 64) max_neighbors (default 32), l_value_ib, product_quantized
Query-time setting ivfflat.probes (default 1) hnsw.ef_search (default 40) diskann.l_value_is (default 100), diskann.iterative_search
Maximum indexed dimensions 2,000 2,000 16,000 with product quantization (pg_diskann 0.6 or later)
Availability pgvector pgvector Azure Database for PostgreSQL flexible server only
CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops)
    WITH (m = 16, ef_construction = 64);
CREATE INDEX ON documents USING diskann (embedding vector_cosine_ops);
SET hnsw.ef_search = 100;

Tuning each index

Query-time parameters can be set for the session with SET or for one transaction with SET LOCAL inside BEGIN ... COMMIT, which is the safer choice with connection poolers.

Building indexes efficiently

Common trap: Building an HNSW or IVFFlat index on vector(3072) embeddings - both are limited to 2,000 dimensions and the build fails. Reduce dimensions (for example with the model’s dimensions option), or use DiskANN with product quantization, which supports up to 16,000.

Common trap: Creating an index with vector_l2_ops and querying with <=> - the operator must match the operator class (<=> needs vector_cosine_ops), otherwise the planner falls back to a sequential scan.

Get the whole book

This note is one section of Ultra Transcenders AI-200: Developing AI Cloud Solutions on Azure, an independent study guide that explains every topic the exam covers by technology, with comparison tables, diagrams and the common traps, plus a glossary linked to Microsoft Learn.

Amazon.co.ukKindle: coming soonPaperback: coming soon
Amazon.comKindle: coming soonPaperback: coming soon

Publishing soon on Amazon in Kindle and paperback editions.

About the book · AI-200 terms in the glossary · All AI-200 study notes

More AI-200 study notes