A vector database stores vector representations and supports similarity search, usually alongside identifiers, metadata, updates and operational controls. Vector search can also be a capability inside a general-purpose database or search engine. A vector-index library supplies algorithms; it does not necessarily supply authentication, replication, backups or a complete database service.
In an interview, select the workload before the product. Ten thousand frequently updated tenant records and a billion mostly immutable image embeddings call for different designs.
Define requirements and scale
Functional requirements for a retrieval service might be:
- Upsert versioned vectors and metadata with stable document/chunk IDs.
- Retrieve the nearest permitted candidates for a query vector.
- Apply tenant, access, version and other business filters.
- Propagate content updates, deletions and permission revocations.
- Support the required dense, sparse, hybrid or multi-vector scoring contract.
Non-functional requirements include:
- Recall against exact neighbors under the chosen metric and filters.
- Relevance to the actual task, measured separately from ANN recall.
- p95/p99 latency and throughput under concurrent reads and writes.
- Freshness and deletion-propagation bounds.
- Availability, recovery time, recovery point and tenant isolation.
- Total cost, including indexing, replicas, storage, operations and migration.
Specify corpus size, dimensions, data type, query rate, filter selectivity and growth. A universal “100 million vectors under 100 ms” requirement is not a substitute for these inputs.
Exact search versus approximate search
An exact scan compares the query with every eligible stored vector. For N vectors of dimension d, a straightforward dense scan costs O(Nd) arithmetic. Hardware, batching and selective filters can make exact search practical for some workloads.
Approximate nearest-neighbor search (ANN) avoids examining every vector and trades resources and latency against the likelihood of recovering the exact nearest neighbors. There is no fixed 95–99% recall range or universal constant-time guarantee.
ANN recall@k compares returned IDs with the exact top-k IDs for the same metric and eligible set. If an approximate top-10 contains eight of the exact top-10, recall is 0.8. Exact search has perfect neighbor recall under that definition; it does not have perfect semantic relevance or answer accuracy.
Choose the metric the embedding expects
| Measure | Definition | Direction and range |
|---|---|---|
| Dot product | a · b |
Larger is more similar; unbounded in general |
| Cosine similarity | (a · b) / (norm(a) × norm(b)) |
Larger is more similar; −1 to 1 for nonzero vectors |
| Cosine distance | 1 − cosine_similarity |
Smaller is closer; 0 to 2 |
| Euclidean distance | sqrt(Σ(a_i − b_i)²) |
Smaller is closer; nonnegative |
| Hamming distance | Number of differing bits | Used for compatible binary representations |
For unit-normalized vectors, dot product equals cosine similarity, and squared L2 distance is 2 − 2(a · b), so they induce the same ordering. Do not select a metric merely because the inputs are text or images. Follow the model and engine contracts, including whether an API returns a distance, similarity or negative inner product.
Understand the index choices
HNSW
Hierarchical Navigable Small World search builds a layered proximity graph. Upper layers provide longer-range navigation; lower layers refine the search among more vectors. Query-time exploration controls how many alternatives are examined. HNSW paper.
| Parameter | Effect of increasing it | Main cost |
|---|---|---|
M or connectivity setting |
More graph connections | Memory and build work |
| Construction exploration | More thorough neighbor selection during build | Index construction time |
| Search exploration | More candidate exploration at query time | Query CPU and latency |
Exact names and limits differ by implementation. HNSW does not need a centroid-training phase, but graph construction is still work. Inserts, deletes and compaction have engine-specific behavior. Do not promise logarithmic worst-case latency or a fixed graph-to-vector memory ratio.
IVF and product quantization
An inverted-file index (IVF) partitions vectors around learned centroids. A query searches selected partitions. nlist commonly denotes the number of partitions and nprobe how many are searched. More probes generally examine more candidates at higher cost.
New vectors can be assigned to existing centroids; every insert does not require retraining. Distribution drift or a poor training sample may eventually warrant rebuilding the partitioning. Evaluate that under the real update workload.
Product quantization (PQ) divides a vector into subvectors and stores codebook IDs. It reduces representation size and enables approximate distance calculations. Retaining original vectors permits later rescoring, but increases storage and fetch cost. IVF and PQ are separate techniques that can be combined. Faiss index documentation.
Disk-oriented ANN
DiskANN is a family of graph-based ANN techniques designed to use SSD capacity while limiting RAM needs. Its implementations include different update and filtering capabilities. It is not restricted to offline search. Benchmark storage latency, I/O parallelism, caching and recall before claiming a cost or speed advantage. Microsoft DiskANN research.
Flat search
A flat index scans vectors without ANN pruning. It provides the reference neighbors needed to measure ANN recall and can be a useful deployment choice for a small or tightly filtered candidate set. “Flat” does not mean storage is free: all vectors still have to be stored or fetched.
Build the service around the index
Read diagram source
flowchart TD
S[Versioned source of truth] --> W[Idempotent indexing worker]
W --> I[Vector and metadata index]
U[Authenticated query] --> A[Resolve tenant and access scope]
A --> E[Compatible query encoder]
E --> R[Filtered candidate search]
I --> R
R --> V[Revalidate access and source versions]
V --> P[Return evidence or rerank candidates]
S --> V
The source of truth owns the document and its current access rules. The index is a derived view with explicit freshness behavior. Query encoders and indexes must use compatible representation versions; see embedding models.
Filtered search is a distinct workload
Filtering may occur before candidate generation, during index traversal, after an ANN candidate stage, or through a combination. The correct performance tradeoff depends on selectivity and implementation.
If only 1% of candidates satisfy a filter, a global top-100 can leave roughly one result under an independence assumption. Real correlations may make it better or worse. Oversampling is not a guaranteed substitute for a filtered-search strategy.
A highly selective filter can make an exact scan of the allowed subset attractive. A graph traversal that simply refuses to traverse excluded nodes may lose routes to eligible neighbors; specialized implementations handle this in different ways. Benchmark restrictive and broad filters rather than assigning one latency to the database.
Security boundary: forbidden payloads must not reach the model, client, logs or an unauthorized reranker. An internal candidate-ID stage followed by enforced authorization can be a legitimate architecture; filtering only after sending full passages to the model is too late. Qdrant documents its filter operators and application patterns; pgvector documents how filtering interacts with approximate scans and iterative scans. Qdrant filtering, pgvector.
Capacity: separate payload, index and replicas
For 10 million vectors, 1,536 dimensions and float32 storage:
vectors: 10,000,000 × 1,536 × 4 = 61.44 GB
metadata at an assumed 500 bytes/record = 5.00 GB
subtotal before graph links, IDs and engine overhead = 66.44 GB
Three complete replicas would store 199.32 GB of that subtotal. Index overhead, logs, temporary build space and backups add more. Quantization changes the vector part; sharding divides a copy across machines; replication creates additional copies. These operations are not interchangeable.
There is no valid general formula that multiplies stored gigabytes by a constant to obtain QPS. Query cost depends on dimensions, index exploration, filters, hardware, cache state, concurrency and updates. Benchmark at the intended load, including a replica failure and an index rebuild.
Choose a deployment category
| Option | Why evaluate it | Responsibilities that remain |
|---|---|---|
| PostgreSQL with pgvector | Existing relational data, SQL filters and transactional integration | Query planning, indexes, vacuum, capacity and database operations |
| Dedicated vector database | Vector-specific indexing, filtering and scale features | Access design, relevance, data lifecycle and measured operations |
| Search engine with vector support | Existing lexical search and hybrid retrieval | Analyzer/index design, ranking and consistency across signals |
| Embedded/index library | Local or custom retrieval requirements | Service, authentication, persistence and recovery as needed |
pgvector supports exact search, HNSW and IVFFlat; it is not limited to brute force or a fixed ten-million-vector ceiling. Its suitability follows the workload and deployment. Managed services and self-hosted engines both require benchmarking.
Pinecone, Qdrant, Weaviate, Milvus/Zilliz and Chroma are examples to evaluate, not a universal ordering of winners. Consult current product documentation for deployment, sparse/multi-vector support, limits and pricing. Do not assume every engine uses HNSW internally, every Milvus deployment needs Kubernetes, or that all hosted services scale instantly to zero. Pinecone, Qdrant, Weaviate, Milvus, Chroma.
Multi-tenancy and lifecycle operations
| Isolation choice | Benefit | Cost or limitation |
|---|---|---|
| Shared collection with enforced tenant filters | Shared capacity and fewer indexes | Filter correctness and noisy-neighbor controls |
| Tenant namespace or partition | Clearer data grouping where supported | Semantics and resource isolation depend on the engine |
| Separate collection or database | Independent policies and operational choices | More objects, overhead and management |
| Separate deployment | Stronger resource/failure separation | Highest operational and capacity cost |
Resolve tenant identity from authentication, not an arbitrary model argument. Apply it to reads, writes, backups and caches. Separate collections alone do not compensate for a server that lets a caller select another tenant's collection.
Make upserts idempotent with source/version IDs. Propagate deletions to every derived record. Track index lag and failed ingestion. Distinguish visibility of a write from durable storage and replicated availability; those guarantees vary by engine and configuration.
For high availability, specify replica placement, failure detection, failover, write consistency and recovery. Asynchronous replication can leave stale reads or lose recently acknowledged data under some failure/configuration combinations. A diagram with three replicas does not define the guarantee. Test backup restoration and a full rebuild from the source of truth.
Cost and release decisions
Compare a realistic managed quote with self-hosted compute, storage, replicas, backups, egress, staff time and incident capacity. Managed service reduces some operational work; it does not make it zero. Self-hosting can offer control but does not automatically remove lock-in or produce lower total cost.
Use a proof of concept that measures exact-neighbor recall, task relevance, filtered p99 latency, update/delete visibility, failure recovery and total cost. Include hybrid and multi-vector queries only if the workload needs them. Keep exportable source data and index configuration so a vendor or model migration is feasible.
Interview practice
Q1: Why not start with a dedicated vector database automatically?
The existing database may already meet the requirements with fewer moving parts. I would compare it using representative scale, filters and concurrency. Add a dedicated service when measured workload needs or operational boundaries justify the extra system.
Q2: What does 100% exact-search recall mean?
It means recovering the nearest vectors under the specified metric and eligible set. Those vectors may still represent irrelevant or outdated passages. Evaluate semantic relevance and answer quality separately.
Q3: How would you tune HNSW?
Establish exact neighbors, then measure search exploration against recall and latency. Build parameters affect memory and graph quality, so changing them may require rebuilding. Include restrictive filters, concurrent updates and realistic hardware.
Q4: When would disk-oriented search help?
When vector/index capacity makes RAM expensive and the SSD-based design meets the latency and recall requirements. Account for I/O, caching, full-vector rescoring and update behavior. Do not infer a fixed 90% cost reduction from the algorithm name.
Q5: Why can filtered ANN return too few results?
A limited approximate candidate set may contain few eligible records. Increase exploration, use an engine's filtered or iterative search, or search a selective allowed subset exactly. The chosen strategy must enforce access before evidence exposure.
Q6: What happens when a document is deleted?
Record a versioned deletion, remove its chunks and derived vectors, invalidate relevant caches and verify visibility within the required bound. Handle indexing retries without resurrecting old content. A successful source deletion does not prove every derived system has applied it.
Q7: How would you justify managed versus self-hosted?
Compare complete workload cost and the team's operational capability, data constraints, availability needs and migration plan. There is no universal vector-count threshold at which the answer changes. An inexpensive node is not a complete highly available service.
Final notes
Recall card: Metric → index → filters → lifecycle → operations → cost. An interview answer should explain the evidence that would make you change the initial choice.