Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Embedding Models

By Anup Rai6 min readReviewed September 2026

An embedding model maps an input into a numerical representation learned for tasks such as similarity search. For retrieval, compatible query and document representations are scored to rank candidate evidence. Similarity is a learned relevance signal, not proof that a passage answers the question.

Start with the geometry of embeddings. This chapter covers the production choices: model contract, representation size, quantization, modalities, evaluation and migration.

Establish the model contract

  1. Choose the supported input modalities and languages.
  2. Follow the documented query/document prompts or task types.
  3. Use compatible model versions, dimensions and normalization.
  4. Select the similarity metric expected by the model.
  5. Count and handle inputs exceeding the model's limit explicitly.
  6. Record model, preprocessing, task, dimension and quantization versions with the index.

Equal vector length does not mean two models share an embedding space. Switching the query encoder while retaining incompatible document vectors can silently degrade retrieval.

Single-vector, sparse and multi-vector representations

Representation Stored object Advantage Tradeoff
Dense single-vector One dense vector per passage or object Compact and widely supported Compresses many details into one representation
Learned sparse Weighted vocabulary dimensions Lexical expansion with sparse indexing Model inference and posting-list costs
Multi-vector late interaction Several token or patch vectors More detailed query-document matching Larger index and more complex scoring
Joint multimodal Compatible representations across supported modalities Text can retrieve images or pages Fine-grained evidence may still need another step

ColBERT is a late-interaction approach. Query and document token vectors are encoded separately; scoring combines the strongest document-token match for each query token. ColBERTv2 compression and the PLAID search engine reduce different parts of this cost; they are not names for a universal compression guarantee.

Matryoshka representations

Matryoshka Representation Learning trains nested representation prefixes to remain useful at multiple dimensions. A supported shorter prefix can be used for inexpensive candidate search, followed by higher-dimensional scoring of those candidates. Arbitrarily truncating any embedding model does not provide this property. Matryoshka Representation Learning.

For a model supporting both dimensions, reducing 1,536 float32 values to 256 reduces raw vector payload from 6,144 to 1,024 bytes: sixfold. Quality loss depends on the model, corpus and task. There is no general “less than 2% loss” promise.

Normalize the shortened output if the model's contract requires it. Changing the norm can affect cosine/dot-product equivalence. Candidate loss in the small representation cannot be repaired by rescoring candidates that never included the relevant document.

Quantization changes storage, not the task definition

Encoding Ideal payload per dimension What to evaluate
Float32 4 bytes Baseline quality and memory
Float16 2 bytes Numeric sensitivity and engine support
Int8 1 byte plus quantization metadata Calibration ranges and recall
Packed binary 1 bit plus metadata Candidate recall and rescoring requirements

Binary payload is ideally 32 times smaller than float32 at the same dimension. End-to-end memory savings also depend on index links, IDs, metadata, replication and retained full-precision vectors. Hamming-distance operations can be efficient, but a universal tenfold search speedup is not justified.

Distinguish output-vector quantization from quantizing the embedding model's own weights. The former compresses stored representations; the latter changes model execution and may alter the produced vectors. A provider returning floats does not automatically provide native int8 embeddings just because a database can quantize them afterward.

A billion-vector capacity exercise

For one billion 1,536-dimensional float32 vectors:

1,000,000,000 × 1,536 × 4 = 6,144,000,000,000 bytes
                             = 6.144 TB of raw vector payload

This is not the full HNSW memory requirement. A 128-dimensional packed-binary candidate representation would contain 16 GB of raw payload, a 384-fold payload reduction from that baseline. It requires a compatible representation and measured recall. If full vectors remain for rescoring, their 6.144 TB still exist somewhere in the system.

Capacity planning must include how those full vectors are fetched, how many candidates are rescored, SSD/network bandwidth and the effect on p99 latency. The smallest candidate index is not necessarily the cheapest complete service.

A dated model comparison, not a leaderboard

The following distinctions were checked against primary documentation for this review on September 24, 2026. Treat model IDs and modality contracts separately from marketing rankings.

Model or family Relevant distinction Source
Gemini Embedding 2 Multimodal embedding space; text, images, audio, video and PDFs under documented limits Google documentation
Gemini Embedding 001 Text embedding model; do not attribute Embedding 2's modalities to it Google model comparison
Qwen3-Embedding Open-weight text embedding sizes with instruction and dimension options Qwen model card
Cohere Embed v4 Text/image support and multiple output representations Cohere model documentation
Voyage Multimodal 3.5 Multimodal retrieval with supported dimension choices Voyage documentation
BGE-M3 Dense, sparse and multi-vector retrieval capabilities BAAI model card
Jina embedding families Separate text and multimodal models; task and context limits vary by model Jina embedding API
NVIDIA embedding families Research and production-oriented model variants require individual contract/license review NVIDIA model card

Jina embeddings v3, specifically, is a text embedding model with an 8,192-token context, not a 128k ColBERT-style model. Newer Jina text and multimodal models have their own contracts. Jina v3 paper.

A public benchmark average helps shortlist candidates. It does not settle accuracy on your identifiers, languages, document layouts, access filters or latency target. Check license and data handling for the actual deployment; open weights do not imply free operation or unrestricted commercial rights.

Evaluate and migrate safely

Architecture / visual model
flowchart LR S[Versioned source passages] --> A[Current model and index] S --> B[Candidate model and separate index] Q[Held-out query set] --> A Q --> B A --> C[Compare evidence recall and answer quality] B --> C C --> D[Shadow traffic and controlled cutover]
Read diagram source
flowchart LR
    S[Versioned source passages] --> A[Current model and index]
    S --> B[Candidate model and separate index]
    Q[Held-out query set] --> A
    Q --> B
    A --> C[Compare evidence recall and answer quality]
    B --> C
    C --> D[Shadow traffic and controlled cutover]

Compare query encoding latency, ingestion throughput, retrieval relevance, filtered ANN recall, multilingual slices and complete-answer quality. Include exact codes and new terms. Hybrid search can preserve lexical signals, but it does not guarantee recovery from every out-of-domain failure.

Re-embed documents when compatibility requires it. Build a separate index, keep ingestion changes synchronized, shadow queries and switch query encoder plus index together. Retain a rollback path until quality and operational checks pass. Reusing old vectors solely because dimensions match is not a migration strategy.

Interview practice

Q1: What matters more than embedding dimension?

Task alignment, correct query/document formatting, supported language and modality, representative evaluation and operating cost. More dimensions can help retain information, but cannot repair a model trained for the wrong matching objective.

Q2: When would you use Matryoshka embeddings?

When supported nested dimensions allow a useful memory/quality tradeoff or a cheap candidate stage followed by rescoring. I would measure candidate recall first, because the full representation cannot rescue a missing candidate.

Q3: Why can a binary index disappoint despite large payload savings?

Index overhead and full-vector storage may dominate. Quantization can reduce recall, requiring deeper search or rescoring. I would compare complete service memory, bandwidth, latency and final relevance rather than the vector byte count alone.

Q4: Does a new product name make an embedding model unusable?

Not necessarily; tokenization and learned composition may still represent it usefully. But exact identifiers and new domain terms need testing. Lexical search, domain data and reranking are possible remedies, each with costs and limitations.

Q5: Does visual retrieval remove all need for OCR?

It can retrieve pages without relying on an OCR-only index. The answering or extraction stage still needs to read values, resolve layout and provide evidence. Compare visual and text paths on the actual document tasks rather than treating them as mutually exclusive.

Q6: Can you replace the embedding API without rebuilding the index?

Only if the representations are documented and validated as compatible. Otherwise re-embed into a separate index and cut over atomically with the query encoder. Equal dimensions or similar leaderboard scores do not establish compatibility.

Final notes

Recall card: Contract, compatibility, quality, capacity, cutover. Select the complete retrieval system, not the largest vector or the current benchmark winner.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Chunking Strategies
NEXT LESSONVector Databases and Search Indexes →

Explore the diagram