Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Hybrid Search

By Anup Rai7 min readReviewed September 2026

Hybrid search combines multiple retrieval signals, commonly lexical and dense-vector relevance, to produce a ranked candidate set. Its purpose is to cover different ways a query and document can match. It is a design to evaluate, not a guarantee that two retrievers always outperform one.

A query for NVIDIA_VISIBLE_DEVICES needs precise identifier handling. A query for “restrict which accelerators a process can see” may benefit from semantic matching. The same corpus can contain both needs.

Understand the signals

Signal How it matches Typical strength Failure to test
Lexical, such as BM25 Analyzed terms and corpus statistics Names, codes and explicit terminology Vocabulary mismatch or unsuitable tokenization
Dense embedding Learned vector similarity Paraphrases and conceptual relations Near meaning but wrong identifier or version
Learned sparse, such as SPLADE Learned weights over vocabulary dimensions Lexical expansion with sparse retrieval Expansion noise and inference/index cost
Structured filter Exact fields or business constraints Tenant, product ID, effective version Wrong metadata or stale access policy

BM25 includes term-frequency saturation and document-length normalization; it is not merely word counting. Its analyzer can normalize words or include synonyms, so “sparse search never understands synonyms” is too broad. Dense retrieval can also retrieve exact terms; the point is to test where each signal fails.

If a product or version is a hard requirement, use a verified structured constraint where possible. A fusion score should not be allowed to silently replace a required exact match with a similar product.

Draw the baseline architecture

Architecture / visual model
flowchart TD Q[Authenticated query and required filters] --> L[Permitted lexical retrieval] Q --> E[Query embedding] E --> D[Permitted dense retrieval] L --> F[Deduplicate and fuse candidate ranks] D --> F F --> R[Optional reranking] R --> V[Validate access and source versions] V --> C[Select evidence within token budget]
Read diagram source
flowchart TD
    Q[Authenticated query and required filters] --> L[Permitted lexical retrieval]
    Q --> E[Query embedding]
    E --> D[Permitted dense retrieval]
    L --> F[Deduplicate and fuse candidate ranks]
    D --> F
    F --> R[Optional reranking]
    R --> V[Validate access and source versions]
    V --> C[Select evidence within token budget]

The two retrieval paths need compatible document IDs, source versions and access policies. Returning the newest policy from one index and an old version from the other creates a lifecycle problem that fusion alone cannot solve.

Reciprocal rank fusion (RRF)

RRF combines ranked lists without assuming their raw scores have comparable scales:

RRF(d) = sum over lists containing d of 1 / (c + rank_i(d))

Ranks start at 1. c is a positive smoothing constant, often written k in the literature; it is not the final result count. The original paper used 60 in its experiments. It is a practical starting point to evaluate, not a universal optimum. Cormack, Clarke and Buettcher.

Consider two candidate lists:

Lexical: A, B, C
Dense:   B, D, A
Document Lexical rank Dense rank RRF with c = 60
A 1 3 1/61 + 1/63 ≈ 0.032266
B 2 1 1/62 + 1/61 ≈ 0.032522
C 3 absent 1/63 ≈ 0.015873
D absent 2 1/62 ≈ 0.016129

The fused order is B, A, D, C. RRF rewards agreement, even though neither raw cosine nor BM25 values appear in the calculation. It discards score-gap information: an overwhelming first-place match and a narrow first-place match receive the same rank contribution.

A runnable fusion example

This example uses only the Python standard library. It takes already authorized document IDs, rejects duplicate IDs within a list, and defines a deterministic tie-break.

from collections import defaultdict

def rrf(rankings, c=60):
    if c <= 0:
        raise ValueError("c must be positive")
    scores = defaultdict(float)
    for ranking in rankings:
        if len(set(ranking)) != len(ranking):
            raise ValueError("duplicate ID inside a ranking")
        for rank, doc_id in enumerate(ranking, start=1):
            scores[doc_id] += 1.0 / (c + rank)
    return sorted(scores.items(), key=lambda item: (-item[1], item[0]))

ranked = rrf([["A", "B", "C"], ["B", "D", "A"]])
assert [doc_id for doc_id, _ in ranked] == ["B", "A", "D", "C"]
print([(doc_id, round(score, 6)) for doc_id, score in ranked])

A document absent from a list contributes zero for that list; it is not assigned an invented last rank. Production code also needs stable identity across versions, response-size bounds and the application's access controls.

Weighted and relative score fusion

Weighted fusion commonly uses:

score(d) = alpha × normalized_dense(d)
         + (1 − alpha) × normalized_lexical(d)

The normalization, missing-candidate convention and score directions are part of the definition. A distance where smaller is better must not be added as if larger were better. Cosine similarity can be negative; it is not universally between zero and one.

Min-max normalization maps each list's score range to a common interval. It preserves relative gaps within that list but can be unstable with outliers or constant scores. Z-score normalization uses the mean and standard deviation and can produce negative values. It is another choice, not the universal meaning of “relative score fusion.” Weaviate's named relative-score method normalizes scores using the result range. Weaviate hybrid search.

Fusion method Keeps score gaps? Main tuning responsibility
RRF No Candidate depth, smoothing and optional list weights
Weighted normalized scores Yes Normalization, direction, weights and missing values
Learned ranking Can use ranks, scores and other features Labels, validation, drift and model complexity

An alpha of 0.5 means equal coefficients under the chosen normalization. It does not prove equal influence or equal usefulness. Do not memorize “technical queries use 0.3” as a general rule.

Learned sparse retrieval

SPLADE produces sparse vocabulary-space representations using a neural model. It can assign weight to related terms that were not literally present, reducing some vocabulary mismatch. That expands the lexical signal but does not guarantee correct matching for every unseen code or identifier. SPLADE paper.

Store vocabulary IDs and weights under a versioned tokenizer/model contract. Decoding each token to a display string is not a safe unique index key: different tokens or forms may be displayed ambiguously. Apply the model's attention mask, pooling and inference configuration correctly.

A database may support BM25, learned sparse vectors and dense vectors in one service. Whether it does so in one API request says little about the number of internal execution stages. Sparse support does not make SPLADE mandatory or BM25 obsolete.

Choose an implementation shape

Architecture Benefit Cost or failure mode
Separate engines, parallel calls Independent retrieval systems and scaling Two update paths, consistency and failure handling
One engine with hybrid queries Shared identity and simpler operations Product-specific fusion and indexing behavior
Staged lexical then semantic scoring Bounded expensive candidate scoring Evidence missed by the first stage cannot be recovered later
Adaptive retriever selection Avoids unnecessary work for some queries Router errors and harder evaluation

Current engines expose different hybrid interfaces. Qdrant's query API supports prefetch stages and fusion; Weaviate has its own hybrid query contract. Pin a client/server version and use its documentation instead of combining old SDK examples into an API that does not exist. Qdrant hybrid queries.

Tune candidate depth and latency together

  1. Create judged queries covering identifiers, paraphrases, mixed intent, languages and missing evidence.
  2. Compare lexical-only and dense-only baselines.
  3. Sweep each retriever's candidate depth and the fusion policy on development data.
  4. Measure relevant-evidence coverage, ranking metrics and final-answer quality.
  5. Inspect performance under filters, stale-index conditions and partial failures.
  6. Confirm the selected policy on held-out data at production-like load.

Fetching 3–5 times the final result count is a possible experiment, not a recall guarantee. Two shallow lists can both miss the same essential evidence. Conversely, very deep lists may add latency and weak candidates.

For an illustrative parallel flow, query embedding takes 25 ms, dense search takes another 20 ms, lexical search takes 30 ms, and fusion takes 3 ms. Ignoring other overhead, completion is max(25 + 20, 30) + 3 = 48 ms. A sequential implementation would take 78 ms. Real p99 latency requires measuring correlated load, queueing and timeouts; percentile values cannot simply be added as if they were fixed durations.

Cache and failure handling

Cache keys need the query, normalized filters, permission scope, corpus/index versions, embedding model and ranking configuration as applicable. A query-only cache can return another tenant's result or a deleted document. Revalidate consequential access changes before exposing evidence.

Set independent retrieval deadlines and define whether one healthy retriever can serve a degraded result. Label and measure that path. If the missing retriever is necessary for exact-identifier accuracy or access correctness, an explicit unavailable outcome is better than silently pretending the full search ran.

Monitor per-arm latency, candidate overlap, zero-result rates, evidence recall, stale-version conflicts and end-to-end success. A higher overlap is not automatically better: the point of hybrid search is often complementary coverage.

Interview practice

Q1: When is hybrid search useful?

When the workload needs both exact lexical cues and semantic matching, and evaluation shows complementary failures. I compare it with each single retriever and account for extra indexing, query cost and lifecycle complexity.

Q2: Why use RRF instead of adding BM25 and cosine scores?

Their scales and meanings differ. RRF uses rank contributions, so it avoids direct scale mixing. The tradeoff is losing score gaps and depending on candidate-list depth. It still requires relevance evaluation.

Q3: Can hybrid search be worse than either component?

Yes. A weak arm can promote irrelevant candidates, poor normalization can dominate the score, or fusion can demote an exact required match. Use hard constraints for hard requirements and inspect failure slices.

Q4: How do you select alpha?

Define the exact score normalization and use development queries to compare a bounded range. Validate on separate data and measure category-level results. An adaptive alpha needs evaluation of the router as well as the retrieval system.

Q5: Does SPLADE replace a vector database?

It is a representation model, not a complete storage service. It produces sparse weights that an appropriate search engine can index. The choice of storage, updates, access controls and hybrid scoring remains.

Q6: Why is a reranker insufficient when retrieval misses the answer?

It only scores candidates it receives unless the system performs another search. Increasing reranker quality cannot reorder a missing document into existence. Diagnose candidate coverage before tuning downstream ranking.

Q7: What do you do when one search arm times out?

Follow a measured degradation policy. If the remaining arm meets the task's minimum requirements, serve that result with correct telemetry; otherwise fail or request another attempt. Preserve permission checks and do not cache a partial result as a fully evaluated one.

Final notes

Recall card: Complementary signals → compatible IDs → explicit fusion → evidence evaluation. Hybrid search earns its complexity only when it improves the complete task under real operating constraints.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Vector Databases and Search Indexes
NEXT LESSONReranking Strategies →

Explore the diagram