Retrieval architecture

Why pure vector search fails on product codes, tickers, and other exact-match queries, and how hybrid retrieval with BM25, reciprocal rank fusion, and reranking fixes it in production RAG.

By HEU.AI ·

Hybrid Search vs Pure Vector Search: An Overview

Vector search made semantic retrieval possible at scale: search by meaning instead of exact keyword match. But as RAG systems move from prototype to production, a pattern shows up repeatedly: pure vector search alone starts failing on exactly the queries that matter most in an enterprise setting. This article looks at why that happens, how hybrid retrieval fixes it, and what a production-ready hybrid pipeline actually looks like.

Understanding Enterprise Search Solutions

Importance of Search in Enterprises

In an enterprise RAG system, search isn't a feature. It's the foundation everything else depends on. A generation model is only as good as what it's given to work with. If retrieval surfaces the wrong document, or misses the right one, no amount of prompt engineering fixes the resulting answer. This is doubly true at enterprise scale, where document volume, format inconsistency, and the cost of a wrong answer are all higher than in a demo.

Key Components of Effective Search

Effective enterprise search needs to handle two different kinds of queries well: conceptual questions ("how do we handle a customer complaint about billing") and precise, exact-match lookups ("order #48291" or "invoice INV-2024-0087"). Most systems are tuned for one and quietly fail at the other, which is the core problem this article addresses.

Definitions and Distinctions

What is Hybrid Search?

Hybrid search combines two retrieval methods that work differently: dense vector search, which finds semantically similar content even when the wording doesn't match, and sparse lexical search (typically BM25), which finds exact keyword and term matches. Results from both are combined (often via a technique called reciprocal rank fusion) so the final result set benefits from both kinds of matching.

What is Pure Vector Search?

Pure vector search relies entirely on dense embeddings: text is converted into a numerical vector representing its meaning, and retrieval finds the closest vectors to the query. It's excellent at conceptual, paraphrased, or loosely-worded queries, but it has no built-in mechanism for exact term matching, which becomes a real liability at scale.

Where Pure Vector Search Breaks Down

Exact Match Failures

Embedding models are trained to capture meaning, not to preserve exact strings. This means a pure vector search system can genuinely fail to retrieve a document containing the exact term a user searched for, if that term's surrounding context pulls its embedding away from the query's embedding. The system isn't broken. It's doing exactly what it was designed to do. It just wasn't designed for this kind of query.

Product Codes, Financial Tickers, and Other Exact-Match Data

This failure mode shows up hardest on data that is inherently exact: product SKUs, part numbers, financial tickers, invoice numbers, legal citation formats. A query for "AAPL Q3 earnings" needs to match the literal string "AAPL," not a semantically similar-sounding term. Pure vector search has no reliable guarantee of surfacing that exact match, because nothing in its architecture prioritizes exact strings over semantic proximity.

How Hybrid Retrieval Works

Combining Dense Semantic Search and Sparse Lexical Search (BM25)

Hybrid retrieval runs both searches in parallel on the same query: a dense vector search for semantic relevance, and a sparse BM25 search (a well-established, term-frequency-based ranking algorithm) for exact and near-exact term matches. Neither replaces the other: they cover each other's blind spots. BM25 catches the product code that vector search missed; vector search catches the paraphrased question that BM25's literal matching missed.

How the Two Signals Are Merged

The two result sets are typically merged using reciprocal rank fusion, a method that combines rankings from multiple retrieval systems without requiring their scores to be on the same scale. Rather than trying to normalize a BM25 relevance score against a cosine similarity score (two fundamentally different measurements), reciprocal rank fusion works from each system's rank order, which is a more robust way to combine genuinely different scoring systems.

A minimal version of this merge looks like this:

def reciprocal_rank_fusion(result_lists, k=60):
    scores = {}
    for results in result_lists:
        for rank, doc_id in enumerate(results, start=1):
            scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)

bm25_results = ["doc_7", "doc_2", "doc_9"]
vector_results = ["doc_2", "doc_5", "doc_7"]
merged = reciprocal_rank_fusion([bm25_results, vector_results])

Each document earns a score of 1 divided by (k plus its rank) in every list it appears in, and the scores are added together. A document found by both systems, like doc_2 and doc_7 here, outranks one found by only one. The constant k (60 is a common default) softens the gap between the top few ranks.

Comparing Indexing Strategies

Traditional Indexing vs. Vector Indexing

Traditional (lexical) indexing builds an inverted index (mapping terms to the documents containing them), which is fast and exact but blind to meaning. Vector indexing stores embeddings and enables approximate nearest-neighbor search, which is powerful for semantic similarity but doesn't inherently support exact filtering.

Hybrid Indexing Models

A hybrid index maintains both structures side by side, or increasingly, a single system capable of both (many modern vector databases now support hybrid indexing natively, rather than requiring two separate systems bolted together).

Information Retrieval Methods

Retrieval Techniques Overview

Beyond the dense/sparse split, retrieval quality also depends on chunking strategy, metadata filtering, and query preprocessing, but the dense-vs-sparse distinction remains the single largest lever for fixing exact-match failures in an existing RAG system.

Strengths and Limitations of Each Approach

Dense retrieval's strength is semantic flexibility; its limitation is unreliable exact matching. Sparse retrieval's strength is precise term matching; its limitation is no understanding of meaning, synonyms, or paraphrasing. Hybrid retrieval's limitation is added system complexity: running and merging two retrieval systems instead of one.

Reranking Strategies

Cohere Rerank and BGE

After initial retrieval (whether hybrid or pure vector) returns a candidate set, a reranking model can reorder those results by relevance more precisely than the initial retrieval scores alone. Cohere's Rerank API and open-source models like BGE-reranker are commonly used for this step. Both take the query and each candidate document as a pair and produce a more accurate relevance score than the original retrieval ranking.

Cross-Encoders

Rerankers are typically built as cross-encoders: models that process the query and document together in a single pass, rather than encoding them separately (as retrieval embeddings do). This joint processing lets the model capture interactions between query and document that separate encoding misses, at the cost of being too slow to run on an entire document set, which is why reranking happens only on the smaller candidate set retrieval already narrowed down.

Latency Overhead and Trade-offs

Reranking adds a real latency cost, since it requires a model inference pass for every candidate document. In practice, this typically means reranking a shortlist of 20 to 50 candidates rather than thousands, trading a small latency increase for a meaningful relevance improvement on the results that actually reach the user.

Architectural Blueprint for a Production-Ready Hybrid Pipeline

Vector Database Options

A production hybrid pipeline needs a vector store that supports both dense and sparse retrieval, ideally natively. Options in production use today include pgvector (adds vector search to PostgreSQL, useful when a team already runs Postgres, and paired with PostgreSQL's built-in full-text search for the keyword side), Qdrant (a dedicated vector database with native hybrid search support), and AWS Bedrock Knowledge Bases (a managed option for teams already inside the AWS ecosystem, where hybrid support depends on the vector store behind it). The right choice depends on existing infrastructure and operational preferences more than raw capability. All three can support a genuine hybrid pipeline. HEU.AI's RAG development work has been built on this exact combination (pgvector, Qdrant, and AWS Bedrock Knowledge Bases), depending on what a given client's stack and hosting requirements call for.

Putting the Pipeline Together

A complete pipeline looks like: ingest and chunk documents, generate both dense embeddings and a sparse (BM25) index, run both retrieval methods on each incoming query, merge results with reciprocal rank fusion, rerank the merged shortlist with a cross-encoder, then pass the final results to the generation model. Each stage is a place where a badly-tuned system loses accuracy, which is why evaluation harnesses that measure retrieval quality directly (rather than just checking whether the final answer "sounds right") matter as much as the pipeline architecture itself.

A hybrid pipeline should also show where each answer came from. When every answer carries a citation to its source document, a wrong answer can be diagnosed by looking at what was retrieved, instead of guessing. Measuring retrieval separately from generation, for example by checking whether the right document appears in the top results for a set of real test queries, shows whether a change to chunking, weighting or reranking actually helped.

Conclusion

Pure vector search's semantic strength is also its exact-match weakness, and at enterprise scale, exact-match queries (product codes, tickers, invoice numbers) are common enough that this weakness becomes a real production problem, not an edge case. Hybrid search, combining dense vector retrieval with sparse BM25 matching and merged through reciprocal rank fusion, closes that gap without giving up the semantic flexibility that made vector search valuable in the first place. Reranking adds a further precision layer on top, at a manageable latency cost when scoped to a shortlist rather than the full candidate set.

For enterprises scaling a RAG system past the prototype stage, this isn't a nice-to-have architectural refinement. It's usually the difference between a system that works in a demo and one that holds up when real users start asking real questions with real product codes and ticker symbols in them. HEU.AI's enterprise RAG development is built around hybrid retrieval, combining keyword and vector search, for US-based teams that need retrieval that's accurate at scale, not just accurate in a demo.

FAQ

Is hybrid search always better than pure vector search?

For most enterprise use cases with any exact-match data (codes, IDs, names, numbers), yes. For narrow use cases that are purely conceptual with no exact-match requirements, pure vector search can be sufficient and simpler to operate.

Does hybrid search slow down retrieval?

It adds some overhead from running two retrieval methods instead of one, but the added latency is typically small compared to the overall RAG pipeline's response time, especially compared to the generation model call itself.

What is reciprocal rank fusion?

A method for combining ranked result lists from different retrieval systems, based on each result's rank position rather than trying to reconcile incompatible relevance scores from different scoring systems.

Do I need a reranker if I already have hybrid search?

Not always, but it helps in high-stakes use cases where the precise ordering of top results matters. Hybrid search improves which documents get retrieved; reranking improves the order they're presented in.

Which vector database should I use for a hybrid pipeline?

It depends on existing infrastructure: pgvector is a natural fit for teams already on PostgreSQL, Qdrant for a dedicated vector-native solution, and AWS Bedrock Knowledge Bases for teams already inside AWS. All three can support hybrid retrieval, though with Bedrock Knowledge Bases it depends on the vector store behind it.

Scaling a RAG system past the prototype?

Book a free consultation and we will look at your documents, the queries that fail today, and the retrieval setup that fits your stack.

More Articles