Finding the right evidence
No single retrieval method performs well for every question. Semantic search can connect “refund” with “can I get my money back?”, while lexical search is strong for a product code, district name or exact policy term. Dolphy uses both rankings.
Two retrieval branches
The vector branch compares the question embedding with passage embeddings. The lexical branch searches meaningful query terms in a full-text index, considering two text representations: a stemmed form for Turkish and a plain form.
Both branches operate within the same tenant and agent boundary; content from another business does not enter the candidate pool. Vector results are returned through an HNSW index (ef_search 120, relaxed-order scan). The query is sent as raw text and is only rewritten on follow-up questions.
Calibrated fusion
The two branches produce scores on different scales, so adding them directly would mislead. The score is computed as:
score = 0.8 × semantic + 0.2 × lexical
The semantic score is scaled against the lowest similarity in the candidate set rather than a fixed base, because every question produces a different candidate set. The lexical score is the fraction of the query's roots present in the candidate set, weighted by inverse document frequency within that set. This keeps the comparison meaningful even when the candidate set changes.
Measured on the same dataset, this change produced:
| Metric | Before | After |
|---|---|---|
| MRR | 0.876 | 0.911 |
| Hit rate at rank 1 | 81% | 86% |
| Hit rate within top 8 | 97% | 99% |
Separately, once the retrieval path was reorganised from two separate database calls into a single calibrated fusion, warm response time fell from 30.4 ms to 4.4 ms and cold response time from 2,806 ms to 259 ms.
Ranking is deterministic end to end: the same query and data produce the same fused order.
RRF is still available
Reciprocal Rank Fusion builds the ranking from position instead of raw score and remains an option, but the production default is calibrated summation. RRF was kept as an alternative because its measured gain did not justify the added latency and external dependency.
Source selection
Exact duplicate passages are removed after fusion. Dolphy keeps enough depth from the strongest source while limiting the ability of one long page to fill the entire answer context, then adds passages from other sources. If limits leave the context short, later candidates fill the remaining slots rather than leaving them empty.
About reranking
Reranking means applying a second model or scoring step to reorder the first candidate set. Dolphy does not claim that its current pipeline uses cross-encoder reranking. Measurements did not show a sufficiently stable benefit for the added latency and dependency, so production ordering uses calibrated fusion. The decision can be measured again on new, broader evaluation sets.