← Back to blog
David16 min read

We Benchmarked Our Own RAG Retrieval — and the First Smoke Run Caught Two Production Bugs

The metric that matters most for a knowledge base product is retrieval quality, and all we had were functional tests. How Langhuan built a retrieval eval with 200 human-annotated real queries, a dual-track channel matrix, and bit-identical reproducibility — plus the production bugs it caught on day one and the hybrid-search hypothesis it finally confirmed with data.

RAG EvaluationRetrievalHybrid SearchKnowledge Base

After we shipped Langhuan v1.1.0, I took stock of what we had and found something a little embarrassing: retrieval is the core of this product, and not a single test in our arsenal could answer “is the retrieval any good?”

Unit tests verified that functions behave. Integration tests verified that pipelines connect. The chunking contract had gone through three revisions, RRF fusion was in, rerank was wired up — every architectural decision had a rationale, and none of them had ever been validated by a number. We had built a car whose dashboard only says “engine running,” with no speedometer.

So this cycle we built the speedometer. This post is about how, and what happened along the way. Spoiler: the very first smoke run caught two production bugs that real users would have hit.

A ruler that doesn’t cheat

The most expensive part of an eval is never the code — it’s the labeled data. We didn’t label a single line. We used MIRACL-zh directly: a Chinese Wikipedia corpus with human-annotated passage-level relevance, Apache-2.0 licensed. From it we deterministically sampled 200 real search queries; the random seed and the dataset fingerprint go into the report, so anyone holding the same data gets the same exam.

A single score isn’t enough either, because retrieval loses quality at different stages, so the exam split into two tracks. One indexes 5,298 single-passage documents to measure embedding, full-text search, and RRF fusion on their own. The other indexes 709 full articles as real long documents to measure the whole pipeline: chunking, parent-child chunks, retrieval. The first is parts-level precision; the second is the car on an actual road — and the gap between the two rulers is itself information.

Each track runs a four-cell matrix: vector-only, FTS-only, hybrid, hybrid-plus-rerank. Metrics are the retrieval-standard recall@10, MRR, and nDCG, with no LLM involved in scoring. The uncertainty of an LLM judge is precisely what this eval exists to eliminate.

The other bottom line is determinism. Same fingerprint — same dataset, chunking parameters, models, and code version — run twice, on a different port with a fresh instance, and the metrics match bit for bit once you strip timestamps. Without that, every fluctuation you see might be noise, and “the change made it better” remains folklore.

First smoke run, two production bugs

The harness itself needed validation, so step one was a smoke run with a semantic-free mock embedding: if the harness were biased, scores should hug the random baseline. They did, to the digit. The ruler was straight.

But the smoke run does more than check the ruler: it has to spin up a real standalone instance, actually create a knowledge base, actually write to it. On that path, two long-hidden bugs surfaced.

Bug one: creating a knowledge base in standalone mode returned a 500, every time — a forward foreign key in SQLite was missing its deferred check. Bug two was sneakier: vector search wasn’t working at all, because the vec extension wasn’t linked into the production binary. Both shipped fixed in v1.1.1.

What’s interesting is why the existing tests never caught them: they never took this path. Unit tests mocked the database away; integration tests ran against PostgreSQL. The eval harness plays the role of a brand-new user — it pulls up a complete system from scratch, exactly as the docs describe, and steps on the mines that only new users step on.

Real models in, full-text search turns in a blank paper

With real bge-m3 embeddings and 200 queries done, the results table showed a row of glaring 0.0000s: the FTS channel, zero recall across the board. Not “weak on some query types” — not a single hit anywhere.

Then I looked at hybrid: 0.9799, identical to vector-only, digit for digit. Meaning the so-called hybrid search had exactly one channel doing any work, and full-text search was a decoration spinning in neutral. In that state the system throws no errors, fires no alerts, logs are green, and the search results users see “seem fine.”

The mechanism and the fix for that bug deserve their own post — the write-up is here.

The same round produced two more findings. One: swapping in Qwen3-Embedding as a control moved the needle by only 0.4 percentage points, which told us the bottleneck of retrieval quality isn’t the model, and saved us who knows how much tinkering. Two: the long-document track dropped roughly 19 points below the passage track, so the chunking pipeline leaks real quality — an account to be settled later, attribution by attribution.

After the fix, an architecture hypothesis closes for the first time

The FTS fix landed and the same-fingerprint rerun read like this: full-text search went from 0 to 0.13 — modest, but alive. Hybrid reached 0.9826, strictly above vector-only’s 0.9799 for the first time. The margin is under three in a thousand, but it’s the first time “hybrid search is worth doing” was confirmed by our own data rather than by an architecture diagram.

With the reranker on, passage-level MRR climbed to 0.9975 — hits pinned to position one, almost always. On the long-document track, hybrid-plus-rerank was also the strongest combination, recall@10 at 0.79. Incidentally, we also ran two chunking-parameter experiments; the perturbation was ±0.5 percentage points, not enough to justify changing the defaults. Tuning chunk size won’t rescue retrieval — a topic for another day.

How to read these numbers (and how not to)

One thing must be said plainly: these absolute scores cannot be compared against public leaderboards. Our passage track retrieves from a sampled pool of about five thousand passages; the official MIRACL leaderboard computes over a full pool of nearly five million. The bigger the pool, the lower the score. Taking 98% from a sampled pool and shouting “beats the leaderboard” would be self-deception.

The value of this eval is measuring yourself with the same ruler: every change gets a same-fingerprint run, a diff of two metrics.json files, and every gain or loss attributed to something concrete. It has already shifted from “validating the architecture” to daily regression — the next change to tokenization or stopword lists has to clear this gate first.

How this lands in Langhuan

The eval system is a standalone binary, outside the main code path. In the repo, make eval is one command: bring up the instance under test, import the dataset, run the four-cell matrix, emit the report. Langhuan’s recommended configuration — hybrid search on by default, workspace-level rerank, default chunking kept as is — grew out of this data rather than out of armchair judgment.

The full evaluation report is public in the GitHub repository, fingerprints and all historical runs included. Bring your own dataset and run it. To just experience the retrieval, the download page has binaries for every platform.

Further Reading

  • MIRACL dataset: multilingual retrieval evaluation with public corpora and human annotations — the data source for this eval.
  • BEIR (arXiv): the standard playbook for zero-shot retrieval evaluation — many tasks, public data, no tuning.
  • RRF paper (Cormack et al., 2009): the original paper on Reciprocal Rank Fusion, the fusion method behind Langhuan’s hybrid search.
  • BGE-M3 (arXiv): the open-source Chinese embedding model used in this eval; runs locally on Ollama.
  • Langhuan retrieval evaluation report: every run, fingerprint, and recommended configuration — the source of the facts in this post.

Related · Further reading