Consider a query: 埃及有哪些民族? — “What ethnic groups does Egypt have?”
A perfectly ordinary Chinese question. Send it as-is to SQLite FTS5 with Chinese full-text search, and the result set is empty. Not poorly ranked — empty.
We ran into this while building Langhuan’s retrieval evaluation: 200 real queries, and the FTS channel posted zero recall across the board, a whole row of 0.0000s in the results table. My first assumption was that the eval code was broken. Tracing it end to end turned up a bug in the product itself — one that had been living comfortably in the system, because nothing ever dies from it.
The lesion: tokenization turns a question into an all-AND
The mechanism, taken apart, is not complicated.
FTS5 tokenizes the query first. Langhuan’s standalone mode uses gse for Chinese segmentation, and “埃及有哪些民族?” comes out as five tokens: 埃及, 有, 哪些, 民族, ?.
FTS5’s query semantics are AND: for a document to match, it must contain every token. Not most of them — all of them.
That’s where it breaks. A body of text about Egypt will certainly contain 埃及 and 民族. It will never contain 哪些, and certainly not the question mark. One missing token out of five makes the whole query an empty set. Question words and punctuation live in questions; they never gather inside any passage that states facts — so every query carrying a question word goes down with the ship.
Here’s the interesting part: no single component is wrong. gse segmented every word correctly. FTS5 faithfully executed AND semantics. Each looks innocent on its own; together they produce zero. That’s what makes this class of bug nasty — it doesn’t hide in a line of code, it hides in the seam between two correct components.
The bigger problem: hybrid search quietly becomes single-channel
If only FTS were dead, it would be a small affair. The real trouble is that it died inside a hybrid retrieval setup.
Langhuan’s default retrieval is vector plus full-text dual-channel recall fused with RRF — the hybrid search post covers the full picture. After the FTS channel went to zero, the fusion formula kept running, results kept coming back ten at a time, except every one of them now came from the vector channel alone. Hybrid in name, single-channel in fact.
The measured numbers show just how invisible this is: before the fix, hybrid’s recall@10 was identical to vector-only, 0.9799, digit for digit. Users saw a search that “works.” Logs showed no errors. Metrics showed no cliff. One channel of retrieval was dead, and the whole system stayed as quiet as if nothing had happened.
This is also why functional testing never caught it. Functional tests ask “does the search return results” — and it does, courtesy of the vector channel, every time. Catching this bug requires a different question: what is this channel actually contributing? We exposed it by turning the vector channel off in the eval matrix and running FTS as the only path.
The fix: strip the question’s noise on the query side
Once the thinking was straight, the fix wasn’t hard. The first decision was which side to fix.
Not the index side: the noise of question words and punctuation lives in the query, not in the documents. The right move is to filter the query before it reaches FTS — drop the punctuation outright, drop single-character function words like 有, 的, 了, drop question filler words like 哪些, 怎么, 为什么. The content words that remain go into the AND match.
“埃及有哪些民族?” filtered down to 埃及 and 民族; a passage containing both is a hit.
We deliberately kept the stopword list conservative — better to under-filter than to butcher a content word. The more aggressively you expand a stopword list, the more likely some domain term gets misclassified as a “function word,” and each misclassification silently kills a batch of keyword queries. So the rule is fixed in stone: touch the list once, rerun the eval for a before/after comparison, and let the numbers talk.
After the fix, FTS came back from 0 to 0.13. Modest-looking, until you understand where FTS earns its keep: vector search is good at “similar in meaning,” full-text search is good at “exactly these characters” — file names, model numbers, IDs, proper nouns. That class of keyword query is precisely the vector channel’s blind spot and FTS’s home turf. Questions get backed by vectors, keywords get backed by FTS; only with both channels alive does hybrid retrieval have a reason to exist. Post-fix, hybrid also came out strictly above vector-only for the first time, 0.9826 against 0.9799.
The same filtering ships for both dialects — SQLite FTS5 and PostgreSQL with zhparser tokenization plus plainto_tsquery; the full PG path is in the pgvector Chinese hybrid search post. The fix shipped in v1.1.1.
Is your FTS channel still alive?
If your system is also hybrid retrieval, it’s worth ten minutes to check whether the full-text path is actually working.
The most direct method is single-channel probing: find a way to turn the vector channel off, leave only full-text search, and try a few everyday questions plus a few keyword queries. Questions returning nothing isn’t alarming; FTS is naturally bad at questions. Keywords returning nothing is when you should worry.
One step further, look at the contribution split: how many results in the fused output come from FTS recall. If that number sits at zero for a long while and you’re sure the corpus contains exact keywords, you’re almost certainly running a channel on idle too.
Langhuan made this a first-class capability: the retrieval channel’s top_k parameters accept 0, where vector_top_k=0 runs full-text only and keyword_top_k=0 runs vector only — channel-level diagnosis without touching code. Every number in this postmortem comes out of the retrieval eval system we just built; the full story is here.
To try it against your own corpus, the download page has binaries for every platform.
Further Reading
- SQLite FTS5 documentation: query syntax and matching semantics of FTS5, including the official wording of AND semantics.
- gse: the Go Chinese segmentation library used by Langhuan’s standalone mode.
- PostgreSQL text search controls: lexing and matching rules of plainto_tsquery, the PG-side counterpart.
- zhparser: the SCWS-based Chinese parser extension for PostgreSQL, the production-side companion.