← Back to blog
David8 min read

Chinese hybrid search with pgvector + PostgreSQL full-text

Store vectors in pgvector and tokenize Chinese with zhparser for full-text search, then fuse with RRF on a single SQL stack without a separate vector database.

pgvectorFull-text searchPostgreSQL

A common fork in the road when building Chinese retrieval for a knowledge base is this: one component for vector search, another for full-text, and then glue to merge the results. But inside PostgreSQL, pgvector plus a Chinese tokenizer gets you both paths on a single stack.

This article explains that path and why it’s often less work than “vector DB plus search engine.”

The Chinese full-text problem

PostgreSQL’s built-in full-text search is complete, but its default parser splits on whitespace and punctuation. Chinese has no spaces, so a whole sentence becomes one token and the inverted index is useless.

The fix is a parser that understands Chinese. zhparser is a PostgreSQL extension for Chinese full-text search, built on SCWS (Simple Chinese Word Segmentation). Install it, configure the parser, and “Chinese hybrid search” gets tokenized into terms like “Chinese / hybrid / search” instead of one blob.

pgvector handles the vector path

pgvector adds vector types and similarity indexes to PostgreSQL. It supports types like vector and halfvec (half precision, roughly half the storage), with HNSW and IVFFlat indexes.

For a knowledge base, halfvec plus HNSW is a common combination: recall quality stays essentially the same, storage drops noticeably, and the vectors live in the same transaction as the business data, so there’s no cross-system sync.

One table holds both vector and full-text

The key is landing both paths on the same table, same row. Each row stores:

  • a halfvec column for vector search, with an HNSW index;
  • a tsvector column for full-text search, generated with a zhparser configuration.

At query time you run both paths, get two ranked result sets, and merge them with RRF. RRF is deterministic fusion (1 / (k + rank)), with no training and no weight tuning.

Why PostgreSQL instead of a separate component

My judgment is this: when the knowledge base already treats PostgreSQL as its source of truth, keeping vectors and full-text in the same database beats “pgvector plus a standalone Elasticsearch or vector database.” Three reasons:

  1. One transaction: documents, versions, vectors, and full-text indexes update consistently in one store, with no cross-system consistency problem.
  2. One thing to operate: one fewer component to deploy, monitor, and upgrade.
  3. Traceability: a retrieved chunk ties directly to the source document, revision, and position on the same row, so troubleshooting doesn’t require stitching evidence across systems.

This isn’t to say dedicated vector databases are wrong. At very large scale, or with very strict retrieval latency, they have their place. But for most knowledge bases, the single-database approach removes complexity first, which is often enough.

Where it lands

Langhuan follows exactly this path: pgvector (halfvec + HNSW) plus PostgreSQL full-text search (zhparser) with deterministic RRF, and every chunk carries a source anchor. The schema, indexes, and retrieval chain are verifiable in the architecture docs of the Langhuan repository.

Since v1.0.0, Langhuan also ships a zero-config standalone path: SQLite + sqlite-vec for vectors + FTS5 full-text + gse Chinese tokenization, with no PostgreSQL at all, for local development and demos. The single-database idea above still holds — the database just becomes SQLite.

If you’re already on PostgreSQL, Chinese hybrid search may not require a new component at all — check whether pgvector plus zhparser solves it first.