Back to releases
v0.7.2

Multi-model and rerank retrieval (v0.7.2)

A model catalog, multi-capability providers, and a rerank loop ship, so results after hybrid recall land closer to the real question.

Multi-modelRerankHybrid retrieval

Langhuan separates model connections from model capabilities, letting different providers offer embedding, rerank, and other abilities, and runs a rerank pass after hybrid recall. Retrieval is no longer bound to a single model or vendor, and connection config no longer dictates how business code is written.

One provider, many capabilities

The same model service may offer embedding, rerank, or future capabilities at once. This version splits “where to connect” from “what the connection can do”:

  • The model catalog shows the capabilities a provider actually offers, not just a connection name.
  • At runtime, provider descriptors are built from registered capabilities; missing capabilities fail early with a stable error.
  • Connection and model config save a snapshot, so later config changes don’t silently alter the behavior of knowledge bases already created.

Judge once more after hybrid recall

Vector recall and PostgreSQL full-text search are good at quickly widening the candidate set; rerank is responsible for re-judging semantic relevance within that set. The retrieval path can now be expressed as:

vector + full-text → RRF hybrid recall → rerank → parent context

The lifecycle, config, and connection management of the rerank model can all be done in the Web Console, and the search API and MCP tools share the same policy contract.

Easier to locate result differences

The retrieval stage logs structured rerank and recall events, helping developers tell apart “nothing was recalled” from “recalled but dropped by rerank.” For debugging RAG, that’s more useful than seeing only the final answer.