Introduces in-KB "web content sources" — three ingestion modes (crawl/sitemap/urls), four fetch engines (http/render/firecrawl/jina with auto escalation), pages indexed alongside uploaded files and retrievable in the same search; each source is independently configurable, re-runnable and deletable, backed by full-stack E2E and real-provider verification.
Changelog
Langhuan product changelog
Tracking how Langhuan evolves across knowledge bases, MCP over HTTP, RAG retrieval, PDF parsing, and the Web Console: why each release shipped, what it solves, and how it makes the knowledge infrastructure more dependable.
Four open retrieval directions closed by controlled experiments (semantic chunking, LLM context headers, chunk parameters, traditional/simplified normalization), then progressive disclosure for search results (detail full/lean) and discovery tools for agents (MCP tools 7→10); matched_children no longer carries child content.
Fixes three defects — standalone knowledge base creation returning 500, vector search unavailable in standalone, and FTS zero recall on Chinese questions — plus a disable-per-channel diagnostic semantic and an offline retrieval eval harness.
Without Redis, tasks now queue in SQLite or PostgreSQL — restarting no longer drops jobs; also fixes a standalone OIDC panic and trims the binary ~22%.
Run as a single binary with a single .db file, no PostgreSQL or Redis; coexists with the production PG path, with complete retrieval and zero PG regressions.
Every retrieval now has a stable identity, verifiable citations, and full lineage. Admins can replay a search with the original query within the retention window — no more guessing at RAG recall drift.
Failure retry, Generation rebuild, OpenTelemetry observability, and asynq task governance ship — giving Langhuan the long-term runnability and recoverability needed to enter release readiness.
Stable incremental document sync, safe increment cursors, forced re-sync, and deletion policies ship, so large collaboration spaces sync without dropping content or deleting the wrong thing.
OIDC login, JIT user creation, external-identity binding, and single-tenant auto-join ship, rounding out the enterprise onboarding path.
Sidebar hierarchy and page backgrounds are tightened and dev-only floats removed, so daily management stays focused.
Feishu app management, source sync, and a file-tree browser ship, bringing team docs into Langhuan's retrieval pipeline.
A built-in OpenAPI live doc and Scalar browser ship, so the REST management surface can be explored and verified directly.
A model catalog, multi-capability providers, and a rerank loop ship, so results after hybrid recall land closer to the real question.
Adaptive parent-child chunking and parent-block aggregation ship, so retrieval hits smaller fragments while still returning full context.
Introduces MinerU Cloud PDF parsing, S3-compatible object storage, and document asset archiving, so complex documents keep a traceable structure.
Langhuan ships under the MIT License, bringing REST, MCP over HTTP, and an embedded console together in one runnable binary.