v1.3.0 fills a long-missing piece of Langhuan’s content matrix: turning websites into knowledge-base content sources. The design ruling is deliberate — a website is not a new knowledge-base type, but a manageable source inside the KB: crawled pages live in the same knowledge base and index as uploaded files, and a single search hits both. Each source carries its own configuration, schedule and lifecycle — re-run it, reconfigure it, disable or delete it independently — and every page document records which source it came from.
Three ingestion modes for three real needs
- Crawl: a bounded BFS from a seed URL — depth/page ceilings, include/exclude path globs, subdomain toggle; discovery and fetching share one pipeline with a per-run page cache to avoid double traffic;
- Sitemap: parses sitemap.xml (including index recursion);
lastmodfeeds the incremental cursor directly; - URL list: a hand-maintained page set — add/remove triggers a sync — for competitor pages or help centers where you know exactly which pages you want.
Four fetch engines + auto escalation
http: SSRF-safe static fetching (dial-time IP checks, per-hop redirect re-validation, private-network blocking) with conditional requests;render: connects to an external CDP headless browser (Obscura and browserless both verified) and captures the rendered DOM — no browser is bundled, the single-binary delivery stands;firecrawl/jina: managed providers with markdown pass-through; self-hosted firecrawl is supported (endpoint overridable, verified);auto: starts on http and escalates to the configured engine when a JS-shell page is detected (empty mount points, low visible text), recording the actual engine inlast_engine.
Extraction runs readability + html-to-markdown on the http/render channel; provider channels pass markdown straight through.
Long crawls never crowd out everyday search
Crawling a site can take minutes to hours. The system automatically limits how many sites crawl at the same time (2 by default), so parsing and search for uploaded documents always have resources to spare; crawl timeouts scale with site size (large sites get up to 4 hours), so big crawls don’t die in a timeout-retry loop. robots.txt is honored by default, with serial same-site fetching and a configurable delay.
Re-crawling is equally frugal: each sync only processes what actually changed — sitemap timestamps, HTTP “not modified” responses, and content fingerprints filter out unchanged pages before anything is re-processed.
Console and API
The content area gains a Web sub-tab: the crawl-request list (health badge, page/failure counts, last-sync summary) → source detail (pages fetched by that request, filtered per source) → add/edit forms (mode cards, advanced options, engine-connection linkage). The integrations page supports three new connection types (firecrawl/jina/browser, credentials encrypted with AES-256-GCM) with one-click connection tests. On the API side, a full /web-sources REST surface lands, and MCP’s document_ingest accepts a url — agents can now drop a URL straight into a knowledge base.
Also fixed: a sync bug affecting older releases
Real-environment verification uncovered that sync-created parse tasks were missing index_generation_id, meaning Feishu sync in every release since v0.7.4 suffered “documents crawl fine but never enter the index” (pages and reconfigurations alike). v1.3.0 fixes it across all three paths (add/update/retry) and adds task-lineage integration tests. If you use Feishu sync on an older release, upgrade as soon as you can.
Verification
go test ./..., make test-integration (ephemeral pgvector containers) and the Web Console pnpm check / test / build all pass; a full-stack E2E smoke (Playwright + standalone mode) covers the user flow “create source → sync → pages visible → search hits”; firecrawl (self-hosted instance) and Obscura (CDP rendering) were verified against real providers end to end. The retrieval eval is not applicable to this release (no retrieval-algorithm changes).