From content to searchable knowledge
The pipeline's goal is not merely to store a document. It produces passages that can be found for a customer question and traced back to their source.
Extraction
For web pages, primary content is separated from navigation, cookie banners and repeated page chrome. Files use format-aware parsing: PDF text and page location, DOCX headings, Excel or CSV rows and plain-text encoding are preserved where possible. An unreadable or partially processed file is not marked ready.
Crawling runs in two phases: pages are fetched and stored with a checkpoint first, then duplicates are removed and the content is chunked. Of several pages carrying identical content only one is kept, the shortest address, because it is the most stable source identity. Pages marked noindex are skipped.
Pages with very little visible text fall back to browser-based rendering; otherwise that step is skipped entirely. PDFs found during the crawl are converted to text. Scanned PDFs are not read; for those documents an accessible text version is more dependable.
Cleaning and structure
Empty sections, repeated templates and elements with little retrieval value are reduced. A heading trail is attached to each passage so context such as “Returns > Conditions” survives when a paragraph is retrieved alone. Repeating column headings when a long table is split keeps values meaningful.
Chunking and embeddings
Content is not blindly cut at fixed character positions. Heading, paragraph and table boundaries help produce passages that fit the answer context. A numerical representation for semantic search is generated for each passage; original text is stored with source identity, location and version details.
A section boundary only opens a new passage once the accumulated text exceeds 250 tokens. This prevents a short introductory line under a heading from splitting an answer on its own. Measured median passage size moved from 143 to 362 tokens after this change; short intro lines no longer enter retrieval as separate passages.
Each passage's content, together with its heading trail, is written into both the numerical representation and the full-text index. When chunking rules change, the pipeline version increases; this guarantees old passages cannot be mixed with content from a new pipeline.
An embedding is not model training. It represents text so that queries with related meaning can be compared. The embedding model identity is read at call time. The original RAG paper describes the foundation of retrieving external knowledge during generation.
Updates
The previous successful version can remain active until a replacement is completely prepared and published. This prevents a failed crawl from replacing working knowledge with partial output.
Synchronisation is idempotent: if content has not changed, nothing is re-embedded. When a source becomes unreachable, automatic refresh is held back for 24 hours so it cannot keep producing failing jobs and cost.
Retest changed questions after synchronisation: a ready source alone does not prove that an answer meets the business need.