Your knowledge base in the app
The knowledge base is the foundation of every answer: this is everything your agent knows. The Knowledge base module shows that knowledge as a list of knowledge items, with the source, the number of text blocks and the status per item; you filter the list by type, segment, health or origin. So you see not only what your agent knows, but also whether that knowledge is in good shape.
Sources and documents
A source is a place knowledge comes from: a website the crawler reads, an API connection through which your own system delivers documents, or an RSS feed. Every source produces documents: one per page or item, split into text blocks the agent cites from.
Every document carries its status. Processed means ready for use; Skipped means the crawler deliberately left the page alone, for example because of a robots rule, a 404 or an unsupported document type. The status always tells you why, so you can fix things precisely. How the crawler reads and chooses is covered in The crawler and your content.
If a source is still reading, you see that too. It shows how many documents have been processed and how many are still queued, and that number climbs while you watch. No need to refresh.
Click a document open and you see everything up close: the text blocks it was split into, the metadata, the crawl information, and how often the document was cited in answers and when it was cited last. So you know not just that a page is in the knowledge base, but whether it is actually being used.
At the bottom of the page you manage the sources themselves: per source you see the type, the starting point and the number of documents, and Add new source connects another one. As soon as a source is updated, vragen.ai processes the changes automatically.
That button opens a short wizard: you pick the type (a website we crawl, an API you send documents to yourself, or an RSS feed), enter the starting point, decide which parts of the site take part and set right away which metadata is collected.
Metadata and smart collections
Per source you decide which metadata is collected, such as an author or publication date; new fields appear automatically as they show up in your content. Those fields are not bookkeeping for its own sake: a search strategy can filter and weigh on them, so that recent documents, for example, get priority in the answers.
Smart collections gather documents automatically based on a description. vragen.ai sets them up for you; ask about it if you want to manage a group of documents as a whole. After that you filter on them easily in the knowledge base and a search strategy can search them specifically.
Watching the health
Per source you see how the content is doing: healthy, skipped, thin content (pages with too little text to build a good answer from) or empty and failed. Documents that really need attention are marked separately. The same health shows up compactly on the dashboard; this is where the tools live to do something about it.
Segments: different reading instructions per part of your site
Large sites are rarely built the same way throughout. With segments you give parts of a source their own reading instructions, based on URL patterns: the news archive differently from the product pages. With Test on this page you try those settings on a single URL first; you see exactly what the crawler extracts, without anything changing in your knowledge base.
Adjusting course
Three buttons do most of the work. A document that does not belong in answers (an outdated page, an internal notice) you disable; it no longer counts, without you having to take the page off your site. Just rewrote a page? Have that one document reindexed and the change is known immediately. And with a new source, Crawl now fetches the first pages right away so you see results immediately; the full crawl then runs automatically every day (with an API connection you deliver documents yourself, so there is no crawl button). On top of that, vragen.ai periodically checks whether existing pages still exist.
Starting over completely is possible too. Clear knowledge base removes every knowledge item with its text blocks, metadata and knowledge areas; your sources stay, so the next indexing run refills the knowledge base. Useful after a big rebuild of your site, but it cannot be undone. That is why the button sits behind a confirmation and is available to the Owner role only.
Who can do what
Everyone from the Annotator role up can view the knowledge items. Adding sources, disabling documents and reindexing starts at the Editor role. More in Managing users and roles.
What you see in the knowledge base is ultimately steered by your content itself. Practical guidelines live in Managing knowledge base content and Best practices.




