Docs The app

Your knowledge base in the app

The knowledge base is the foundation of every answer: this is everything your agent knows. The Knowledge base module shows that knowledge as a list of knowledge items, with the source, the number of text blocks and the status per item; you filter the list by type, segment, health or origin. So you see not only what your agent knows, but also whether that knowledge is in good shape.

The knowledge base in vragen.ai: a list of knowledge items with the last update, the source, the number of text blocks, the type and the status per item.

Sources and documents

A source is a place knowledge comes from: a website the crawler reads, an API connection through which your own system delivers documents, or an RSS feed. Every source produces documents: one per page or item, split into text blocks the agent cites from.

Every document carries its status. Processed means ready for use; Skipped means the crawler deliberately left the page alone, for example because of a robots rule, a 404 or an unsupported document type. The status always tells you why, so you can fix things precisely. How the crawler reads and chooses is covered in The crawler and your content.

If a source is still reading, you see that too. It shows how many documents have been processed and how many are still queued, and that number climbs while you watch. No need to refresh.

Click a document open and you see everything up close: the text blocks it was split into, the metadata, the crawl information, and how often the document was cited in answers and when it was cited last. So you know not just that a page is in the knowledge base, but whether it is actually being used.

A document's detail page in the knowledge base: the source, the last update and how often the document was cited at the top, the text blocks below and the information and metadata on the right.

At the bottom of the page you manage the sources themselves: per source you see the type, the starting point and the number of documents, and Add new source connects another one. As soon as a source is updated, vragen.ai processes the changes automatically.

The Sources section in the knowledge base: the type, starting point and number of documents per source, with the Add new source button.

That button opens a short wizard: you pick the type (a website we crawl, an API you send documents to yourself, or an RSS feed), enter the starting point, decide which parts of the site take part and set right away which metadata is collected.

The Create source wizard: the choice between website, API or RSS feed, with the follow-up steps for the address and which pages take part at the top.

Metadata and smart collections

Per source you decide which metadata is collected, such as an author or publication date; new fields appear automatically as they show up in your content. Those fields are not bookkeeping for its own sake: a search strategy can filter and weigh on them, so that recent documents, for example, get priority in the answers.

The Metadata section in the knowledge base: the type and the number of documents per field, with the New metadata field button.

Smart collections gather documents automatically based on a description. vragen.ai sets them up for you; ask about it if you want to manage a group of documents as a whole. After that you filter on them easily in the knowledge base and a search strategy can search them specifically.

Watching the health

Per source you see how the content is doing: healthy, skipped, thin content (pages with too little text to build a good answer from) or empty and failed. Documents that really need attention are marked separately. The same health shows up compactly on the dashboard; this is where the tools live to do something about it.

Segments: different reading instructions per part of your site

Large sites are rarely built the same way throughout. With segments you give parts of a source their own reading instructions, based on URL patterns: the news archive differently from the product pages. With Test on this page you try those settings on a single URL first; you see exactly what the crawler extracts, without anything changing in your knowledge base.

Adjusting course

Three buttons do most of the work. A document that does not belong in answers (an outdated page, an internal notice) you disable; it no longer counts, without you having to take the page off your site. Just rewrote a page? Have that one document reindexed and the change is known immediately. And with a new source, Crawl now fetches the first pages right away so you see results immediately; the full crawl then runs automatically every day (with an API connection you deliver documents yourself, so there is no crawl button). On top of that, vragen.ai periodically checks whether existing pages still exist.

Starting over completely is possible too. Clear knowledge base removes every knowledge item with its text blocks, metadata and knowledge areas; your sources stay, so the next indexing run refills the knowledge base. Useful after a big rebuild of your site, but it cannot be undone. That is why the button sits behind a confirmation and is available to the Owner role only.

Who can do what

Everyone from the Annotator role up can view the knowledge items. Adding sources, disabling documents and reindexing starts at the Editor role. More in Managing users and roles.

What you see in the knowledge base is ultimately steered by your content itself. Practical guidelines live in Managing knowledge base content and Best practices.

Still have a question?

Ask it here. The answer comes from this documentation and points to the right page.

What exactly is vragen.ai?

Example answer by vragen.ai

vragen.ai lets visitors ask their question on your website and gives them a reliable answer straight from your own content, with the source included. You decide which sources the AI uses.

Source: How it works

This is an example. The interactive widget could not load here, for instance because of a script blocker or a slow connection.

You are talking to an AI assistant from vragen.ai. Why we mention this (Dutch)