Documents
Documents are the long-form knowledge base your AI uses to answer user questions. Anything you add — terms of service, FAQ, product catalog descriptions, internal SOPs — becomes searchable context the AI can ground its answers in.
Creating a Document
Click + Add Document and pick one of three methods:
| Method | What it does |
|---|---|
| Manual entry | Write the content directly in the editor. Best for short, hand-curated text. |
| Upload file | Convert a PDF, DOCX, or HTML file (max 10 MB) into a document. |
| Generate with AI | Give the AI a source (pasted text or an uploaded file) and let it produce a draft document — see AI Document Generation below. |
For Manual entry and Upload file, every document also needs:
| Field | What it does |
|---|---|
| Title | Required. Shown in the document list and used in retrieval context. |
| Type | Required — see Document Types below. |
| Category | Required. One of Navigation, Features, Rules, Payment, Account, Support, General — organizational, used for filtering the document list. |
| Keywords | Required for manual entry (comma-separated terms that help the AI and the dashboard search match this document); optional when uploading a file. |
| Content | Required for manual entry. For an uploaded file, this is the extracted text — while extracting, the dashboard shows “Uploading file…” → “Extracting text…”, and for a PDF specifically, also “Formatting content with AI…” (the raw extraction is cleaned up before you see it). |
Manually-created and directly-uploaded documents are published immediately — they skip the draft phase and are indexed for retrieval as soon as you save.
Document Types
Every document carries a type that tells the AI how to use the content. Picking the right type is more important than it looks: the AI follows different rules per type, so a contract loaded as Information is paraphrased while the same contract loaded as Rule / Policy is delivered verbatim.
| Type | What it tells the AI to do | Use it for |
|---|---|---|
| Information (default) | Use the content naturally to form an answer; rephrase and summarize while preserving accuracy | Reference material, product descriptions, narrative explanations |
| FAQ | Match the user’s question to the closest entry and deliver that answer directly | Question-answer content with discrete entries — see FAQ Items below for the special structure |
| Guide | Preserve step order; never skip steps; if asked about a specific step, continue forward from there | How-to procedures, onboarding flows, troubleshooting trees |
| Rule / Policy | Deliver exactly as written — don’t soften, shorten, add commentary, or rephrase | Legal terms, official policies, regulatory text where the wording matters |
The type is a strong signal — the wrong type leads to the wrong tone (a terms-of-service document loaded as Information gets paraphrased into something legal didn’t approve; a chatty FAQ loaded as Rule / Policy reads like a manifesto). If a single document mixes content kinds, split it into multiple typed documents.
Rule / Policy documents have an extra setting: Preserve Original Language — the AI delivers this content in its original language, without translating it.
FAQ Items
FAQ documents have a special structure: instead of free text, you add FAQ items — explicit question/answer pairs. Each item also accepts question variants, alternate phrasings of the same question (e.g. “how do I cancel?”, “I want to cancel my order”). Variants are embedded individually so semantic search finds them regardless of how the user asks.
An Answer Style setting applies to the whole FAQ document:
| Setting | What it does |
|---|---|
| Exact Delivery | The AI delivers the FAQ answer as-is, with no additions or modifications. |
| Natural Integration | The AI uses the FAQ answer as reference and integrates it naturally into the conversation. |
This structure typically gives sharper retrieval than dumping Q&A pairs into a long Information document, because each entry is its own embedding rather than a sliver of a chunked text.
AI Document Generation
For sources you don’t want to type by hand — long PDFs, contracts, scraped web pages, chat transcripts — the dashboard can generate the document for you in two output modes:
- FAQ mode — extract the source into structured question/answer pairs, saved as a FAQ document.
- Summary mode — produce a clean narrative summary that drops marketing fluff, saved as an Information document.
Provide a source (paste text, or upload a PDF, DOCX, HTML, or plain text file up to 10 MB), pick the mode, optionally add a focus instruction and a target output language (leave empty to match the source language). As you type or pick a file, the dashboard shows a live estimate of the tokens the source will cost against your plan; click Analyze to run generation — extraction, analysis and generation happen server-side in one request, and the result lands as a draft for review.
Generation Limits
- Source size — how large a source you can submit depends on your plan; the live estimate shown while you type or upload a file flags it before you submit if the source is too large for your plan. See Billing.
- Generated output is capped internally per generation. On a very long source, generation can fail outright (FAQ mode) or come back as incomplete text (Summary mode) if the output hits that cap. If that happens, split the source into a few smaller generations rather than one large one.
If your source is too large, trim it, split it into multiple generations, or upload the file directly as a regular document if you don’t need AI restructuring.
Draft → Publish Lifecycle
AI-generated documents land as drafts. Drafts are visible in the dashboard but not indexed for retrieval — the AI can’t see them yet. This is deliberate: AI output needs review before it shapes user-facing answers. A banner on the draft explains this and offers a Publish button.
Publishing a draft activates it and triggers embedding generation. Once embedded, it joins the retrieval pool for the next user question.
You can edit a draft before publishing freely. Editing a published document re-embeds the changed content automatically.
Activate / Deactivate
Once published, a document can be toggled Active or Deactivated without deleting it. Deactivating removes it from the retrieval pool immediately — use it to temporarily pull a document out of circulation (e.g. content under review) without losing the content or its history.
Editing Content — Markdown Edit / Preview
The document editor’s content field has two tabs: Edit (raw Markdown) and Preview (rendered). Switch to Preview to check formatting before saving.
Chunking
Long documents are auto-chunked at upload so each piece fits in the embedding model’s context window. Chunks share the same parent document but each carries its own embedding, and retrieval surfaces the relevant chunk rather than the whole parent.
Use View Chunks on a document to inspect how it was split: chunk count, original word count, average chunk size, the embedding model used, and — per chunk — its position, character count, overlap with adjacent chunks, and the raw embedding vector.
You don’t configure chunking — it happens automatically. The implication is that very long documents can have multiple chunks ranking against each other in retrieval; if you find one chunk consistently misleading the AI, splitting the source into smaller, focused documents gives you finer control.
Tool Knowledge Documents
A document can be marked tool knowledge. Tool-knowledge documents are never used for general answers — they reach the AI only when a tool that links them is selected for the current turn. Use this for reference material that only makes sense in the context of a specific tool (field dictionaries, formatting rules, edge cases) rather than as general grounding.
Linking happens from the tool’s editor, under Knowledge Documents — see Tools › Knowledge Documents.
Sub-Project Documents
Each sub-project has its own Documents page, managed the same way as the parent project’s. Sub-project documents are scoped to that sub-project only.
How Documents Differ from Critical Instructions
Both shape what the AI says, but they work differently:
- Critical Instructions apply to every message. They cost tokens on every turn and are best for short, always-on rules (“respond in the user’s language”, “keep replies under 3 sentences”).
- Documents are retrieved on-demand based on the user’s question. The AI only “sees” the documents relevant to the current query, so size doesn’t directly cost tokens on unrelated turns. Best for content the AI doesn’t need on every message.
If you find yourself repeating policy text in Critical Instructions, it usually belongs in a Rule / Policy document instead.
Tuning Retrieval Quality
The AI’s answer quality is roughly bounded by the quality of the documents it retrieves. Three levers, in order of impact:
- Pick the right type. A Rule / Policy document forces verbatim delivery; an Information one allows paraphrasing. Picking wrong is a bigger quality hit than any other lever.
- Fewer, sharper documents. Many overlapping documents create retrieval ambiguity and dilute the relevant signal. Consolidating duplicate content into one canonical document usually beats adding more.
- Use FAQ items, not long FAQ narratives. Discrete Q&A entries with variants embed and retrieve better than a wall of “Q: … A: …” text in an Information document.