# How agents use knowledge

Loading a knowledge base does not make an agent use it. Two things have to be true: the
agent has to have it *selected*, and the content has to reach the model. These are
separate mechanisms, and most reports of "the agent ignored my knowledge" come down to
one of them rather than to the model.

## Two paths into the prompt

The type of a knowledge base decides how its content travels:

| Type | How the model sees it |
|---|---|
| `form`, `singlePage` | **Inlined.** Every record is written into the system prompt, on every request. |
| `files` | **Retrieved.** Only the passages most similar to the current message are pulled in. |

Inline knowledge is always considered and takes up room in the prompt on every message.
Retrieved
knowledge scales to a document library but is present only when the question resembles
it.

That asymmetry drives the design decision. A twelve-line refund policy that must never be
missed belongs inline. Three hundred PDFs cannot go inline, because there is not room, so
they go in a files base and are retrieved.

A files-type base contributes its **name and description only**, as a pointer telling the
model that a library exists and what is in it; the passages themselves arrive separately,
per message. This is why `ds_description` matters: it is the model's only guide to when a
knowledge base is worth drawing on.

`GET /agents/{id}/preview-system-prompt` shows what a specific agent resolved, without
running anything. Check it first when an agent behaves oddly: if a knowledge base is not in
the output, the problem is selection rather than the model.

## How selection resolves

Two fields, and they are not alternatives: one is a default and the other an override.

| Field | Contains | Set by |
|---|---|---|
| `agent_ds` | Knowledge base **names**, kebab-case | The agent's own definition |
| `agent_tenant_ds` | Knowledge base **ids** | Your workspace, layered on top |

```bash
curl -X PATCH https://api.genuineai.app/api/v1/agents/<agent-id>/ds \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{"agent_tenant_ds": ["3f2a9c1e-…", "8a1b6e5c-…"]}'
```

When `agent_tenant_ds` is set, it takes precedence outright and `agent_ds` is not
consulted. Send `null` to drop the override and fall back to the agent's own set.

<Callout type="caution">
**An empty array is not the same as `null`.** `agent_tenant_ds: []` means "this workspace
selects no knowledge", and the agent runs with none at all. `null` means "use the agent's
default." Sending `[]` when you meant to clear the override is the most common way to
silently strip an agent of everything it knew.
</Callout>

### Why names are matched loosely

`agent_ds` holds names in kebab-case, but `ds_name` accepts only letters, numbers and
spaces. They meet in the middle: each `agent_ds` entry is lowercased and its hyphens
become spaces, then matched against the lowercased `ds_name`.

```
agent_ds: "product-catalog"   →   ds_name: "Product catalog"   ✓
```

The naming rule on knowledge bases exists to keep this mapping unambiguous. If an agent's
knowledge silently does not appear, check this first: a base named `Product Catalog 2026`
cannot be referenced as `product-catalog`, and a rename breaks the link with no error
anywhere.

Selecting by id through `agent_tenant_ds` avoids the issue entirely, and is what an
integration should prefer.

Only **active** knowledge bases resolve. Deactivating one silently removes it from every
agent that referenced it.

## How retrieval works

For files-type knowledge, each message triggers a similarity search, and the passages that
come back are added to the model's context for that message, marked as retrieved reference
material rather than as part of the conversation.

Documents are split into overlapping chunks, on paragraph breaks where possible and on
smaller boundaries when a paragraph runs long. The overlap exists so that a fact straddling
a chunk boundary survives whole in at least one chunk. That is why chunks are not a clean
paragraph split.

The mechanics are fixed rather than configurable per request. A small number of passages is
retrieved per message, and passages that are not similar enough are dropped rather than
included to fill the quota.

That budget is the constraint to design against. A question whose answer is spread across
forty pages will not be answered well by retrieval, which is a sign the knowledge should be
restructured into a knowledge template, where it is all present at once.

<Callout type="note">
**Retrieval fails soft.** If the similarity search errors, the conversation continues
without knowledge rather than returning an error. The agent still answers, but from the
model alone. An answer that looks as though the knowledge base does not exist is
consistent with retrieval having failed, not only with nothing matching.
</Callout>

## Attach knowledge to one message

Knowledge selection is agent-level, but a single message can point at specific files. A
`reference` part in `msg_content` names a file to retrieve against for that message only:

```json
{
  "msg_content": [
    { "type": "text", "text": "Does this contradict our warranty policy?" },
    { "type": "reference", "file_id": "8a1b6e5c-…" }
  ]
}
```

A `reference` is retrieved against, while a `file` part attaches the document whole. Use a
reference for something long, and a file for something the model should read end to end.

Threads carry this too: `thread_ref` scopes retrieval for a whole conversation, and spaces
contribute their own knowledge to every thread inside them.

## Why an agent ignores your knowledge

Work down this list, which is roughly ordered by how often each is the cause.

| Check | |
|---|---|
| **Is it selected?** | `GET /agents/{id}/preview-system-prompt`. Not in the output means not selected. |
| **Is `agent_tenant_ds` an empty array?** | That means "no knowledge", not "default". |
| **Do the names match?** | `agent_ds` kebab-case must map onto `ds_name` exactly once hyphens become spaces. |
| **Is the knowledge base active?** | Inactive bases resolve to nothing. |
| **For files: is the file `ready`?** | Text extraction and embedding happen after finalize. Documents mid-processing are invisible to retrieval. |
| **For files: is anything close enough?** | Below 0.2 similarity nothing is returned. Phrasing that shares no vocabulary with the document retrieves nothing. |
| **Is the description specific?** | The model decides relevance from `ds_description`, and "Misc data" gives it nothing to work with. |

The first check needs `agent:edit`. For a platform-published agent it returns only the parts
your own workspace contributed, with the base prompt withheld. The knowledge section is one
of those parts, so it still answers this question.

## Knowledge outside chat

Two uses of knowledge bases have nothing to do with agents answering questions:

- **Analysis vocabulary.** The Media Tags and Flags knowledge bases are the closed
  vocabulary [image analysis](/ai-enrichment#your-controlled-vocabulary) may use. Editing
  them changes how every future upload is described.
- **Brand identity.** A base with `ds_category: "brand-identity"` feeds design and brand
  tooling rather than only chat.

A knowledge base is closer to shared organizational data than to a chatbot feature, which
is why it is its own resource rather than part of the agent.

## Next steps

- [Knowledge bases](/knowledge-bases) covers creating and filling the three kinds.
- [Agents and conversations](/agents-and-conversations) covers the agent that consumes it.
- [AI enrichment](/ai-enrichment) covers vocabulary-driven image analysis.
