# Knowledge bases

A knowledge base is what an agent knows that the model does not: product facts, brand
rules, policies, and the documents your team works from.

There are three kinds, and picking the right one is most of the work. They are filled
differently, and an agent uses them differently.

| `ds_type` | Called | Holds | An agent |
|---|---|---|---|
| `form` | Knowledge template | Structured records, all sharing a field schema | Reads all of it, every time |
| `singlePage` | Document | One block of text | Reads all of it, every time |
| `files` | Files | Uploaded documents | Retrieves the relevant passages |

The rule of thumb: **if the content must always be considered, use `form` or
`singlePage`; if it is too large to always be considered, use `files`.** Brand rules and
a product list belong in the first two, and a folder of 300 PDFs belongs in the third.
[How agents use knowledge](/knowledge-retrieval) explains why this matters.

## Create a knowledge base

<CodeTabs syncKey="lang">

```bash title="cURL"
curl -X POST https://api.genuineai.app/api/v1/knowledge-bases \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{
    "ds_name": "Product catalog",
    "ds_description": "Every product we sell, with positioning and pricing tier",
    "ds_type": "form"
  }'
```

```js title="JavaScript"
const kb = await fetch("https://api.genuineai.app/api/v1/knowledge-bases", {
  method: "POST",
  headers: { ...headers, "Content-Type": "application/json" },
  body: JSON.stringify({
    ds_name: "Product catalog",
    ds_description: "Every product we sell, with positioning and pricing tier",
    ds_type: "form",
  }),
}).then(r => r.json());
```

```python title="Python"
kb = requests.post(
    "https://api.genuineai.app/api/v1/knowledge-bases",
    headers=headers,
    json={
        "ds_name": "Product catalog",
        "ds_description": "Every product we sell, with positioning and pricing tier",
        "ds_type": "form",
    },
).json()
```

</CodeTabs>

<Callout type="caution">
**`ds_name` accepts letters, numbers and spaces only.** "Product catalog" is fine;
"product-catalog" and "Q3/2026 pricing" are rejected with a 422. Names must also be
unique in the workspace, because they are how agents refer to a knowledge base.
</Callout>

Two fields carry more weight than they appear to:

- **`ds_description` is read by the model.** It is not internal documentation: it tells
  the agent what this knowledge base contains and when to use it. A vague description
  produces a knowledge base the agent reaches for at the wrong moment. Write it as an
  instruction: "Every product we sell, with positioning and pricing tier."
- **`ds_category`** is `general` (the default), `brand-identity` or `studio`. Categories
  decide which parts of the platform draw on the knowledge base; brand identity feeds
  design and brand tooling rather than only chat.

Creating a `form` knowledge base without naming an existing schema creates a blank schema
for it, which you then design. Pass `ds_config.formId` instead to reuse a schema that
already exists.

## Add records to a knowledge template

Records go in one at a time, and their fields are whatever the schema declares:

```bash
curl -X POST https://api.genuineai.app/api/v1/knowledge-bases/<kb-id>/records \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{
    "data_title": "Harbor Series 40",
    "data_content": {
      "sku": "HS-40",
      "positioning": "Mid-range cruiser for coastal work",
      "tier": "premium"
    }
  }'
```

`data_title` is the record's label, and `data_content` holds the schema's fields. `PATCH`
and `DELETE` on `/records/{recordId}` edit and remove one record, and `DELETE
/knowledge-bases/{id}/records` empties the knowledge base.

To read records back, `GET /knowledge-bases/{id}/records` lists one base's records, and
`summary=true` returns them without the full content, the right call when rendering a
list rather than reading the data.

### Search across knowledge bases

```
GET /knowledge-base-records?ds_category=general&data_title=harbor
```

This collection is flat rather than nested, because the queries it exists for span
several knowledge bases at once: by name, by category, or by a list of ids. It is
read-only by design, since records are created and edited through the knowledge base that
owns them.

## Schemas

A knowledge template's fields come from a schema, called a form in platform terms. Its
`form_config.fields` declare each field's name, control type and options, and
`data_content` must match them. Fields of type `files` are excluded from import and
export, because a file reference does not survive a JSON round trip.

Standard schemas are shared: a knowledge base built from a platform template points at a
template everyone uses. To change the shape without affecting anyone else:

```
POST /knowledge-bases/{id}/fork-form
```

This copies the schema into one your workspace owns and re-points the knowledge base and
its existing records at the copy. **Fork before customizing**, or a later template update
overwrites your changes. It requires both `knowledge-base:edit` and `knowledge-base:schema:edit`.

## Fill a document knowledge base

A `singlePage` base holds one field, `text`. Use the same record endpoint with that one
field:

```bash
curl -X POST https://api.genuineai.app/api/v1/knowledge-bases/<kb-id>/records \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{"data_content": {"text": "Our refund policy is …"}}'
```

Omit `data_title` and the record takes the knowledge base's name.

## Fill a files knowledge base

Files-type bases are filled by uploading rather than by posting records. Follow the
standard [three-step upload](/uploading-files) with two differences: `purpose` is
`knowledge`, and `file_repo` points the file at the knowledge base.

<CodeTabs syncKey="lang">

```bash title="cURL"
curl -X POST https://api.genuineai.app/api/v1/files/upload-urls \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{
    "files": [{
      "file_name": "warranty-policy.pdf",
      "file_type": "application/pdf",
      "file_size": 184320,
      "purpose": "knowledge",
      "file_repo": { "ds": "<kb-id>" }
    }]
  }'
```

```js title="JavaScript"
const [entry] = await fetch(
  "https://api.genuineai.app/api/v1/files/upload-urls",
  {
    method: "POST",
    headers: { ...headers, "Content-Type": "application/json" },
    body: JSON.stringify({
      files: [{
        file_name: "warranty-policy.pdf",
        file_type: "application/pdf",
        file_size: 184320,
        purpose: "knowledge",
        file_repo: { ds: kb.ds_id },
      }],
    }),
  },
).then(r => r.json());
```

```python title="Python"
entry = requests.post(
    "https://api.genuineai.app/api/v1/files/upload-urls",
    headers=headers,
    json={"files": [{
        "file_name": "warranty-policy.pdf",
        "file_type": "application/pdf",
        "file_size": 184320,
        "purpose": "knowledge",
        "file_repo": {"ds": kb_id},
    }]},
).json()[0]
```

</CodeTabs>

PUT the bytes and finalize, and processing extracts the text, splits it and embeds it.
The file becomes retrievable once its `file_status` reaches `ready`. Before that the agent
cannot see it, so a document uploaded seconds before a question will not be in the answer.

List what a knowledge base holds from its own collection:

```
GET /knowledge-bases/<kb-id>/files
```

The same listing is reachable from `GET /files` by naming the repository, which is how
you ask for any of the others:

```
GET /files?repo_type=ds&repo_id=<kb-id>
```

A repository is required on `GET /files`, because files are always scoped to one rather
than listed across the workspace. `repo_type` names the kind and `repo_id` the specific
one. `repo_type=media` is the exception that takes no id, because the media library is a
library rather than a record.

`POST /files/actions/reindex` re-runs extraction and embedding for a whole repository,
named the same way:

```bash
curl -X POST https://api.genuineai.app/api/v1/files/actions/reindex \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{"repo_type": "ds", "repo_id": "<kb-id>"}'
```

Use it after replacing documents in bulk, or when extraction has improved since the files
were loaded.

## Import and export

**Import** takes an array of items:

```bash
curl -X POST https://api.genuineai.app/api/v1/knowledge-bases/<kb-id>/import \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{
    "items": [
      { "data_title": "Harbor Series 40", "data_content": { "sku": "HS-40", "tier": "premium" } },
      { "data_title": "Harbor Series 28", "data_content": { "sku": "HS-28", "tier": "standard" } }
    ],
    "sourceUrl": "https://example.com/catalog"
  }'
```

The caps, which are enforced rather than advisory:

| | |
|---|---|
| Items per call | 1,000 |
| Per item | 100 KB |
| Per call | 10 MB |
| Per field value | 50,000 characters. Longer values are **truncated, not rejected** |

Import is transactional and tolerant in specific ways: unknown fields are dropped rather
than failing the item, items left with no valid fields are skipped, and oversized items
are reported in an `errors` array while the rest are imported. Check the response counts
rather than treating a 200 as confirmation that everything landed.

Import and export work on `form` and `singlePage` bases only. A files-type base returns
400, because its content is files, which you upload instead.

**Export** (`GET /knowledge-bases/{id}/export`) returns the base's name, type,
description and every record's title and content as a downloadable JSON attachment. File
fields are stripped. Export followed by import copies a knowledge base between
workspaces and serves as a reasonable backup.

## Build a knowledge base from a website

```bash
curl -X POST https://api.genuineai.app/api/v1/knowledge-bases/<kb-id>/init-knowledge \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/products", "preview": true}'
```

Crawls the URL and extracts records **shaped by the schema**: it reads your field names,
labels and options and fills them, rather than dumping page text. That is why it works
only on `form` and `singlePage` bases, and why a well-designed schema produces a far
better crawl than a vague one.

Always run it with `preview: true` first, which returns the proposed records without
storing anything. Without preview it runs in the background, and the response returning
does not mean the crawl has finished.

## Listing and deleting

`GET /knowledge-bases` lists them, filtering on `ds_category`, `ds_type` and
name. Where a base is backed by a schema, `form_name` is included. `GET
/knowledge-bases/stats` returns counts for a dashboard.

Deletion is deliberately restricted. `DELETE /knowledge-bases/{id}` refuses while any
agent still uses the base, or while any file still sits in it, so detach the agents and
remove the files first. A failed delete is better than an agent silently losing the
knowledge it was answering from.

## Next steps

- [How agents use knowledge](/knowledge-retrieval) covers inline versus retrieved, and why an agent ignores what you loaded.
- [Agents and conversations](/agents-and-conversations) shows how knowledge is attached to an agent.
- [AI enrichment](/ai-enrichment) names the knowledge bases that control image analysis vocabulary.
