# Agents and conversations

An agent is a configured model: a prompt, a model, a set of tools, and the knowledge it
answers from. A thread is a conversation with an agent, and messages run the model.

```
agent  →  thread  →  message  →  the reply streams back
```

## Create an agent

<CodeTabs syncKey="lang">

```bash title="cURL"
curl -X POST https://api.genuineai.app/api/v1/agents \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{
    "agent_name": "Archive researcher",
    "agent_description": "Answers questions about the photo archive",
    "agent_model": "genai-standard",
    "agent_temperature": "0.3",
    "agent_top_p": "0.95",
    "agent_audience": "internal"
  }'
```

```js title="JavaScript"
const agent = await fetch("https://api.genuineai.app/api/v1/agents", {
  method: "POST",
  headers: { ...headers, "Content-Type": "application/json" },
  body: JSON.stringify({
    agent_name: "Archive researcher",
    agent_description: "Answers questions about the photo archive",
    agent_model: "genai-standard",
    agent_temperature: "0.3",
    agent_top_p: "0.95",
  }),
}).then(r => r.json());
```

```python title="Python"
agent = requests.post(
    "https://api.genuineai.app/api/v1/agents",
    headers=headers,
    json={
        "agent_name": "Archive researcher",
        "agent_description": "Answers questions about the photo archive",
        "agent_model": "genai-standard",
        "agent_temperature": "0.3",
        "agent_top_p": "0.95",
    },
).json()
```

</CodeTabs>

`agent_name`, `agent_description`, `agent_model`, `agent_temperature` and `agent_top_p`
are all required. **`agent_model` is restricted**: a workspace may use the platform's own
`genai-` models, and foundation models only where they have been enabled for it. `GET
/agents` on an existing agent is the quickest way to see what your workspace accepts.

`agent_audience` is `internal` or `public`, and `agent_tools` names the capabilities the
agent may call. `GET /agents/{id}/preview-system-prompt` shows the composed system prompt,
everything the agent's configuration, knowledge and tools add up to, without running
anything. It is the fastest way to diagnose an agent that answers oddly.

`GET /agents/{id}/history` lists edits over time, so a change in behavior can be traced to
a change in configuration.

## Tools

`agent_tools` is what turns an agent from something that answers into something that acts.
Each entry names a capability the model may call mid-conversation, on its own initiative:

```json
{ "agent_tools": ["photo_selector", "image_generation", "logo_selector"] }
```

| Tool | Does |
|---|---|
| `image_generation` | Creates a new image. Not bound when the agent's own model is already an image model. |
| `create_design` | Creates an editable, layered design in the design library. Unlike an image, its headlines stay real text layers. |
| `photo_selector` | Searches the existing [media library](/organizing-assets) and picks photos. Selects; never creates. |
| `logo_selector` | Chooses the brand logo that suits the content and the background it sits on. |
| `brand_style` | Reads the workspace's brand colors and typography, so what the agent produces is styled in the brand. |
| `kb_lookup` | Looks up records and reference images in the agent's knowledge bases on demand. |
| `external_data` | Calls a connected third-party service. See [external data](/external-data). |
| `web_search` | Searches the web. |
| `url_context` | Reads the content at a URL. |
| `maps` | Looks up places and directions. |

The distinction between `image_generation` and `photo_selector` is the one to get right,
because users notice it: "make me a picture of a harbor" and "find me a picture of a
harbor" are different requests, and an agent carrying only one of the two tools will
confidently do the wrong thing. Give it both if it should be able to do either.

Unknown tool keys are ignored rather than refused, so a typo fails silently. Check
`GET /agents/{id}` to confirm what was stored.

### Tool calls in a thread

Tool calls appear in the thread as part of the exchange, and results that reference files
(a selected photo, a generated image) carry the `file_id`, so you can render or download
them like any other library file. `GET /agents/{id}/preview-system-prompt` shows the
instructions each tool adds to the agent's prompt, which is the fastest way to understand
why an agent reaches for one tool over another.

## Knowledge bases

A knowledge base is a set of records an agent retrieves from. Names are unique in the
workspace, because names are how agents refer to them:

```bash
curl -X POST https://api.genuineai.app/api/v1/knowledge-bases \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{"ds_name": "Brand guidelines", "ds_description": "Voice, tone and usage rules"}'
```

Building and filling one is a subject of its own. [Knowledge bases](/knowledge-bases)
covers the three types, schemas, importing and crawling, and [how agents use
knowledge](/knowledge-retrieval) covers how an agent uses what you put there.

Attach knowledge to an agent:

```bash
curl -X PATCH https://api.genuineai.app/api/v1/agents/<agent-id>/ds \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{"agent_tenant_ds": ["3f2a9c1e-…", "8a1b6e5c-…"]}'
```

Two fields, deliberately separate. `agent_ds` is the agent's own default set, named in
kebab-case. `agent_tenant_ds` is *your workspace's* selection layered on top. Send `null`
to drop the override and fall back to the agent's default.

<Callout type="tip">
Knowledge bases do more than answer questions. The Media Tags and Flags knowledge
bases are the controlled vocabulary [image analysis](/ai-enrichment#your-controlled-vocabulary)
is allowed to use. Editing them changes how every future upload is described.
</Callout>

## Threads

**You supply the thread id.** `thread_id` is required in the body rather than generated
for you, so mint a UUID client-side. That makes thread creation idempotent from your
side: a retry after a timeout reuses the id rather than leaving two conversations behind.

<CodeTabs syncKey="lang">

```bash title="cURL"
curl -X POST https://api.genuineai.app/api/v1/threads \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{
    "thread_id": "d9f1c2b4-6a55-4f0e-9c3a-1b2d3e4f5a6b",
    "thread_agent": "3f2a9c1e-…",
    "thread_title": "Q3 archive questions"
  }'
```

```js title="JavaScript"
const threadId = crypto.randomUUID();

await fetch("https://api.genuineai.app/api/v1/threads", {
  method: "POST",
  headers: { ...headers, "Content-Type": "application/json" },
  body: JSON.stringify({
    thread_id: threadId,
    thread_agent: agent.agent_id,
    thread_title: "Q3 archive questions",
  }),
});
```

```python title="Python"
import uuid

thread_id = str(uuid.uuid4())

requests.post(
    "https://api.genuineai.app/api/v1/threads",
    headers=headers,
    json={
        "thread_id": thread_id,
        "thread_agent": agent["agent_id"],
        "thread_title": "Q3 archive questions",
    },
)
```

</CodeTabs>

`thread_share_mode` decides who else can see the thread: `private`, `space` (everyone in
the space) or `collab`. `thread_space` puts it in a space, which requires write access to
that space.

For a conversation attached to something else, such as a file or a task, `POST
/threads/actions/find-or-create` takes a `thread_ref` and returns the existing thread or
starts one. It is a POST because it writes. To read without creating, use `GET
/threads?thread_ref=…`.

### A thread's files

```
GET /threads/{id}/files
```

Everything belonging to the conversation: what was attached to it, and what was generated
in it. Reading these files means reading the conversation, so its access rules apply on
top of each file's.

An image generated inside a conversation belongs to two places at once: that thread, and
the `generated` library where every AI output lands. It is returned by this listing
and by `GET /media-files?repo_type=generated` alike. See [organizing
assets](/organizing-assets) for the repository vocabulary.

## Send a message

```
POST /threads/{id}/messages
```

`msg_content` is an array of parts rather than a string, which is how a message carries
files from your library alongside its text:

```json
{
  "msg_content": [
    { "type": "text", "text": "Which of these would work as a homepage hero?" },
    { "type": "image_url", "file_id": "8a1b6e5c-…", "file_name": "harbor-sunset.jpg", "file_type": "image/jpeg" }
  ]
}
```

| Part `type` | |
|---|---|
| `text` | Carries `text`. |
| `image_url` | An image from the library, by `file_id`. |
| `file` | A document from the library, by `file_id`. |
| `reference` | A file to retrieve against rather than attach whole. |

<Callout type="caution">
**This is the endpoint that runs a model.** It consumes credits and counts against the
`expensive` [rate-limit tier](/rate-limits).
</Callout>

### The reply streams on the response

The response body is the assistant's reply, written as it is generated. Read it as a
stream of text rather than waiting for a complete JSON document:

<CodeTabs syncKey="lang">

```js title="JavaScript"
const res = await fetch(
  `https://api.genuineai.app/api/v1/threads/${threadId}/messages`,
  {
    method: "POST",
    headers: { ...headers, "Content-Type": "application/json" },
    body: JSON.stringify({
      msg_content: [{ type: "text", text: "Summarize this album." }],
    }),
  },
);

const reader = res.body.getReader();
const decoder = new TextDecoder();
let reply = "";

for (;;) {
  const { done, value } = await reader.read();
  if (done) break;
  const chunk = decoder.decode(value, { stream: true });
  reply += chunk;
  process.stdout.write(chunk);
}
```

```python title="Python"
with requests.post(
    f"https://api.genuineai.app/api/v1/threads/{thread_id}/messages",
    headers=headers,
    json={"msg_content": [{"type": "text", "text": "Summarize this album."}]},
    stream=True,
) as res:
    reply = ""
    for chunk in res.iter_content(chunk_size=None, decode_unicode=True):
        reply += chunk
        print(chunk, end="", flush=True)
```

</CodeTabs>

<Callout type="note">
The response carries `Content-Type: text/event-stream`, but the body is **plain text
chunks, not SSE frames**. There are no `event:` or `data:` lines to parse. Append
them as they arrive. The structured events are on the separate stream below.
</Callout>

The finished message is persisted either way, so a dropped connection loses the live
output but not the reply, which remains available from `GET /threads/{id}/messages`.

### Retry a message

`retry: true` re-runs the last exchange. It removes the previous assistant message only
if that message failed, so a retry after an error does not duplicate a good answer.

For image generation, `imageQuality` (`1K`, `2K`, `4K`) and `aspectRatio` are accepted
and remembered on the thread.

## The event stream

```
GET /threads/{id}/events
```

A properly framed SSE stream, for following a conversation that more than one client is
watching:

| Event | |
|---|---|
| `hello` | Sent on connect: the thread's current status, who it is generating for, and who is present. |
| `user_message` | Someone posted a message. |
| `assistant_message` | A reply completed. |
| `lock` | The thread started or stopped generating, and for whom. |
| `presence` | Who is viewing or typing. |

A `: keepalive` comment arrives every 25 seconds. Ignore it, and do not treat the absence
of events as a dead connection any sooner than that.

<Callout type="note">
**Events are broadcast only for shared threads.** A private, single-user thread publishes
nothing but `hello` and presence; its reply arrives on the POST response described above.
Do not wait on the event stream for a conversation only your integration is having,
because the events never arrive.
</Callout>

`POST /threads/{id}/presence` reports that you are viewing or typing, which is what
populates other clients' `presence` events.

### The generating lock

A shared thread accepts one generation at a time. Posting while it is busy returns
**409**, because the reply to someone else's message is still being written. Wait for the
`lock` event reporting `thread_status: "idle"` rather than retrying immediately.

Private threads do not take the lock, so they never return 409 for this reason.

## Reusable prompts

A prompt is a saved instruction with declared inputs, so the same task is not retyped
differently every time. `GET /prompts/{id}/form` returns the inputs it asks for, and
wildcards (`/wildcards`) define how each is collected: a free-text box, a list, or a file
picker.

Send one by including a `form` part in `msg_content` naming the `prompt_id`. The platform
resolves it into text and attachments before the model sees it.

Prompts are organized with categories rather than free-text tags. `GET
/prompt-categories` returns the workspace's own terms plus the product-shipped ones, and
`PUT /prompts/{id}/categories` files a prompt. Filing is per workspace, so filing a
product prompt affects nobody else. Filter the catalog with the
`prompt_tenant_categories` parameter on `GET /prompts`.

## Memories

`/memories` holds what an agent remembers about the calling user between conversations.
`GET` lists them, `POST` adds one, and `DELETE /memories` clears them all. Memories are
per user rather than per workspace, so one person's memories never inform another's
conversations.

## Next steps

- [AI enrichment](/ai-enrichment) covers the analysis that makes library files answerable.
- [Organizing and finding assets](/organizing-assets) is how you find the files to attach.
- [Rate limits](/rate-limits) describes the `expensive` tier that governs every model call.
