# Uploading files

Bytes never pass through the API. An upload is three calls, and the middle one goes
straight to storage:

```
1. POST /files/upload-urls    → a signed URL per file
2. PUT  <the signed URL>      → the bytes, direct to storage
3. PUT  /files/{id}/finalize  → admits the file and starts processing
```

The file record exists from step 1, but **a file that is never finalized never appears**:
it stays invisible to every list, and nothing processes it. Step 3 is required rather
than bookkeeping.

## Step 1: request upload URLs

Batched, up to **300 files** per call:

<CodeTabs syncKey="lang">

```bash title="cURL"
curl -X POST https://api.genuineai.app/api/v1/files/upload-urls \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{
    "files": [
      {
        "file_name": "harbor-sunset.jpg",
        "file_type": "image/jpeg",
        "file_size": 2481923,
        "purpose": "media"
      }
    ]
  }'
```

```js title="JavaScript"
const [entry] = await fetch(
  "https://api.genuineai.app/api/v1/files/upload-urls",
  {
    method: "POST",
    headers: { ...headers, "Content-Type": "application/json" },
    body: JSON.stringify({
      files: [{
        file_name: file.name,
        file_type: file.type,
        file_size: file.size,
        purpose: "media",
      }],
    }),
  },
).then(r => r.json());
```

```python title="Python"
entry = requests.post(
    "https://api.genuineai.app/api/v1/files/upload-urls",
    headers=headers,
    json={"files": [{
        "file_name": "harbor-sunset.jpg",
        "file_type": "image/jpeg",
        "file_size": 2481923,
        "purpose": "media",
    }]},
).json()[0]
```

</CodeTabs>

| Field | |
|---|---|
| `file_name` | Required. The extension decides the stored object's extension. |
| `file_type` | Required. The MIME type. The signed URL is pinned to it, so the `PUT` must send the same one. |
| `file_size` | Required, in bytes. Used for the quota check and to bound the signed URL. |
| `purpose` | Required. Which library the file belongs to. See below. |
| `file_repo` | Optional. Which repository the file belongs to, and how a document joins a [knowledge base](/knowledge-bases). See below. |
| `album_ids` | Optional. Pre-assigns the file to albums, so it lands already filed. |
| `mediaTags` | Optional. Media tags applied on arrival. |

### Repository

`file_repo` is one key naming what the file belongs to. You must be able to write to
whatever it names: filing a document into a knowledge base you cannot edit is refused
rather than silently accepted.

| Key | Value | |
|---|---|---|
| `library` | `["media"]` or `["docs"]` | The shared media library, or the document library. An array, because a library is a place rather than an object; `generated` is the third name in it, and the platform sets that one. |
| `ds` | knowledge base id | The document joins that knowledge base and becomes retrievable. |
| `thread` | conversation id | The file belongs to that conversation. |
| `design` | design id | The file belongs to that design. |
| `prompt` | prompt id | A reference attached to that saved prompt. |

More than one key may apply. An image generated inside a conversation carries both the
`generated` library and its `thread`, and is returned by either listing.

Omit `file_repo` and the file belongs to nothing in particular, which is correct for a
font or a profile picture. Such a file is fetched by id and never appears in a listing,
because every listing names a repository.

A file read back may carry other keys: the presentation job that generated it, the form
submission it arrived on, the workflow run behind it. These record how the file came to
exist, are set only by the platform, and are rejected if you send them. Drop them before
echoing a file's own `file_repo` back on a copy.

Move a file afterward with `POST /files/{id}/move` rather than a field on `PUT
/files/{id}`: where a file sits decides who can reach it, so both ends are checked.

### Purpose

`purpose` decides where the file lives and which types are accepted. The three most common
values:

| Purpose | For | Accepts |
|---|---|---|
| `media` | The media library, the DAM proper | Images, video, audio |
| `document` | The document library | Documents: PDF, office formats, text |
| `file-repo` | General file storage | Any type |

The rest cover narrower cases: `knowledge` (files backing a knowledge base),
`media-knowledge` (reference images for AI, images only), `font` and `chat`. Those seven
are the complete list, and anything else is rejected with a 422.

Each purpose is gated by its own permission, so a key that may upload media is not
thereby allowed to upload knowledge.

### The response mixes accepted and rejected files

The response is an array in the order you sent, but **it contains both accepted and
rejected files**. An accepted entry is the created file record plus an `uploadUrl`, and a
rejected one carries `rejection_reason` and no URL:

```json
[
  {
    "file_id": "8a1b6e5c-0d31-4c0d-9d2f-7b444c0d9d2f",
    "file_name": "harbor-sunset.jpg",
    "file_status": "pendingUpload",
    "uploadUrl": "https://…"
  },
  {
    "file_name": "notes.txt",
    "file_type": "text/plain",
    "rejection_reason": "Only image, video, and audio files are allowed for media purpose"
  }
]
```

The call returns 200 either way. A rejection here concerns *that file* (its type does not
suit its purpose, it exceeds the per-file size limit, or your key lacks the permission that
purpose requires) rather than the request being wrong. A malformed request, including a `purpose`
outside the seven above, returns 422 for the whole call.

Branch on the presence of `uploadUrl` rather than on the status code:

<CodeTabs syncKey="lang">

```js title="JavaScript"
const accepted = entries.filter(e => e.uploadUrl);
const rejected = entries.filter(e => !e.uploadUrl);
```

```python title="Python"
accepted = [e for e in entries if e.get("uploadUrl")]
rejected = [e for e in entries if not e.get("uploadUrl")]
```

</CodeTabs>

One condition does fail the whole batch: exceeding your storage entitlement returns
**403** before any URL is issued, because the check runs against the batch total.

## Step 2: PUT the bytes

Send them straight to the signed URL with no API credentials, because the URL is the
authorization:

<CodeTabs syncKey="lang">

```bash title="cURL"
curl -X PUT "<uploadUrl>" \
  -H "Content-Type: image/jpeg" \
  --data-binary @harbor-sunset.jpg
```

```js title="JavaScript"
await fetch(entry.uploadUrl, {
  method: "PUT",
  headers: { "Content-Type": file.type },   // must match file_type exactly
  body: file,
});
```

```python title="Python"
with open("harbor-sunset.jpg", "rb") as fh:
    requests.put(
        entry["uploadUrl"],
        data=fh,
        headers={"Content-Type": "image/jpeg"},
    )
```

</CodeTabs>

Three constraints are encoded in the URL itself, so getting them wrong fails at storage
rather than at the API:

- **It expires after one hour.** For a long queue, request URLs in batches as you go
  rather than all at once up front.
- **The `Content-Type` must match** the `file_type` you declared.
- **The size is bounded** to your declared `file_size` plus 1%, capped at the per-file
  maximum.

If a URL expires before you use it, `PUT /files/{id}/update-url` issues a fresh one for
the same file, so you do not lose the record or its album assignments.

<Callout type="note">
This is also why the API's 10 MB request body cap does not apply to uploads: it governs
JSON bodies sent to the API, and your bytes never go there.
</Callout>

## Step 3: finalize

<CodeTabs syncKey="lang">

```bash title="cURL"
curl -X PUT https://api.genuineai.app/api/v1/files/<file-id>/finalize \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>"
```

```js title="JavaScript"
const file = await fetch(
  `https://api.genuineai.app/api/v1/files/${entry.file_id}/finalize`,
  { method: "PUT", headers },
).then(r => r.json());
```

```python title="Python"
file = requests.put(
    f"https://api.genuineai.app/api/v1/files/{entry['file_id']}/finalize",
    headers=headers,
).json()
```

</CodeTabs>

Finalize confirms the bytes arrived, records their true size, type and MD5, and hands the
file to processing: thumbnails, [AI analysis](/ai-enrichment), and indexing. It answers
with the file record.

Finalize each file with its own `PUT` as soon as its bytes land, rather than finalizing
the batch at the end. Processing starts sooner and one slow upload no longer holds up the
rest.

## Check when a file is ready

Processing is asynchronous. Poll `GET /files/{id}` and watch `file_status`:

| `file_status` | Means |
|---|---|
| `pendingUpload` | The row exists, the bytes have not been confirmed. Pre-finalize. |
| `pending` | Finalized, queued for or undergoing processing. |
| `ready` | Processing finished. Thumbnails, analysis and embeddings are in place. |
| `error` | Processing failed. The file stays, but its renditions do not. |

<CodeTabs syncKey="lang">

```js title="JavaScript"
async function waitUntilReady(fileId, { timeoutMs = 120000 } = {}) {
  const deadline = Date.now() + timeoutMs;
  while (Date.now() < deadline) {
    const file = await fetch(
      `https://api.genuineai.app/api/v1/files/${fileId}`,
      { headers },
    ).then(r => r.json());

    if (file.file_status === "ready") return file;
    if (file.file_status === "error") throw new Error(`Processing failed: ${fileId}`);

    await new Promise(r => setTimeout(r, 2000));
  }
  throw new Error(`Timed out waiting for ${fileId}`);
}
```

```python title="Python"
import time

def wait_until_ready(file_id, timeout=120):
    deadline = time.time() + timeout
    while time.time() < deadline:
        file = requests.get(
            f"https://api.genuineai.app/api/v1/files/{file_id}",
            headers=headers,
        ).json()

        if file["file_status"] == "ready":
            return file
        if file["file_status"] == "error":
            raise RuntimeError(f"Processing failed: {file_id}")

        time.sleep(2)
    raise TimeoutError(file_id)
```

</CodeTabs>

Poll on the order of seconds rather than milliseconds. A small image is usually ready in a
few seconds, while video and long documents take proportionally longer, and every poll
spends rate-limit quota.

Most integrations do not need to wait at all. The file is listed and downloadable as soon
as it is finalized. Waiting matters only when your next step depends on something
processing produces, such as a thumbnail, extracted text, or the embeddings that
[similarity search](/ai-enrichment) needs.

## Import from Dropbox or Google Drive

Importing skips the local round trip: you supply links, and the platform fetches them.
Between 1 and 500 files per call.

```bash
curl -X POST https://api.genuineai.app/api/v1/file-imports \
  -H "X-Api-Key: gai_…" -H "X-Tenant-Id: <workspace-id>" \
  -H "Content-Type: application/json" \
  -d '{
    "provider": "dropbox",
    "files": [{ "name": "brochure.pdf", "link": "https://…" }]
  }'
```

`provider` is `dropbox` or `google-drive`. Dropbox entries carry a direct `link`, while
Drive entries carry `driveFileId` and an `accessToken` instead. Poll `GET
/file-imports/{id}` for progress. Imported files move through `cloudImport` and
`importing` before reaching the same `ready` state as any other file.

## Versions

Uploading a corrected file over an existing one keeps the history rather than replacing
it. Upload the new file normally, then attach it as a version of the original:

```
POST /files/{id}/versions
{ "source_file_id": "<the file you just uploaded>" }
```

`GET /files/{id}/versions` lists them, `POST /files/{id}/versions/{version_id}/revert`
makes an older version current again, and each version has its own download URL.
Reverting is a swap rather than a delete: the version you were on becomes a version in
the history.

Versions count against your storage entitlement, because the bytes are retained.

## Delete and recover files

`DELETE /files/{id}` is a soft delete: the file leaves the library but the bytes remain
for a grace period. `GET /files/trash` lists what is recoverable, and `POST
/files/trash/{id}/recover` restores one.

## What to check when an upload fails

| Symptom | Usual cause |
|---|---|
| 403 on `upload-urls` | Storage entitlement exceeded for the batch total. |
| `rejection_reason` on an entry | The type is not allowed for that `purpose`, or the file exceeds the per-file size limit. |
| The `PUT` to storage fails | `Content-Type` does not match `file_type`, the body is larger than declared, or the URL expired. |
| Finalize returns an error | The bytes never landed, meaning the `PUT` did not succeed. |
| Stuck at `pending` | Processing is queued or slow. A permanent failure ends at `error` rather than `pending`. |
| File never appears in a list | It was never finalized. |

## Next steps

- [Organizing and finding assets](/organizing-assets) covers albums, tags, and the filter vocabulary.
- [AI enrichment](/ai-enrichment) covers what analysis produces, and searching by meaning.
- [Delivering assets](/delivering-assets) covers download URLs, bulk ZIPs and share links.
