# Rate limits

Limits are applied at the edge, before a request reaches the API. Several tiers can
govern one request at once, and the first tier to trip is the one that stops it.

| Tier | Quota | Window | Applies to |
|---|---|---|---|
| `api-key` | 600 | 60s | Any request carrying `X-Api-Key`. |
| `tenant` | 3000 | 60s | All authenticated traffic for one workspace. |
| `expensive` | 600 | 60s | AI generation (see below). |
| `ip` | 1500 | 60s | Per source address. |
| `ip-burst` | 300 | 10s | Per source address, short window. |
| `public` | 120 | 60s | Unauthenticated share and widget endpoints. |
| `public-session` | 60 | 60s | Minting a widget session. |
| `register` | 10 | 60s | Sign-up. |
| `delivery` | 6000 | 60s | Published asset URLs, per workspace and source address. |
| `delivery-burst` | 600 | 10s | Published asset URLs, short window. |
| `transform` | 120 | 60s | First render of an image variant not yet cached. |

An integration is normally governed by `api-key`, `tenant` and the two `ip` tiers
together. The 600 requests a minute on the key is the limit you will reach first.

The two `delivery` tiers are separate from the rest: they govern the public asset URLs on
the delivery hostname, not the API, so serving a page full of published images never counts
against an integration's budget. `transform` governs only the first render of an image
variant; once cached, requests for it count against `delivery` alone. Over the limit the
original bytes are served in place of the variant rather than an error.

## The expensive tier

Anything that runs a model is additionally counted against `expensive`: sending a message
to a thread, and the AI generation endpoints across designs, presentations, content and
inline assistance. These calls carry a real per-call cost, and the tier exists so that one
integration cannot spend a workspace's budget in a minute.

## Reading the policy header

```http
RateLimit-Policy: "api-key";q=600;w=60, "tenant";q=3000;w=60, "ip";q=1500;w=60
```

This is the [IETF structured
field](https://datatracker.ietf.org/doc/draft-ietf-httpapi-ratelimit-headers/), listing
every tier governing *this* request with its quota and window. Read it rather than
hard-coding the table above.

**There is deliberately no `RateLimit-Remaining` or `RateLimit-Reset`.** Counting happens
at the edge location nearest the caller and is eventually consistent, so a remaining
count would not be reliable. We publish the policy rather than a counter we cannot stand
behind.

## Handle a 429

```http
HTTP/1.1 429 Too Many Requests
Retry-After: 12
```

`Retry-After` is the value to act on: wait that many seconds. Retrying sooner consumes
quota without succeeding.

<CodeTabs syncKey="lang">

```js title="JavaScript"
if (res.status === 429) {
  const wait = Number(res.headers.get("Retry-After") ?? 5);
  await new Promise(r => setTimeout(r, wait * 1000));
  return retry();
}
```

```python title="Python"
import time

if res.status_code == 429:
    wait = int(res.headers.get("Retry-After", 5))
    time.sleep(wait)
    return retry()
```

</CodeTabs>

Back off rather than retrying immediately. A tight retry loop against a limit you have
already hit keeps you blocked.

## Limiter failures

The limiter fails open. If it breaks, requests are unlimited rather than rejected, so you
will not see 429s caused by our own infrastructure being unavailable.

Asset delivery is the one exception, and only for policy rather than for limits: a published
URL whose distribution rules cannot be read is refused with a 503 rather than served. A rule
that cannot be checked is never assumed to permit.
