GenuineAIGenuineAI
For usersFor developers
  • Overview
  • Guides
  • How-to
  • API reference
Documentation
  • Quickstart
  • Authentication
  • Errors
  • Rate limits
  • API reference
Platform
  • Platform overview
  • Solutions
  • Pricing
  • Sign in
Company
  • Status
  • Trust Center
  • Contact
  • Terms of Service
  • Privacy Policy

© 2026 GenuineAI Ventures LLC

support@genuinehq.com
Start here
QuickstartAuthentication
Working with the API
Objects and fieldsErrorsPaginationRate limitsMaintenance windowsVersioning and deprecation
Security
Security and tenancyRoles and permissionsPermission referenceSharing and access
Your plan
Plans and modules
Guides

Rate limits

Limits are applied at the edge, before a request reaches the API. Several tiers can govern one request at once, and the first tier to trip is the one that stops it.

TierQuotaWindowApplies to
api-key60060sAny request carrying X-Api-Key.
tenant300060sAll authenticated traffic for one workspace.
expensive60060sAI generation (see below).
ip150060sPer source address.
ip-burst30010sPer source address, short window.
public12060sUnauthenticated share and widget endpoints.
public-session6060sMinting a widget session.
register1060sSign-up.
delivery600060sPublished asset URLs, per workspace and source address.
delivery-burst60010sPublished asset URLs, short window.
transform12060sFirst render of an image variant not yet cached.

An integration is normally governed by api-key, tenant and the two ip tiers together. The 600 requests a minute on the key is the limit you will reach first.

The two delivery tiers are separate from the rest: they govern the public asset URLs on the delivery hostname, not the API, so serving a page full of published images never counts against an integration's budget. transform governs only the first render of an image variant; once cached, requests for it count against delivery alone. Over the limit the original bytes are served in place of the variant rather than an error.

The expensive tier

Anything that runs a model is additionally counted against expensive: sending a message to a thread, and the AI generation endpoints across designs, presentations, content and inline assistance. These calls carry a real per-call cost, and the tier exists so that one integration cannot spend a workspace's budget in a minute.

Reading the policy header

Code
RateLimit-Policy: "api-key";q=600;w=60, "tenant";q=3000;w=60, "ip";q=1500;w=60

This is the IETF structured field, listing every tier governing this request with its quota and window. Read it rather than hard-coding the table above.

There is deliberately no RateLimit-Remaining or RateLimit-Reset. Counting happens at the edge location nearest the caller and is eventually consistent, so a remaining count would not be reliable. We publish the policy rather than a counter we cannot stand behind.

Handle a 429

Code
HTTP/1.1 429 Too Many Requests Retry-After: 12

Retry-After is the value to act on: wait that many seconds. Retrying sooner consumes quota without succeeding.

Back off rather than retrying immediately. A tight retry loop against a limit you have already hit keeps you blocked.

Limiter failures

The limiter fails open. If it breaks, requests are unlimited rather than rejected, so you will not see 429s caused by our own infrastructure being unavailable.

Asset delivery is the one exception, and only for policy rather than for limits: a published URL whose distribution rules cannot be read is refused with a 503 rather than served. A rule that cannot be checked is never assumed to permit.

Last modified on October 8, 2026
PaginationMaintenance windows
On this page
  • The expensive tier
  • Reading the policy header
  • Handle a 429
  • Limiter failures
if (res.status === 429) { const wait = Number(res.headers.get("Retry-After") ?? 5); await new Promise(r => setTimeout(r, wait * 1000)); return retry(); }