Limits

Rate limits

Limits are set by your tier and apply per API key. Four things are counted independently: requests per minute, tokens per minute, per-modality daily caps, and how many requests you may have in flight at once. Exceeding any of them is a 429.

Response headers
HTTP/1.1 200 OK
X-RateLimit-Limit: 600
X-RateLimit-Remaining: 587
X-RateLimit-Reset: 1700000060
X-RateLimit-Limit-Tokens: 2000000
X-RateLimit-Remaining-Tokens: 1984210
X-Request-Id: 9c1f4e7a-2b8d-4e31-9a0c-5f7b1d2e3a44

// On a 429 the same headers are present, plus:
Retry-After: 43
HeaderMeaning
X-RateLimit-LimitRequests per minute allowed on your tier
X-RateLimit-RemainingRequests left in the current minute
X-RateLimit-ResetUnix seconds at which the request window resets
X-RateLimit-Limit-TokensTokens per minute allowed, when your tier caps them
X-RateLimit-Remaining-TokensTokens left in the current window
Retry-AfterSeconds to wait. Sent on every 429; honour it rather than retrying at once.
X-Request-IdThe id this request is recorded under. Quote it in any support request.
Back off politely. Every 429 carries Retry-After in seconds. The official SDKs already honour it; if you call the API directly, wait that long rather than retrying immediately — a tight retry loop makes the limit last longer. Your current limits for a given model are in GET /v1/models/{model}.