Limits
Rate limits
Limits are set by your tier and apply per API key. Four things are counted independently: requests per minute, tokens per minute, per-modality daily caps, and how many requests you may have in flight at once. Exceeding any of them is a 429.
Response headers
HTTP/1.1 200 OK X-RateLimit-Limit: 600 X-RateLimit-Remaining: 587 X-RateLimit-Reset: 1700000060 X-RateLimit-Limit-Tokens: 2000000 X-RateLimit-Remaining-Tokens: 1984210 X-Request-Id: 9c1f4e7a-2b8d-4e31-9a0c-5f7b1d2e3a44 // On a 429 the same headers are present, plus: Retry-After: 43
| Header | Meaning |
|---|---|
X-RateLimit-Limit | Requests per minute allowed on your tier |
X-RateLimit-Remaining | Requests left in the current minute |
X-RateLimit-Reset | Unix seconds at which the request window resets |
X-RateLimit-Limit-Tokens | Tokens per minute allowed, when your tier caps them |
X-RateLimit-Remaining-Tokens | Tokens left in the current window |
Retry-After | Seconds to wait. Sent on every 429; honour it rather than retrying at once. |
X-Request-Id | The id this request is recorded under. Quote it in any support request. |
Back off politely. Every
429 carries Retry-After in seconds. The official SDKs already honour it; if you call the API directly, wait that long rather than retrying immediately — a tight retry loop makes the limit last longer. Your current limits for a given model are in GET /v1/models/{model}.