| Limit | Default | Counted as |
|---|---|---|
| Requests | 40 per second | Every call to /v1/decide |
| Input tokens | 100,000 per second | usage.input_tokens across all requests |
Both limits are shared by every key in a workspace and by both models. Your current values are on the Rate limits tab of API keys in the console.
When you hit a limit#
The API answers 429 with a Retry-After header in seconds. Wait that long and retry the same request; nothing was charged for the refused one. Retrying sooner only earns another 429.
HTTP/1.1 429 Too Many RequestsRetry-After: 2x-request-id: req_01J9Z6Q8T4content-type: application/json{"detail": {"error_type": "rate_limited", "message": "Too many requests. Retry after 2 seconds."}}Staying under them#
- Bound concurrency in batch jobs (see Fan-out) instead of firing everything at once.
- Ask all questions about an input in one request: one request instead of many, and the state counted once.
- Trim states. Tokens per second is usually the limit a backlog hits first.
- Need a higher ceiling or reserved GPUs? Dedicated capacity is set to your traffic; ask through the console.