| Limit | Value | Notes | Over the limit |
|---|---|---|---|
| Context per request | 64k tokens (65,536) | State, images and every question together | 400 max_tokens_exceeded |
| Images per request | Up to 4 | JPEG, PNG or WebP, base64 in data | 400 api_usage_error |
| Image size | Up to 10 MB each | Measured on the image bytes; base64 text is about a third longer | 400 api_usage_error |
| Options per choice | 1 to 255 | Two or more recommended | 400 Too many choices |
| Levels per score | 1 to 10 | Lowest first; two or more recommended | 400 Too many score levels |
| Values per number question | Up to 256 | Whole-number grids in the first release; halves, decimals and wider grids follow | 400 api_usage_error |
| Response | One buffered JSON body | On /v1/decide every answer arrives together; sessions coming soon stream one reply per frame | |
A request over a limit is refused before any model work and costs nothing. The state is never cut to make it fit, and the body names the limit you crossed. Bodies and fixes are on Errors.
Sessions coming soon#
Sessions stream decisions over one open connection, for control loops that decide more than 10 times a second. Their limits apply on top of the request limits above.
| Limit | Value | Notes |
|---|---|---|
| Frame | At most 4,096 tokens, or one image | Text, a JSON value or an image |
| Window | 8 recent frames by default | Set with window when the session opens |
| Lifetime | 300 seconds by default, at most 3,600 | Set with ttl_seconds |
| Sessions per key | Published at /v1/models | One session per socket |
Throughput#
New workspaces start at 40 requests per second and 100,000 input tokens per second (provisional defaults). Your current limits are in the console. Over the limit, the API answers 429 with a Retry-After header. See Rate limits.