When to use it#
- Your steady traffic is above what the default rate limits of a few keys cover, and you want headroom that is yours alone.
- You need capacity that other customers' bursts cannot take, so a
529never reaches your hot path. - You are a defence or security body of an eligible state and need the models inside your own network.
What stays the same#
A dedicated deployment answers the same API: POST /v1/decide, sessions, batch jobs and GET /v1/models, with the same request and response shapes, the same errors and the same models. The same request to the same pinned version returns the same answer as on the shared API, so tests, thresholds and replays carry over unchanged.
What changes#
| Shared API | Dedicated | |
|---|---|---|
| Capacity | Shared with every customer | Reserved for your workspace on our own GPUs |
| Rate limits | The defaults per key and model | Set for your deployment, sized to your traffic |
| Limits | As on Limits | Set per deployment and published at its /v1/models: read them there, not from constants in your code |
| Price | $0.042 and $0.021 per 1M input tokens | Agreed for the capacity you reserve |
| Safety check | On every request | On every request, as on the shared API |
On your premises#
The models can run inside your own network only under the Defence Contract Addendum, which defence and security bodies of the member states of the North Atlantic Treaty Organization, the member states of the European Union, Australia, Japan, New Zealand and the Republic of Korea may sign. Military use is banned outside it. The addendum's rules are part of the model licence. Under it you:
- run our safety-check software beside the models, with its self-harm part always on;
- keep the records the addendum names for 2 years, and allow the audits it sets out;
- never use the models for what the addendum forbids in every case, whatever the contract says.
Responses carry x-decisionnode-safety wherever the check runs. A self-hosted server that runs without it sends no safety header at all, so code that reads the header should treat it as optional.
Health check#
A dedicated or on-premises deployment also answers GET /healthz, for your load balancer's readiness checks. It needs no key and costs nothing. It returns 200 once the deployment is ready to answer, and 529 with Retry-After while the models are still loading. On the shared API, check a key and your connection with GET /v1/models instead.
curl -i https://<your-deployment>/healthzHow to ask for it#
Write to hello@bynn.com with your model, your peak requests and input tokens a minute, and where it must run. We size the deployment with you.