state
“Customer: I was charged twice and nobody has replied for 3 days.”
questions
- routechoice
- urgencyscore
- refundtruth
POST /v1/decide3 questions, 1 call
- route0.87
- urgency2.31
- refund0.94
- one pass
- no text generated
- our own GPUs
- input
- 62 tokens
- output
- 0, free
answers
- route"billing"
- urgency2.31
- refund0.94true
your code acts
- if refund >= 0.8issue_refund()
- elif refund >= 0.4recheck()
- elseassign(route)
What you send and what comes back#
One request carries a state (the thing to decide about) and a map of questions. The response carries one answer per question, under the same keys, in the shape its type defines. Here is a support message routed, ranked and checked for an automatic refund in one call.
{
"model": "decisionnode-flash-latest",
"state": "Customer: I was charged twice and nobody has replied for 3 days.",
"questions": {
"route": {
"type": "choice",
"instructions": "Where should this go?",
"criteria": { "billing": "money", "bug": "broken", "account": "login" }
},
"urgency": {
"type": "score",
"instructions": "How urgent is this?",
"criteria": ["routine", "today", "urgent", "critical"]
},
"refund": {
"type": "truth",
"instructions": "Refund this automatically?"
}
}
}The model read the message once and answered three questions in one pass. It generated no text, so output_tokens is 0 and you pay only for the 62 input tokens. This request goes to DecisionNode-1.0 Flash, about 5 ms server side. A refund in the unsure middle (0.40 to 0.80) is re-checked by the full model, decisionnode-latest, as the diagram shows; everything else is settled by the first answer.
Four question types#
Every type names its answer: choice (one of your labels), score (a level on your scale), truth (a probability), number (a value on your grid).
| Type | In the API | What it asks | What comes back |
|---|---|---|---|
| Choice | "type": "choice" | Pick one option from the list you give | choice (always one of your keys), probabilities per option, confidence |
| Score | "type": "score" | Place the input on your ordered scale | score (expected level), probabilities per level, legend, confidence |
| Truth | "type": "truth" | Is this statement true? | truth, one number from 0.00 to 1.00: the probability that the statement is true |
| Number | "type": "number" | How many, or what value? | number (a value on your grid), expected, confidence, probabilities per value |
Reading a Truth answer. One calibrated probability from 0.00 to 1.00 that the statement is true. 0.80 means true about 8 times in 10. Your code picks the cut-off: 0.50 for a plain yes or no, higher when acting on a false yes is costly.
Reading a Number answer. One value on a grid you set (min, max and an optional step), with a calibrated probability for every value: the most probable value, the expected value and the confidence. Count the cars in a drone frame, read the year off a contract, count the overdue invoices in a statement. See Number.
Streaming decisions coming soon. More than 10 decisions a second, for a drone, a simulator or a monitoring loop? Open a session: send the context once, stream frames over a WebSocket and get a typed decision for each. Control loops shows where it fits.
Why a decision model#
Rules break on real input: a regex does not know that "charged twice" is a billing problem. A chat model knows, but it writes text you have to parse, and it does not answer the same way twice. DecisionNode is trained to decide, not to write.
- Always valid. Every answer has the shape its type defines. A choice is always one of the keys you sent, so there is nothing to parse or repair.
- Calibrated. Probabilities are calibrated per question type, so a threshold on confidence means what it says. See Confidence.
- Deterministic. The same request to the same model version returns the same answer. There is no sampling. See Determinism.
- State read once. The state and images are encoded once and shared by every question, so ten questions cost little more than one.
The model and the machine#
We train the models and we run them, on our own inference stack and our own GPUs. There is no third-party API between your request and the answer, so nothing waits in someone else's queue. DecisionNode-1.0 Flash answers a short request in about 5 ms, measured server side in October 2026: fast enough to sit inside every request, every payment and every frame of a robot's loop.
| Model | Model id | Price per 1M input tokens | Output |
|---|---|---|---|
| DecisionNode-1.0 | decisionnode-latest | $0.042 | Free |
| DecisionNode-1.0 Flash | decisionnode-flash-latest | $0.021 (provisional) | Free |
Start here#
- Quickstart5 minGet a key, send one request, branch on the answer. Under five minutes.Read
- With coding agentsagentsA prompt to paste into your coding agent and a decide tool for your own agents.Read
- ExamplesrecipesComplete recipes: support triage, moderation, receipts, listing photos, lead scoring.Read
- POST /v1/decidereferenceEvery field of the request and the response, headers and errors.Read