A batch job takes work nothing is waiting on: an archive to re-score, a backlog to route, a catalog to tag. You upload the requests, finalize the batch, and poll until it has ended. Then you download one result line per request, each with the same status, body and answers as a call to /v1/decide on the same model version. Completed requests are billed at half the live price per input token; nothing else is billed.
When to use a batch#
| The work | Use | Why |
|---|---|---|
| Something waits on the answer: a payment, a ticket, an agent's next step | POST /v1/decide | One request, one answer, in milliseconds |
| A large set that can wait: backfills, re-scoring after a policy change, tagging a catalog, testing thresholds on last month's data | A batch job | Half the price per input token, no rate limits to pace, and the same answers |
| More than 10 decisions a second about one evolving input | A session | The context stays loaded; each frame costs only its own tokens |
The flow#
- 1
Write the requests#
One request per line of a JSON Lines file, each a
/v1/decidebody plus acustom_idyou choose: 1 to 512 characters, unique in the batch. It is how you match each result to its input. A JSON array of the same objects works too.{"custom_id":"ticket-48213","model":"decisionnode-latest","state":"Customer: I was charged twice and nobody has replied for 3 days.","questions":{"route":{"type":"choice","instructions":"Where should this go?","criteria":{"billing":"money","bug":"broken","account":"login"}},"refund":{"type":"truth","instructions":"Refund this automatically?"}}} {"custom_id":"ticket-48214","model":"decisionnode-latest","state":"Customer: The app crashes every time I open the invoices tab.","questions":{"route":{"type":"choice","instructions":"Where should this go?","criteria":{"billing":"money","bug":"broken","account":"login"}},"refund":{"type":"truth","instructions":"Refund this automatically?"}}} {"custom_id":"ticket-48215","model":"decisionnode-latest","state":"Customer: My login stopped working after I changed my email address.","questions":{"route":{"type":"choice","instructions":"Where should this go?","criteria":{"billing":"money","bug":"broken","account":"login"}},"refund":{"type":"truth","instructions":"Refund this automatically?"}}} - 2
Create the batch#
POST /v1/batcheswith the file as the body. You get a batch in statedraft. Lines the API refuses for their shape (an unknown model, a question it cannot read) are kept as failed lines with the live error, so one bad line never stops the rest. See Create a batch.API=https://api.decisionnode.com/v1 curl "$API/batches" \ -H "Authorization: Bearer $DECISIONNODE_API_KEY" \ -H "Content-Type: application/x-ndjson" \ --data-binary @tickets.jsonl - 3
Add more requests, if you need to#
While the batch is a draft,
POST /adds lines, up to 10 MiB a call and 10,000 requests in all. Useful when you build a batch as records arrive. See Add requests.v1/ batches/ {id}/ requests - 4
Finalize it#
POST /locks the batch and queues it. Nothing runs before this, and nothing is billed. From now on the 24-hour expiry counts. See Finalize a batch.v1/ batches/ {id}/ finalize API=https://api.decisionnode.com/v1 BATCH=batch_01J9ZB8Q3T curl -X POST "$API/batches/$BATCH/finalize" \ -H "Authorization: Bearer $DECISIONNODE_API_KEY" - 5
Poll until it has ended#
GET /v1/batches/{id}returns the state and the counters. Poll every few minutes; polling more often does not make a batch run sooner, and every poll counts against your request rate. The batch has ended whenstateiscompleted,failed,cancelledorexpired. See Retrieve a batch.API=https://api.decisionnode.com/v1 BATCH=batch_01J9ZB8Q3T curl "$API/batches/$BATCH" \ -H "Authorization: Bearer $DECISIONNODE_API_KEY" - 6
Download the results#
GET /v1/batches/{id}/resultsreturns one JSON line per request, in upload order:custom_id,status,bodyandheaders. Read each line'sstatusbefore its body: a completed batch can hold failed lines. Results are kept 7 days. See Get batch results.API=https://api.decisionnode.com/v1 BATCH=batch_01J9ZB8Q3T curl "$API/batches/$BATCH/results" \ -H "Authorization: Bearer $DECISIONNODE_API_KEY" - 7
Resubmit what expired#
Lines with status
"expired"were still unfinished 24 hours after you finalized. They were not billed. Put the same requests in a new batch. Keep your own copy of every request: the API deletes the requests of a batch once its results are written.
A complete program#
Reads a backlog, splits it into batches of 10,000, submits them 10 at a time (your key's active batches), waits for each, acts on every answer and resubmits what expired or failed on our side. Every call rides out a network fault or a server error with a growing wait; a create that fails after it went through leaves no draft behind, and a finalize retried after it went through reads the batch instead. Run it as a job: it sleeps between polls.
"""Route a backlog of support tickets with a DecisionNode batch job.
backlog.jsonl holds one ticket per line: {"id": "...", "text": "..."}
"""
import json
import os
import time
import requests
API = "https://api.decisionnode.com/v1"
AUTH = {"Authorization": f"Bearer {os.environ['DECISIONNODE_API_KEY']}"}
MAX_LINES = 10000 # requests per batch
MAX_ACTIVE = 10 # active batches per key
POLL_EVERY = 300 # seconds between polls: results take hours
QUESTIONS = {
"route": {
"type": "choice",
"instructions": "Where should this go?",
"criteria": {"billing": "money", "bug": "broken", "account": "login"},
},
"refund": {"type": "truth", "instructions": "Refund this automatically?"},
}
def retryable(r):
"""408, a rate limit, 529 and any 5xx are worth another try; nothing else is."""
if r.status_code == 429:
return r.json()["detail"]["error_type"] == "rate_limit_error"
return r.status_code in (408, 529) or r.status_code >= 500
def call(method, path, headers=None, accept=(), before_retry=None, **kwargs):
"""One API call. Retries a network fault, 408, a rate limit, 529 and any 5xx
with a growing wait (about 8 minutes in all), returns a status in
`accept`, and raises on any other error."""
for attempt in range(10):
if attempt and before_retry:
before_retry()
try:
r = requests.request(
method,
f"{API}{path}",
headers={**AUTH, **(headers or {})},
timeout=300,
**kwargs,
)
except (requests.ConnectionError, requests.Timeout):
if attempt == 9:
raise
time.sleep(min(2**attempt, 300)) # the network is down or slow: retried
continue
if r.ok or r.status_code in accept:
return r
if not retryable(r) or attempt == 9:
r.raise_for_status()
wait = float(r.headers.get("Retry-After", 0))
time.sleep(max(wait, min(2**attempt, 300)))
def cancel_drafts():
"""Cancels this key's drafts. The program finalizes every batch it creates at
once, so a draft is one a failed create left behind (run one copy per key)."""
for batch in call("GET", "/batches", params={"limit": 100}).json()["batches"]:
if batch["state"] == "draft":
call("POST", f"/batches/{batch['id']}/cancel", accept=(409,))
def submit(lines):
"""Create a batch from request lines and finalize it."""
body = "\n".join(json.dumps(line) for line in lines)
# a 5xx or a network fault can come after the batch was created: clear the
# draft it may have left before creating again, so none counts against your limits
batch = call(
"POST",
"/batches",
data=body.encode(),
headers={"Content-Type": "application/x-ndjson"},
before_retry=cancel_drafts,
).json()
r = call("POST", f"/batches/{batch['id']}/finalize", accept=(409,))
if r.status_code == 409: # a retried finalize that had gone through
r = call("GET", f"/batches/{batch['id']}")
return r.json()
def wait(batch):
"""Poll until the batch has ended."""
while batch["state"] in ("queued", "running"):
time.sleep(POLL_EVERY)
batch = call("GET", f"/batches/{batch['id']}").json()
return batch
def results(batch):
"""The result lines, in upload order, a page at a time: a fault costs one page."""
offset = 0
while offset is not None:
page = call(
"GET",
f"/batches/{batch['id']}/results",
params={"format": "json", "offset": offset, "limit": 1000},
).json()
yield from page["results"]
offset = page["next_offset"]
def act(custom_id, answers):
route, refund = answers["route"], answers["refund"]
print(custom_id, route["choice"], route["confidence"], refund["truth"])
def run(tickets):
# keep your own copy: the API deletes the requests once results exist
lines = {
t["id"]: {
"custom_id": t["id"], # unique in the batch, 1 to 512 characters
"model": "decisionnode-latest",
"state": t["text"],
"questions": QUESTIONS,
}
for t in tickets
}
todo = list(lines)
for _ in range(3): # the first run, then what expired, at most twice
chunks = [todo[i : i + MAX_LINES] for i in range(0, len(todo), MAX_LINES)]
retry = []
for g in range(0, len(chunks), MAX_ACTIVE): # within your active batches
group = chunks[g : g + MAX_ACTIVE]
batches = [submit([lines[c] for c in chunk]) for chunk in group]
for batch in batches:
for line in results(wait(batch)):
if line["status"] == 200:
act(line["custom_id"], line["body"]["answers"])
elif line["status"] in ("expired", 500):
retry.append(line["custom_id"]) # not billed: send again
else:
# the live error status and body, as /v1/decide sends them
print("failed", line["custom_id"], line["status"], line["body"])
if not retry:
return
todo = retry
if __name__ == "__main__":
with open("backlog.jsonl") as f:
run([json.loads(line) for line in f if line.strip()])Billing#
batch = sum(input_tokens of completed lines)
× live price × 0.5
- input_tokens
- the
usage.input_tokensof each line whose status is 200, exactly as the live call counts them - live price
- the price per input token of the line's
model: $0.042 per million on DecisionNode-1.0, $0.021 (provisional) on DecisionNode-1.0 Flash
| Model | Live | Batch | Output |
|---|---|---|---|
| DecisionNode-1.0 | $0.042 | $0.021 | Free |
| DecisionNode-1.0 Flash | $0.021 (provisional) | $0.0105 (provisional) | Free |
Failed, expired and cancelled lines cost nothing, and a line the safety check refuses is not charged, as on the live call. For scale: 1,000,000 tickets of 47 tokens each cost $1.97 live on DecisionNode-1.0 and $0.9870 as batches. The batch object's usage shows what a batch has billed so far.
Cancelling#
POST /v1/batches/{id}/cancel stops a batch. A draft, or a batch none of whose requests has started, is cancelled at once. Otherwise requests that have not started are cancelled and never billed, the ones already running finish and are billed, and the batch shows queued or running until they are done, then cancelled. Its results hold every answer that finished. See Cancel a batch.
What to watch out for#
- Keep your requests. The API deletes a batch's requests once its results are written, and the results after 7 days. Store each request under its
custom_iduntil you have acted on its result. - Read every line's status. A
completedbatch can hold400,403or500lines. Each carries the live status and body, so the same error handling as your live calls applies. - Plan for expiry. Requests still unfinished 24 hours after a batch is finalized expire: they are not billed, and you can resubmit them. The program above does it for you.
- Keep credit in the balance. Lines run against your prepaid balance like live calls: a line answered while the balance is empty gets the live
402. Turn on auto reload before a large batch. - Stay within your key's limits. 10 active batches and 100,000 requests across them, drafts included in both. A draft you will not finalize: cancel it.
- Use your own idempotency. A
5xxor a timeout on a create can come after the batch was created. Before you create again, list your batches and cancel the draft it left: a draft never runs and costs nothing, but it counts toward your active batches until you cancel it. The program above does this. - Save batch ids before you wait. A program that stops while it waits and starts again from the top submits every request again, and each completed request is billed again. Write each batch id to disk once it is finalized, and on a restart poll the saved batches instead of submitting.
- Batches do not count against your rate limits; the calls that create, poll and download them do.