{"state":{"claim":{"merchant":"Nordic Rail","total":"740.00","currency":"SEK","receipt_total":"740.00","prior_claims_90d":0}},"questions":{"fraud":{"type":"truth","instructions":"Is this claim fraudulent?"},"route":{"type":"choice","instructions":"What should happen to this claim?","criteria":{"pay":"pay the claim in full","partial":"pay only the part the receipt documents","decline":"refuse the claim"}},"risk":{"type":"score","instructions":"How risky is this claim?","criteria":["none","low","medium","high"]}},"answers":{"fraud":0,"route":"pay","risk":"none"}}
{"state":{"claim":{"merchant":"Hotel Lindgren","total":"1980.00","currency":"SEK","receipt_total":"980.00","prior_claims_90d":3}},"questions":{"fraud":{"type":"truth","instructions":"Is this claim fraudulent?"},"route":{"type":"choice","instructions":"What should happen to this claim?","criteria":{"pay":"pay the claim in full","partial":"pay only the part the receipt documents","decline":"refuse the claim"}},"risk":{"type":"score","instructions":"How risky is this claim?","criteria":["none","low","medium","high"]}},"answers":{"fraud":1,"route":"decline","risk":3}}
{"state":{"claim":{"merchant":"Taxi Stockholm","total":"512.40","currency":"SEK","receipt_total":"312.40","prior_claims_90d":1}},"questions":{"fraud":{"type":"truth","instructions":"Is this claim fraudulent?"},"route":{"type":"choice","instructions":"What should happen to this claim?","criteria":{"pay":"pay the claim in full","partial":"pay only the part the receipt documents","decline":"refuse the claim"}},"risk":{"type":"score","instructions":"How risky is this claim?","criteria":["none","low","medium","high"]}},"answers":{"fraud":0,"route":"partial","risk":"medium"}}Three records of a claims dataset. Every record asks the same three questions about a different claim, and answers holds what your team decided. Write the questions exactly as your service will send them: the model learns your answers to those questions, in those words.
A record#
{
"state": {
"claim": {
"merchant": "Nordic Rail",
"total": "740.00",
"currency": "SEK",
"receipt_total": "740.00",
"prior_claims_90d": 0
}
},
"images": [
{ "id": "receipt", "upload_id": "upl_3f6c1a9e0b2d4c7f8a1e5d9b2c4f6a80" }
],
"questions": {
"fraud": { "type": "truth", "instructions": "Is this claim fraudulent?" },
"route": {
"type": "choice",
"instructions": "What should happen to this claim?",
"criteria": {
"pay": "pay the claim in full",
"partial": "pay only the part the receipt documents",
"decline": "refuse the claim"
}
},
"risk": {
"type": "score",
"instructions": "How risky is this claim?",
"criteria": ["none", "low", "medium", "high"]
},
"matches": {
"type": "truth",
"instructions": "Does the total on `receipt` equal the claim's total?"
}
},
"answers": { "fraud": 0, "route": "pay", "risk": "none", "matches": 1 },
"weight": 2,
"split": "held_out"
}statestring | object | array- The input, as in a request. Leave it out for questions that carry everything.
imagesarray- Up to 16 images, each a live
upload_idor inline base64datawith itsmedia_type; an image byurlis refused. See Images in records. The sameidrules as in a request. videosarray- Videos as frames with timestamps, as in a request.
questionsobjectrequired- Your question ids mapped to questions, exactly as in a request:
choice,score,truthandnumber.instructionsmay be left out; the type's default instruction is used. Point and box questions are not trained. answersobjectrequired- One answer per question, keyed by the question's id.
choicestring- One of the question's option names, exactly as written in
criteria. scoreinteger | string- A level's index from 0, lowest first, or the level itself as written in
criteria. truthnumber1for true,0for false, or a probability in between when your own label is not certain.numbernumber- A value on the question's grid: from
mintomaxin steps ofstep, as a number question in a request.
weightnumberdefault1- How much the record counts, from 0.1 to 10. Raise it for records that matter more; never copy a record to give it weight, since duplicates are dropped.
split"train" | "calibration" | "held_out"- Where the record goes. Leave it out and the run deals the records itself, which is right for almost every dataset. See Splits.
- One JSON object per line, UTF-8, no line breaks inside a record. Blank lines are not records.
- The request rules apply to every record: the same types, option and level limits, token limit and nesting as on
/v1/decide. - Mixed questions are fine. Records can ask different questions; each question id keeps its meaning across records.
Write it from your own decisions#
Most datasets come from decisions your team already made: a claims system's outcomes, a moderation log, tickets and where they were routed. Read them out of your system, keep the questions fixed, and write one record per decision:
import json
QUESTIONS = {
"fraud": {
"type": "truth",
"instructions": "Is this claim fraudulent?"
},
"route": {
"type": "choice",
"instructions": "What should happen to this claim?",
"criteria": {
"pay": "pay the claim in full",
"partial": "pay only the part the receipt documents",
"decline": "refuse the claim"
}
},
"risk": {
"type": "score",
"instructions": "How risky is this claim?",
"criteria": ["none", "low", "medium", "high"]
}
}
def record(claim, outcome):
"""One labelled decision: the request as you send it, plus the answers."""
return {
"state": {"claim": claim},
"questions": QUESTIONS,
"answers": {
"fraud": 1 if outcome["fraud"] else 0,
"route": outcome["route"], # one of the criteria keys
"risk": outcome["risk"], # a level name, or its index from 0
},
}
with open("claims-2026-09.jsonl", "w", encoding="utf-8") as out:
for claim, outcome in labelled_claims(): # your own past decisions
out.write(json.dumps(record(claim, outcome), ensure_ascii=False) + "\n")Label your history#
If your own system does not keep its decisions, DecisionNode can keep them for you, so you can label them in the console. An Owner or Admin turns on Keep requests for labelling for the workspace. From then on, every /v1/decide request answered with 200 is kept with its answer for 30 days, encrypted, in your workspace's storage location. Nothing is kept while the setting is off, and nothing before it was turned on.
- 1
Turn on keeping#
In the console, an Owner or Admin turns on Keep requests for labelling. It takes effect within a second. A workspace with zero retention cannot turn it on.
- 2
Let traffic run#
Send requests as usual. Each answered
/v1/deciderequest is kept for 30 days from the moment it was answered, with its images: uploaded images are copied with it and inline images are in its body. Images sent byurlare not kept, so a request whose image came by URL cannot become a training record: send images you may want to label as uploads. Batch lines and session frames are not kept. - 3
Label in the console#
Under Fine-tuning, list the kept requests by date, model or key, open one with the answer the model gave, and record the answer your policy gives to each question. Make them a dataset turns the labelled requests into records.
- 4
Upload the labelled set#
The labelled requests become records in the format above and go up as a dataset, through the same upload and validation as a file you wrote yourself.
- Only your workspace sees them. Kept requests are shown only to its members (Owners, Admins and Developers can label), never leave its storage location, and are used only to build its own datasets.
- They are deleted 30 days after each request, all at once when the setting is turned off, when zero retention is turned on, and with the workspace.
- Zero retention and keeping are exclusive. A zero-retention workspace keeps nothing; upload records you kept yourself.
What makes a good record#
| Good | Bad | Why it matters |
|---|---|---|
| The question is worded as your service will send it | Questions reworded for the dataset | The model learns your answers to those exact words |
| The state holds what the decider saw | The answer depended on something outside the record, such as a phone call | The model cannot learn a reason it never sees; it learns to guess |
| The same input always has the same answer | Two people labelled similar claims differently | Conflicting records teach the model to hedge; validation counts them |
| Every option and level appears, in proportion to real traffic | Nearly every record answered pay | A model shown one answer learns to give it |
Weight set with weight | A record copied ten times | Copies are dropped as duplicates |
| A probability in a truth answer only where your own label was unsure | Every truth answer written as 0.7 | The answer is what the model learns to say |
How many records#
A dataset holds 200 to 50,000 records. Start with a few hundred labelled consistently; more records give the run more to learn from and give the gate a sharper measure. The run sets two parts aside before it trains: 10% of the records, at least 50, to calibrate the probabilities, and the same again, held out, to judge the result. So the smallest dataset trains on half its records:
| Records | Train | Calibration | Held out | One held-out record, in accuracy |
|---|---|---|---|---|
| 200 | 100 | 50 | 50 | 0.020 |
| 500 | 400 | 50 | 50 | 0.020 |
| 2,000 | 1,600 | 200 | 200 | 0.005 |
| 10,000 | 8,000 | 1,000 | 1,000 | 0.0010 |
| 50,000 | 40,000 | 5,000 | 5,000 | 0.0002 |
The gate asks for a held-out accuracy at least 0.02 above the base's, and a gain whose interval lies clear of zero. With 50 held-out records, one record is 0.02 of accuracy, so a small dataset has to win by whole records; with a few thousand, a real but small gain can show.
Splits#
- train: what the model learns from.
- calibration: used to stop training at the right time and to calibrate the new version's probabilities, so a 0.80 still means 8 times in 10.
- held_out: never seen until the end; the gate scores your version and the base on it.
- Records with the same state go to the same split, so a near-copy in training cannot inflate the held-out score.
- Set
splityourself only when you need a fixed test set, such as last month's decisions held out. Keep records with the same state in the same split.
Images in records#
A record carries each image in one of two forms, as a request does; an image by url is refused, since the training run cannot fetch your address later.
| Form | When to use it | What to know |
|---|---|---|
upload_id | Most datasets, and any image over a few hundred KB | Put the image in your workspace's storage with POST /v1/uploads (up to 16 files a call, each JPEG, PNG or WebP of up to 10 MB) and write its upload_id into the record. An upload can be named for one hour, and validation holds it for the dataset's life: upload the images, write the file, upload it and start validation within that hour |
data with media_type | A few small images | The base64 bytes inside the record, as in a request. They count toward the dataset file's 64 MiB, about a third larger than the image |
import base64, hashlib, os, pathlib, requests
API = "https://api.decisionnode.com"
KEY = os.environ["DECISIONNODE_API_KEY"]
TYPES = {".jpg": "image/jpeg", ".jpeg": "image/jpeg", ".png": "image/png", ".webp": "image/webp"}
def upload(path):
"""Puts one image in your workspace's storage; returns its upload_id."""
raw = pathlib.Path(path).read_bytes()
md5 = base64.b64encode(hashlib.md5(raw).digest()).decode()
r = requests.post(
f"{API}/v1/uploads",
headers={"Authorization": f"Bearer {KEY}"},
json={"files": [{
"content_type": TYPES[pathlib.Path(path).suffix.lower()],
"byte_size": len(raw),
"checksum": md5,
}]},
timeout=10,
)
r.raise_for_status()
slot = r.json()["uploads"][0]
# the PUT goes straight to storage, with exactly the headers given
requests.put(slot["url"], data=raw, headers=slot["headers"], timeout=60).raise_for_status()
return slot["upload_id"]
image = {"id": "receipt", "upload_id": upload("receipts/4411.jpg")}Check the file before you upload#
Validation checks everything again, but a local check saves a round trip on the common mistakes: the record count, the file size, one answer per question, a choice answer that is not an option, a score answer out of range, a truth answer outside 0 to 1.
import json, os, sys
path = sys.argv[1]
problems = []
with open(path, encoding="utf-8") as f:
lines = [line for line in f if line.strip()]
if not 200 <= len(lines) <= 50_000:
problems.append(f"{len(lines)} records: a dataset holds 200 to 50,000")
if os.path.getsize(path) > 64 * 1024 * 1024:
problems.append("the file is over 64 MiB")
for n, line in enumerate(lines, start=1):
r = json.loads(line)
questions, answers = r["questions"], r.get("answers", {})
if set(answers) != set(questions):
problems.append(f"record {n}: one answer per question, keyed by its id")
continue
for qid, q in questions.items():
a = answers[qid]
if q["type"] == "choice" and a not in q["criteria"]:
problems.append(f"record {n}: {qid} = {a!r} is not an option")
if q["type"] == "score" and not (
a in q["criteria"] or (isinstance(a, int) and 0 <= a < len(q["criteria"]))
):
problems.append(f"record {n}: {qid} = {a!r} is not a level or its index")
if q["type"] == "truth" and not (isinstance(a, (int, float)) and 0 <= a <= 1):
problems.append(f"record {n}: {qid} = {a!r} is not 0, 1 or a probability")
if q["type"] in ("point", "box"):
problems.append(f"record {n}: {qid}: point and box questions are not trained")
print("\n".join(problems[:100]) or f"{len(lines)} records look right")