- What you bring
- 200 to 50,000 labelled decisions: requests as you send them, plus the answers your policy gives
- Where
- The console, under Fine-tuning. The API sees only the model's name
- Bases
- DecisionNode-1.0 (
decisionnode-1.0) and DecisionNode-1.0 Flash (decisionnode-1.0-flash) - What you get
- A model named
decisionnode-1.0-flash@acme/, private to your workspace, with the same question types, answer shapes and limits as its basefraud - Quality gate
- A version ships only if it beats the base on your held-out records and does not get worse elsewhere
- Price
- 3 × the run's compute time + $50 a run; $50 only when the gate says no; the compute so far, without the $50, when you cancel; nothing when a run fails or is refused. Requests: the base's price + $0.02 per million input tokens
When to fine-tune#
Fine-tune when the base answers your questions well in general but not the way your organisation decides: your threshold for "risky", the words your customers use, a label convention that lives in your team's heads rather than in your instructions. A fine-tuned model learns that judgment from decisions you have already made.
| The gap | Try first | Why |
|---|---|---|
| The answer is wrong because the question is vague | Sharper instructions and criteria | Free, immediate, and the model reads them on every request. See Questions |
| The model misses something in a picture: an age, a generated image | A booster | A specialist reading added as evidence, from the next request on |
| The answers are sensible but not your policy, and you have a few hundred decisions labelled the same way | Fine-tuning | It learns your judgment from your labels, and the gate proves the gain before it ships |
- Measure first. Keep a set of labelled requests and score the base on it. Without that number you cannot tell whether fine-tuning helped; the gate does this for you on the held-out records, but your own set tells you what to expect.
- Fix the questions first. A fine-tuned model answers the questions it was trained on. If you change the instructions or the options later, train again.
- Label consistently. Two people deciding the same record should give the same answer. Records that disagree with each other teach the model to hedge; validation counts them for you.
When not to#
- To teach new perception. Fine-tuning changes judgment, not eyesight: a reading the base cannot make from a picture is a job for a booster.
- To change the answer shapes. Question types, answers, calibration, limits and errors stay those of the base.
- With fewer than 200 labelled records. The run needs at least 50 records to calibrate and 50 to judge the result, and more to learn from.
- To make text. DecisionNode answers typed questions and generates no text, fine-tuned or not.
How it works#
- 1
Prepare a dataset#
JSON Lines, one record per line: a request as the API takes it, plus
answers. 200 to 50,000 records, written from your own system or labelled in the console from requests you kept. See Prepare your dataset. - 2
Create a model and upload#
In the console, name the model, pick its base, and upload the file. Validation checks every record before anything is trained. See Upload and validation.
- 3
Start a run#
The console shows the estimate first. The run trains, calibrates and scores a new version of the base. See Start a training run and Watch a run.
- 4
Pass the gate#
The version is compared with the base on records it never saw. It becomes
readyonly if it is better and not worse elsewhere. See The quality gate. - 5
Deploy and call it#
Deploy a
readyversion and send its name asmodel:decisionnode-1.0-flash@acme/. See Use your fine-tuned model.fraud
What stays the same#
| Fine-tuned model | |
|---|---|
| Endpoint and request | POST /v1/decide, the same body; only model changes |
| Question types and answers | The same types and shapes, rounding included |
| Calibration | Measured again for your version, on records the training never saw |
| Safety check | The same check, outside the model: it refuses exactly what the base refuses |
| Limits and errors | Those of the base |
| Boosters | Work as on the base |
| Determinism | A pinned version gives the same answer to the same request, every time |
Your data#
- A fine-tuned model belongs to one workspace. Only that workspace's keys can call it; to any other key its name does not exist.
- Your records train your model only. They are never used to train anything else, and the base is never updated from them.
- They stay in your workspace's storage location, the same one your uploads use, so your data-residency choice holds.
- Datasets are deleted 30 days after their last run unless you keep them, and whenever you delete them. The scorecard holds aggregate numbers only, never a record.
- Requests are kept for labelling only if you turn it on, and then for 30 days, encrypted, in your workspace's storage location; turning it off deletes them. See Label your history.