Start it#
- 1.Open the model in Fine-tuning, choose New training run and a
readydataset. - 2.See the estimate: the compute time the run should take, the expected wall time, and the fee.
- 3.Start training. The run gets the next version number, 1, 2, 3 and so on, and appears as
queued.
| Reason | What to do |
|---|---|
The dataset is not ready | Wait for validation, or fix it |
| The dataset belongs to another model | Upload it for this model |
| A run is already queued or running in the workspace (1 at a time) | Wait for it to end |
| The prepaid balance is below the estimate | Add credit under Billing, then start again |
The estimate and the fee#
fee = 3 × compute time × compute price
+ $50
- compute time
- how long the run trains and scores, measured when it ends
- compute price
- the price of the machine the run uses, recorded when it starts
| The run ends | Charged |
|---|---|
ready | The full fee: 3 × the compute + $50 |
no_gain | $50 only: the compute part is not charged |
failed | Nothing |
refused | Nothing: the safety check refused records before any training |
cancelled | The compute used so far × 3, without the $50; nothing if it had not started. See Cancel a run |
The estimate comes from the dataset's size and the base, and the fee is charged from the time the run really took, so the charge can differ a little from the estimate. For example, a run whose compute costs $0.70 is charged $52.10 when it ends ready and $50 when it ends no_gain. The fee is drawn from the prepaid balance; every figure is on Limits and pricing.
What a run does#
- 1
Check#
Runs the safety check over every record's state and images, as a request is checked, before any training. If it refuses a record, the run ends
refused: nothing is trained, nothing is charged, and the dataset is marked refused with the record numbers only, never their content. - 2
Split#
Deals the records into train, calibration and held out (10% each for the last two, at least 50), keeping records with the same state together.
- 3
Train#
Trains a new version of the base on the training records, checking itself against the calibration records as it goes and stopping when they stop improving.
- 4
Calibrate#
Fits the new version's probabilities on the calibration records, so its confidences mean what they say.
- 5
Score#
Scores the new version and the base side by side: on your held-out records, and on our standard test sets for the same question types.
- 6
Gate#
Passes the version as
ready, or marks itno_gainwith the reason. See The quality gate.
A run takes minutes for a small dataset on DecisionNode-1.0 Flash and longer for large datasets and for DecisionNode-1.0; the estimate says how long yours should take. A run is stopped after 6 hours and ends failed, at no charge.
Train again#
- Every run is a new version. Earlier versions stay as they are, deployed or not, so you can compare and roll back.
- A model keeps its 5 newest versions; older ones are archived.
- When the base gets a new release, your deployed versions stay on the base they were trained on and keep answering until that base is retired, announced in the console with a date. Train the same dataset again on the new base, and the gate decides again.