States#
| State | Meaning | Charged |
|---|---|---|
queued | Waiting to start | Nothing yet |
validating | Checking the dataset and setting up; the console shows it as preparing | Nothing yet |
training | Learning from the training records | Nothing yet |
evaluating | Calibrating and scoring against the base | Nothing yet |
ready | Passed the gate; can be deployed | The full fee |
no_gain | Did not pass the gate; the scorecard says why; the console shows it as no gain | $50 |
failed | Stopped; the reason is shown | Nothing |
refused | The safety check refused records before training; the dataset lists their numbers. The console shows it as refused | Nothing |
cancelled | Cancelled by a member, or by deleting its model, before it ended | The compute used so far, without the $50; nothing if it had not started |
Live progress#
| Reading | What it tells you |
|---|---|
| Step and total steps | How far through training the run is |
| Epoch | How many passes over the training records so far, such as 1.4 |
| Training loss | How wrong the version still is on the records it learns from; lower is better |
| Evaluation so far | Scores measured during the run, where there are any yet |
| Tokens per second | How fast the run reads your records |
| Elapsed | Time since the run started, beside the estimate it started with |
The loss falls fast at first and then flattens. A loss that keeps falling while the evaluation stops improving means the version is learning the training records by heart; the run watches the calibration records for exactly that and stops in time. The loss does not decide whether the version ships: the gate does, on records the run never trained on.
When a run fails#
| Cause | What happens | Charged |
|---|---|---|
| The machine running it is lost | The run is started again once, automatically, on a new machine | Only the attempt that finishes; the lost one is free |
| Training breaks down, such as a loss that grows without bound | failed, with the reason; it is not retried | Nothing |
| Still running after 6 hours | Stopped, failed | Nothing |
Cancel a run#
An Owner, Admin or Developer can cancel a run while it is queued, validating, training or evaluating: Cancel run, then confirm with Cancel run (or Keep training to let it go on). It stops at once, nothing is published, and the version ends cancelled. You pay for the compute it used so far at the usual rate, 3 × its compute time, without the $50; a run cancelled before it started costs nothing. Deleting a model while one of its runs is going cancels that run on the same terms. If DecisionNode stops a run, it ends cancelled the same way, you are told why, and the charge may be waived.
When a run ends, whichever way, the member who started it gets an email.