Create a model#
In the console, open Fine-tuning and choose New fine-tuned model: a name, a base, and an optional description of up to 200 characters. The name is 3 to 40 lower-case letters, digits and hyphens, unique in the workspace, and it becomes part of what you send as model: fraud on DecisionNode-1.0 Flash in the workspace acme is decisionnode-1.0-flash@acme/. Pick the base by what you need: DecisionNode-1.0 for accuracy, DecisionNode-1.0 Flash for speed and price. A model keeps its base; to try the other, create a second model.
Upload a dataset#
Choose Create and add a dataset, give it a name (1 to 80 characters) and Choose a file. The browser sends it straight into your workspace's storage location, so a 64 MiB file never passes through anything else. Validation starts when the upload is complete.
What validation checks#
| Check | What happens |
|---|---|
| Format and limits | Every line is a record with the fields of a record, within the request limits; 200 to 50,000 records |
| Answers | One per question, each a valid answer for its type |
| Images | Every image is a live upload_id of your workspace or inline data; an image by url is refused. From then on the dataset keeps its images for its life |
| Duplicates | Records with the same state, questions and answers are dropped, and counted |
| Conflicts | Records with the same input but different answers are kept, and counted, so you can fix them |
| Splits | 10% of the records, at least 50, for calibration and the same for held out; the rest trains |
| Safety check | Not here: it runs as the first step of a training run, on every record's state and images (see Start a training run) |
The result#
| State | Meaning | Next |
|---|---|---|
uploading | The file is on its way | Wait |
validating | The checks are running; the console shows it as checking | Wait; the console shows the result when they end |
ready | Every record passed | Start a run |
invalid | Records broke the format or the limits | Fix the issues listed and upload again |
refused | A training run's safety check refused records; the run ended refused, at no charge | See the refused records, remove them, upload again and start a new run |
A validated dataset shows its record count, the three splits, the duplicates dropped, the conflicts, how many images it holds, and how many questions of each type it asks. An invalid dataset lists its issues, the first 100 of them, each with the record's line number, the path to the field (such as answers.risk) and what is wrong; the total count is shown beside them. A refused dataset, marked by a training run's safety check, names the record numbers only, never their content.
| Issue | Fix |
|---|---|
| A choice answer that is not an option | Use an option name exactly as written in criteria, case and spaces included |
| A score answer that is neither a level nor its index | Use the level's index from 0, or the level as written |
| An answer for a question the record does not ask, or a question without an answer | One answer per question id |
An expired or unknown upload_id | Upload the image again and validate within the hour |
An image by url | Send it as an upload, or inline as data |
| Fewer than 200 records after duplicates are dropped | Add records; copies do not count |
| Many conflicts | Decide which answer is your policy and relabel; conflicts are kept, but they cost accuracy |
Keeping and deleting datasets#
- A dataset is deleted 30 days after its last run unless you choose Keep on it (Stop keeping undoes it).
- Delete a dataset at any time, except while a queued or running run uses it.
- A dataset belongs to the model it was uploaded for. To train another model on the same records, upload them for that model.
- Deleting a model deletes its datasets with it, and cancels a run that is going on the terms of a cancel.