The four tests#
| Test | Passes when | Why |
|---|---|---|
| Held-out accuracy | Both: your version's accuracy on the held-out records is at least 0.02 above the base's, and the interval of the difference lies wholly above 0 | It must be better on your decisions, by more than chance: random labels cannot pass |
| Held-out calibration | Its calibration error is at most 0.01 above the base's | A 0.80 must still mean true about 8 times in 10 |
| Standard test sets | Its mean score on our test sets for the same question types is no more than 0.01 below the base's; doing better never fails | It must not forget what the base does well |
| Refusals | Identical to the base's | The safety check sits outside the model, so this always holds; it is checked anyway |
The held-out records are the ones the run set aside and never trained on, so the comparison is fair to both. Calibration error is how far a model's stated probabilities are from how often they come true; lower is better.
The scorecard#
| Part | Numbers |
|---|---|
| Held out | The number of records; accuracy of your version and the base, the difference and its interval; the calibration error of each |
| By question | Each question id: its type, its records, and the accuracy of your version and the base |
| Standard test sets | Each set by name and question type, your version against the base |
| Gate | Each test, passed or not, with its value and its threshold; and the reason when the version did not pass |
The scorecard holds aggregate numbers only, never a record. Read it by question: a version can gain a lot on one question and nothing on another, which tells you where your labels carry a policy the base did not know.
When the gate says no_gain#
no_gain means the version was trained and scored, and the base is still the better choice on your data, for example "Your held-out accuracy 0.81 against the base's 0.80: not enough to ship." The version keeps its scorecard but no trained weights, so it can never be deployed; the run costs $50, and the compute part is not charged.
| The scorecard shows | Likely cause | Try |
|---|---|---|
| Accuracy close to the base's on every question | The base already decides the way you do | Keep the base and save the cost; sharpen instructions where it errs |
| Accuracy up on some questions, down on others | Some questions have inconsistent labels | Check the conflicts count and relabel those questions |
| A gain too small to pass, on a small held-out split | One held-out record is 0.02 of accuracy at 50 records | Add records: a larger held-out split measures a real gain |
| Calibration worse than the base's | Truth answers written as guesses, or heavy weights on few records | Use 0 and 1 where your label is certain; even out the weights |
| Standard test sets lower | The dataset is narrow and repetitive | Add variety: more states, more options answered |