Skip to content
DecisionNodeDecisionNdedocs
  • Guides
  • API reference
  • Examples
  • Playground

start here

  • QuickstartGet startedGet an API key, send one request with three questions, and branch your code on the typed answers. Plain HTTPS, no SDK to install.
  • POST /v1/decideAPI referenceAnswer typed questions about a state, optional images and video frames. One request, one buffered JSON response, one answer per question.
  • QuestionsConceptsQuestions say what to decide. Each one has a type that fixes the shape of its answer: a choice from your options, a score on your scale, a…
  • ConfidenceConceptsProbabilities are calibrated per question type, so a threshold means what it says on the data we measure.
  • ImagesConceptsSend up to 16 images and text in the same request. The model reads printed and handwritten text, amounts, dates, objects and layout, and…
  • Batch jobsPatternsSend up to 10,000 requests in one file and collect the answers later, at a lower price than live calls.
  • Pricing and billingYou pay for input tokens only. Output is free because the model generates no text.
↑↓ moveopen7 suggestions
Get API keyGet API key
DecisionNodeDecisionNde

Get started

  • Introduction
  • Quickstart
  • Playground
  • Console and keys
  • With coding agents
  • MCP server
  • Examples

Concepts

  • State
  • Questions
  • Choice
  • Score
  • Truth
  • Number
  • Points and boxes
  • Images
  • Video
  • Boosters
  • Confidence
  • Determinism

Models

  • DecisionNode-1.0
  • DecisionNode-1.0 Flash
  • Limits
  • Versions
  • Dedicated capacity

Fine-tuning

  • Overview
  • Prepare your dataset
  • Upload and validation
  • Start a training run
  • Watch a run
  • The quality gate
  • Use your model
  • Limits and pricing

Patterns

  • Confidence-gated routing
  • Fan-out
  • Guardrails
  • Control loops
  • Batch jobs

API reference

  • Overview
  • POST/v1/decide
  • GET/v1/models
  • POST/v1/uploads
  • Errors
  • Safety check
  • Rate limits

Sessions API

  • Sessions overview
  • POSTOpen a session
  • WSStream frames
  • DELEnd a session

Batch API

  • The batch object
  • POSTCreate a batch
  • POSTAdd requests
  • POSTFinalize a batch
  • GETRetrieve a batch
  • GETGet batch results
  • POSTCancel a batch
  • GETList batches

Pricing and billing

  • Pricing and billing
  • Refer & earn

Policies

  • Responsible use
  • Data and privacy
  • Benchmarks
  • Pricing
  • Playground
Get API key
  • Guides
  • API reference
  • Examples
  • Playground

Get started

  • Introduction
  • Quickstart
  • Playground
  • Console and keys
  • With coding agents
  • MCP server
  • Examples

Concepts

  • State
  • Questions
  • Choice
  • Score
  • Truth
  • Number
  • Points and boxes
  • Images
  • Video
  • Boosters
  • Confidence
  • Determinism

Models

  • DecisionNode-1.0
  • DecisionNode-1.0 Flash
  • Limits
  • Versions
  • Dedicated capacity

Fine-tuning

  • Overview
  • Prepare your dataset
  • Upload and validation
  • Start a training run
  • Watch a run
  • The quality gate
  • Use your model
  • Limits and pricing

Patterns

  • Confidence-gated routing
  • Fan-out
  • Guardrails
  • Control loops
  • Batch jobs

API reference

  • Overview
  • POST/v1/decide
  • GET/v1/models
  • POST/v1/uploads
  • Errors
  • Safety check
  • Rate limits

Sessions API

  • Sessions overview
  • POSTOpen a session
  • WSStream frames
  • DELEnd a session

Batch API

  • The batch object
  • POSTCreate a batch
  • POSTAdd requests
  • POSTFinalize a batch
  • GETRetrieve a batch
  • GETGet batch results
  • POSTCancel a batch
  • GETList batches

Pricing and billing

  • Pricing and billing
  • Refer & earn

Policies

  • Responsible use
  • Data and privacy
  1. docs
  2. /
  3. Get started

The fine-tuning quality gate

A fine-tuned version ships only when it helps. Four tests compare it with its base: held-out accuracy at least 0.02 better and clear of chance, calibration no more than 0.01 worse, our standard test sets no more than 0.01 lower, and the same refusals. Pass all four and it is ready; miss one and it is no_gain, which costs $50.

on this page3 sections
  1. The four tests
  2. The scorecard
  3. When the gate says no_gain

The four tests#

  • Held-out accuracy

    Passes when
    Both: your version's accuracy on the held-out records is at least 0.02 above the base's, and the interval of the difference lies wholly above 0
    Why
    It must be better on your decisions, by more than chance: random labels cannot pass
  • Held-out calibration

    Passes when
    Its calibration error is at most 0.01 above the base's
    Why
    A 0.80 must still mean true about 8 times in 10
  • Standard test sets

    Passes when
    Its mean score on our test sets for the same question types is no more than 0.01 below the base's; doing better never fails
    Why
    It must not forget what the base does well
  • Refusals

    Passes when
    Identical to the base's
    Why
    The safety check sits outside the model, so this always holds; it is checked anyway
Your version against its base
TestPasses whenWhy
Held-out accuracyBoth: your version's accuracy on the held-out records is at least 0.02 above the base's, and the interval of the difference lies wholly above 0It must be better on your decisions, by more than chance: random labels cannot pass
Held-out calibrationIts calibration error is at most 0.01 above the base'sA 0.80 must still mean true about 8 times in 10
Standard test setsIts mean score on our test sets for the same question types is no more than 0.01 below the base's; doing better never failsIt must not forget what the base does well
RefusalsIdentical to the base'sThe safety check sits outside the model, so this always holds; it is checked anyway

The held-out records are the ones the run set aside and never trained on, so the comparison is fair to both. Calibration error is how far a model's stated probabilities are from how often they come true; lower is better.

The scorecard#

  • Held out

    Numbers
    The number of records; accuracy of your version and the base, the difference and its interval; the calibration error of each
  • By question

    Numbers
    Each question id: its type, its records, and the accuracy of your version and the base
  • Standard test sets

    Numbers
    Each set by name and question type, your version against the base
  • Gate

    Numbers
    Each test, passed or not, with its value and its threshold; and the reason when the version did not pass
What the scorecard shows
PartNumbers
Held outThe number of records; accuracy of your version and the base, the difference and its interval; the calibration error of each
By questionEach question id: its type, its records, and the accuracy of your version and the base
Standard test setsEach set by name and question type, your version against the base
GateEach test, passed or not, with its value and its threshold; and the reason when the version did not pass

The scorecard holds aggregate numbers only, never a record. Read it by question: a version can gain a lot on one question and nothing on another, which tells you where your labels carry a policy the base did not know.

When the gate says no_gain#

no_gain means the version was trained and scored, and the base is still the better choice on your data, for example "Your held-out accuracy 0.81 against the base's 0.80: not enough to ship." The version keeps its scorecard but no trained weights, so it can never be deployed; the run costs $50, and the compute part is not charged.

  • Accuracy close to the base's on every question

    Likely cause
    The base already decides the way you do
    Try
    Keep the base and save the cost; sharpen instructions where it errs
  • Accuracy up on some questions, down on others

    Likely cause
    Some questions have inconsistent labels
    Try
    Check the conflicts count and relabel those questions
  • A gain too small to pass, on a small held-out split

    Likely cause
    One held-out record is 0.02 of accuracy at 50 records
    Try
    Add records: a larger held-out split measures a real gain
  • Calibration worse than the base's

    Likely cause
    Truth answers written as guesses, or heavy weights on few records
    Try
    Use 0 and 1 where your label is certain; even out the weights
  • Standard test sets lower

    Likely cause
    The dataset is narrow and repetitive
    Try
    Add variety: more states, more options answered
What a no_gain usually means
The scorecard showsLikely causeTry
Accuracy close to the base's on every questionThe base already decides the way you doKeep the base and save the cost; sharpen instructions where it errs
Accuracy up on some questions, down on othersSome questions have inconsistent labelsCheck the conflicts count and relabel those questions
A gain too small to pass, on a small held-out splitOne held-out record is 0.02 of accuracy at 50 recordsAdd records: a larger held-out split measures a real gain
Calibration worse than the base'sTruth answers written as guesses, or heavy weights on few recordsUse 0 and 1 where your label is certain; even out the weights
Standard test sets lowerThe dataset is narrow and repetitiveAdd variety: more states, more options answered
  • Use your fine-tuned modelDeploy it and call it by name.Read
  • Prepare your datasetThe record format, every rule, and how many records you need.Read
nextIntroduction

DecisionNode is built and run by Bynn Intelligence, Inc.

  • Home
  • Playground
  • Examples
  • Console
  • Responsible use
  • Terms
  • Acceptable use
  • Privacy
  • Data processing
  • Defence addendum
  • Cookies
  • Affiliate program

on this page

  1. The four tests
  2. The scorecard
  3. When the gate says no_gain