Skip to content
DecisionNodeDecisionNde
  • Model
  • Inference
  • Benchmarks
  • Examples
  • Docs
  • Pricing

api all systems normal

Get API keyGet API key
  • dnModelTyped answers, calibrated confidence
  • msInferenceOur own stack and GPUs, answers in ms
  • %BenchmarksAccuracy per suite, with intervals
  • exExamplesBuilds you can start from today
  • /v1DocsQuickstart, API reference, recipes
  • $PricingPay per input token, output is free
benchmarks

Measured. Wins and losses.

DecisionNode-⁠1.0 against TypeSafe Jev, Cloudflare Clef and four open decision models: 15 text suites, 25,134 items, every suite shown, including the ones we lose. Then two image suites, eight general LLMs on the same decisions, and our own latency.

Open playgroundOpen playgroundHow we measured↓
DecisionNode-⁠1.0, 15-suite mean
0.817
DecisionNode-1.0 Flash
0.780
Flash, short request
5 ms

Preliminary results, October 2026

The 15-suite mean

Accuracy averaged over the 15 text suites, one dot per system, with its 95% interval. DecisionNode-⁠1.0 scores 0.817 (0.809 to 0.825) and TypeSafe Jev 0.803 (0.795 to 0.811). DecisionNode-⁠1.0 Flash, the tier built for speed and price, scores 0.780.

  • our models
  • other systems
  • 95% interval

Preliminary results, October 2026

0.600.650.700.750.800.85
  1. 1DecisionNode-1.0Bynn Intelligenceours0.8170.809 to 0.825, 95% interval 0.809 to 0.825
  2. 2TypeSafe Jev 1.13TypeSafe AI0.8030.795 to 0.811, 95% interval 0.795 to 0.811
  3. 3Cloudflare Clef 27BCloudflare0.7860.778 to 0.794, 95% interval 0.778 to 0.794
  4. 4DecisionNode-1.0 FlashBynn Intelligenceours0.7800.771 to 0.789, 95% interval 0.771 to 0.789
  5. 5Cloudflare Clef-flash 9BCloudflare0.7590.750 to 0.767, 95% interval 0.750 to 0.767
  6. 6pplx-decider-27bopen model0.7460.738 to 0.755, 95% interval 0.738 to 0.755
  7. 7Cygnet (Gemma 4 12B)open model0.7180.710 to 0.726, 95% interval 0.710 to 0.726
  8. 8Kev-4Bopen model0.7170.708 to 0.726, 95% interval 0.708 to 0.726
  9. 9Strands Decider 2Bopen model0.6210.612 to 0.630, 95% interval 0.612 to 0.630

accuracy, 15 text suites

Where we lead, where we trail

DecisionNode-⁠1.0 minus TypeSafe Jev on every text suite. A bar is solid where the two 95% intervals do not overlap and hollow where the difference is inside the noise.

  • DecisionNode-1.0 ahead, intervals apart
  • TypeSafe Jev ahead, intervals apart
  • inside the noise
−15−7.50+7.5+15
  1. Rules and policies0.921 vs 0.831+9.0 pts, DecisionNode-1.0 0.921, TypeSafe Jev 0.831, intervals apart
  2. Web tasks0.781 vs 0.698+8.3 pts, DecisionNode-1.0 0.781, TypeSafe Jev 0.698, intervals apart
  3. Large option sets0.846 vs 0.775+7.1 pts, DecisionNode-1.0 0.846, TypeSafe Jev 0.775, intervals apart
  4. Short classifications0.955 vs 0.904+5.1 pts, DecisionNode-1.0 0.955, TypeSafe Jev 0.904, inside the noise
  5. Familiar task families0.812 vs 0.774+3.8 pts, DecisionNode-1.0 0.812, TypeSafe Jev 0.774, intervals apart
  6. Multilingual0.808 vs 0.782+2.6 pts, DecisionNode-1.0 0.808, TypeSafe Jev 0.782, inside the noise
  7. Public decision set0.872 vs 0.855+1.7 pts, DecisionNode-1.0 0.872, TypeSafe Jev 0.855, inside the noise
  8. Open decision tasks0.812 vs 0.798+1.4 pts, DecisionNode-1.0 0.812, TypeSafe Jev 0.798, inside the noise
  9. Held-out decisions0.958 vs 0.948+1.0 pts, DecisionNode-1.0 0.958, TypeSafe Jev 0.948, inside the noise
  10. Typed decisions0.741 vs 0.735+0.6 pts, DecisionNode-1.0 0.741, TypeSafe Jev 0.735, inside the noise
  11. Long inputs0.848 vs 0.853−0.5 pts, DecisionNode-1.0 0.848, TypeSafe Jev 0.853, inside the noise
  12. New task families0.812 vs 0.823−1.1 pts, DecisionNode-1.0 0.812, TypeSafe Jev 0.823, inside the noise
  13. Transfer tasks0.836 vs 0.852−1.6 pts, DecisionNode-1.0 0.836, TypeSafe Jev 0.852, inside the noise
  14. Prompt injection0.565 vs 0.600−3.5 pts, DecisionNode-1.0 0.565, TypeSafe Jev 0.600, inside the noise
  15. Knowledge (MMLU-Pro)0.690 vs 0.820−13.0 pts, DecisionNode-1.0 0.690, TypeSafe Jev 0.820, intervals apart

← TypeSafe Jev aheadDecisionNode-1.0 ahead →

difference in accuracy

Leads

Clear leads on rules and policies, web tasks and large option sets. That is the everyday work of a workflow: apply a written policy, act on a page, pick from many options.

Trails

Jev is clearly ahead on knowledge (−13.0 pts), and ahead by a margin inside the noise on prompt injection (−3.5 pts) and transfer tasks (−1.6 pts). If your decisions lean on broad world knowledge or on resisting adversarial text, weigh those suites first.

Every suite, every system

Accuracy per suite with the 95% interval under it. The highest score on each row is marked. Scroll sideways on a small screen.

Accuracy by suite and system, with 95% intervals
SuiteDecisionNode-⁠1.0Bynn IntelligenceDecisionNode-⁠1.0 FlashBynn IntelligenceTypeSafe Jev 1.13TypeSafe AICloudflare Clef 27BCloudflareCloudflare Clef-flash 9BCloudflarepplx-⁠decider-⁠27bopen modelCygnet (Gemma 4 12B)open modelKev-⁠4Bopen modelStrands Decider 2Bopen model
Held-out decisions636 items0.9580.94 to 0.980.9350.91 to 0.960.9480.93 to 0.960.9330.91 to 0.950.9020.88 to 0.930.959 (highest on this suite)0.94 to 0.970.9240.90 to 0.940.8920.87 to 0.920.8360.81 to 0.86
Public decision set117 items0.872 (highest on this suite)0.80 to 0.940.8380.76 to 0.920.8550.78 to 0.920.8550.79 to 0.920.8210.75 to 0.880.8500.78 to 0.910.8680.80 to 0.920.7440.66 to 0.820.7090.62 to 0.79
Transfer tasks730 items0.8360.81 to 0.870.7920.76 to 0.830.852 (highest on this suite)0.81 to 0.890.8100.77 to 0.850.8070.77 to 0.840.8290.79 to 0.870.7860.75 to 0.820.7620.72 to 0.800.6380.59 to 0.69
Web tasks1,106 items0.781 (highest on this suite)0.75 to 0.810.7520.72 to 0.780.6980.66 to 0.730.6970.67 to 0.730.6410.60 to 0.680.6720.63 to 0.710.6320.59 to 0.670.6120.58 to 0.650.5040.47 to 0.54
Knowledge (MMLU-Pro)1,479 items0.6900.66 to 0.720.5850.56 to 0.610.820 (highest on this suite)0.80 to 0.840.6340.61 to 0.660.6350.61 to 0.660.6270.60 to 0.650.5440.52 to 0.570.4940.47 to 0.520.2960.27 to 0.32
Short classifications178 items0.955 (highest on this suite)0.92 to 0.990.9120.86 to 0.960.9040.86 to 0.940.9210.88 to 0.950.7330.68 to 0.790.9410.90 to 0.970.9130.87 to 0.950.7920.74 to 0.850.5080.47 to 0.54
Open decision tasks5,011 items0.8120.80 to 0.820.7900.78 to 0.800.7980.79 to 0.810.8340.82 to 0.840.845 (highest on this suite)0.84 to 0.860.7900.78 to 0.800.7590.75 to 0.770.8110.80 to 0.820.8090.80 to 0.82
Typed decisions955 items0.741 (highest on this suite)0.71 to 0.770.7120.68 to 0.740.7350.70 to 0.770.7050.67 to 0.740.6970.67 to 0.730.7200.69 to 0.750.6630.63 to 0.700.6630.63 to 0.690.5990.56 to 0.63
New task families3,629 items0.8120.80 to 0.830.7810.77 to 0.800.823 (highest on this suite)0.81 to 0.830.8000.79 to 0.810.7840.77 to 0.800.8000.79 to 0.810.7580.74 to 0.770.7380.72 to 0.750.7100.69 to 0.72
Familiar task families4,100 items0.812 (highest on this suite)0.80 to 0.830.7760.76 to 0.790.7740.76 to 0.790.7720.76 to 0.790.7480.73 to 0.760.7770.76 to 0.790.7270.71 to 0.740.7070.69 to 0.720.6120.60 to 0.63
Large option sets1,677 items0.846 (highest on this suite)0.83 to 0.870.8080.79 to 0.830.7750.76 to 0.800.8420.82 to 0.860.8340.81 to 0.850.7750.76 to 0.790.6940.67 to 0.720.7010.68 to 0.720.7090.69 to 0.73
Long inputs1,147 items0.8480.82 to 0.870.7950.77 to 0.820.853 (highest on this suite)0.83 to 0.870.8450.82 to 0.870.8000.77 to 0.820.4920.44 to 0.540.6340.60 to 0.670.7580.73 to 0.790.6680.64 to 0.70
Multilingual1,812 items0.808 (highest on this suite)0.79 to 0.830.7720.75 to 0.790.7820.76 to 0.800.7600.74 to 0.780.7440.72 to 0.760.8000.78 to 0.820.7560.73 to 0.780.7270.71 to 0.750.6870.66 to 0.71
Rules and policies1,759 items0.921 (highest on this suite)0.91 to 0.940.8720.85 to 0.890.8310.81 to 0.850.7980.78 to 0.820.7330.71 to 0.750.8450.83 to 0.860.7780.76 to 0.800.7040.68 to 0.730.5560.53 to 0.58
Prompt injection798 items0.5650.53 to 0.600.5800.54 to 0.620.6000.56 to 0.630.5820.55 to 0.620.659 (highest on this suite)0.63 to 0.690.3180.28 to 0.350.3270.30 to 0.360.6520.62 to 0.680.4810.45 to 0.52

Text and images, one request

Two image suites: receipts and documents, and photos. DecisionNode reads images and text together and answers through the same three question types. TypeSafe Jev reads text only, so it is shown as not supported rather than as zero.

Receipts and documents

Totals, dates and currencies read from document images. 600 items.

DecisionNode-1.0
0.861
DecisionNode-1.0 Flash
0.823
TypeSafe Jev 1.13
not supported, reads text only

Not run on the image suites: Cloudflare Clef 27B, Cloudflare Clef-flash 9B, pplx-decider-27b, Cygnet (Gemma 4 12B), Kev-4B and Strands Decider 2B.

Photos

Objects, conditions and layout judged from photos. 800 items.

DecisionNode-1.0
0.834
DecisionNode-1.0 Flash
0.794
TypeSafe Jev 1.13
not supported, reads text only

Not run on the image suites: Cloudflare Clef 27B, Cloudflare Clef-flash 9B, pplx-decider-27b, Cygnet (Gemma 4 12B), Kev-4B and Strands Decider 2B.

Against general LLMs

Most teams start by asking a chat model for JSON. Here is what that buys you: one short decision sent 2,000 times to eight general LLMs from OpenAI, Anthropic and Google, TypeSafe Jev, and both our models, plus the 15 accuracy suites above.

2,000 requests, each sent twice

2,000 answers. Zero your code has to catch.

One dot per request. A filled dot is an answer that did not fit the requested shape or picked an option never offered; a ring is a request answered differently the second time. Pick a general LLM to compare.

DecisionNode-⁠1.0 and Flashours
of 2,000 invalid
0
of 2,000 changed on repeat
0

Every answer in the requested shape, every repeat the same: true by construction.

GPT-5.6 LunaOpenAI
of 2,000 invalid
77
of 2,000 changed on repeat
172
  • invalid answer
  • different answer the second time
Compare with

Color shows the maker

  • Bynn Intelligenceours
  • TypeSafe AI
  • OpenAI
  • Anthropic
  • Google

Preliminary results, October 2026

How every chart was run

Task
A support ticket routed to one of four teams (Choice), plus a refund call (Truth).
Prompt
General LLMs get the question, the options and the JSON shape to answer in, with JSON mode on where the API has it. Decision models get the same questions as a typed request.
Requests
2,000 per system.
Date
October 2026, preliminary.
Invalid answers
0.00%DecisionNode-⁠1.0 and Flash0.30% at bestClaude Fable 5.1
Same answer twice
100%DecisionNode-⁠1.0 and Flash97.3% at bestClaude Fable 5.1
Time to answer
19 msDecisionNode-⁠1.0 Flash540 ms at bestGemini 3.8 Flash
Per 1,000 decisions
$0.0038DecisionNode-⁠1.0 Flash$0.057 at bestGPT-5.6 Luna
15-suite accuracy
0.817DecisionNode-1.00.826, higherClaude Fable 5.1

Invalid answer rate

Answers that do not parse into the requested shape, or pick an option that was never offered. Each one is a retry, a fallback or a bug in your workflow.

0%1%2%3%4%
  1. DecisionNode-1.0, Bynn Intelligence, ours0.00%. 0 of 2,000 answers invalid.
  2. DecisionNode-1.0 Flash, Bynn Intelligence, ours0.00%. 0 of 2,000 answers invalid.
  3. TypeSafe Jev 1.13, TypeSafe AI0.00%. 0 of 2,000 answers invalid.
  4. Claude Fable 5.1, Anthropic0.30%. 6 of 2,000 answers invalid.
  5. Claude Opus 5.5, Anthropic0.45%. 9 of 2,000 answers invalid.
  6. GPT-5.6 Sol, OpenAI0.60%. 12 of 2,000 answers invalid.
  7. Claude Sonnet 5.5, Anthropic0.90%. 18 of 2,000 answers invalid.
  8. Gemini 3.1 Pro, Google1.10%. 22 of 2,000 answers invalid.
  9. GPT-5.6 Terra, OpenAI1.45%. 29 of 2,000 answers invalid.
  10. Gemini 3.8 Flash, Google2.40%. 48 of 2,000 answers invalid.
  11. GPT-5.6 Luna, OpenAI3.85%. 77 of 2,000 answers invalid.

share of 2,000 answers, lower is better

0.00%

DecisionNode-⁠1.0, Flash and Jev. A typed answer always has its type's shape, and a choice is always one of the keys you sent.

General LLMs: 0.30% (Claude Fable 5.1) to 3.85% (GPT-5.6 Luna). At 10,000 decisions a day, 3.85% is 385 answers your code has to catch.

Counts
Not valid JSON, a missing or extra field, or an option that was not in the request.

Invalid answer rate, how it was run: October 2026, preliminary Support routing and refund, 2,000 requests per system, LLMs prompted for JSON.

Same request, same answer

Every request sent twice. Anything under 100% means the same ticket can be routed two different ways.

90%95%100%
  1. DecisionNode-1.0, Bynn Intelligence, ours100%. 2,000 of 2,000 requests answered the same way twice.
  2. DecisionNode-1.0 Flash, Bynn Intelligence, ours100%. 2,000 of 2,000 requests answered the same way twice.
  3. TypeSafe Jev 1.13, TypeSafe AI98.4%. 1,968 of 2,000 requests answered the same way twice.
  4. Claude Fable 5.1, Anthropic97.3%. 1,946 of 2,000 requests answered the same way twice.
  5. Claude Opus 5.5, Anthropic96.6%. 1,932 of 2,000 requests answered the same way twice.
  6. GPT-5.6 Sol, OpenAI96.1%. 1,922 of 2,000 requests answered the same way twice.
  7. Claude Sonnet 5.5, Anthropic95.2%. 1,904 of 2,000 requests answered the same way twice.
  8. Gemini 3.1 Pro, Google95.0%. 1,900 of 2,000 requests answered the same way twice.
  9. GPT-5.6 Terra, OpenAI94.5%. 1,890 of 2,000 requests answered the same way twice.
  10. Gemini 3.8 Flash, Google92.8%. 1,856 of 2,000 requests answered the same way twice.
  11. GPT-5.6 Luna, OpenAI91.4%. 1,828 of 2,000 requests answered the same way twice.

repeats that matched, of 2,000, higher is better

100%

DecisionNode-⁠1.0 and Flash are deterministic: the same request to the same model version returns the same answer, every time.

General LLMs: 91.4% to 97.3%, at temperature 0 where the API has one. TypeSafe Jev: 98.4%.

Counts
A repeat matches when it gives the same choice and lands on the same side of 0.50 on the Truth question.
Settings
Temperature 0 where the API has it; reasoning models at their provider's defaults. Each request sent twice, an hour apart.

Same request, same answer, how it was run: October 2026, preliminary Support routing and refund, 2,000 requests, each sent twice, LLMs prompted for JSON.

Median time to answer

From sending the request to holding the complete answer. Log scale: each gridline is ten times longer than the last.

10 ms100 ms1 s10 s
  1. DecisionNode-1.0 Flash, Bynn Intelligence, ours19 ms. Reads the request once and answers in one pass, no text written.
  2. DecisionNode-1.0, Bynn Intelligence, ours24 ms. Reads the request once and answers in one pass, no text written.
  3. Gemini 3.8 Flash, Google540 ms. Writes the JSON answer token by token.
  4. GPT-5.6 Luna, OpenAI610 ms. Writes the JSON answer token by token.
  5. GPT-5.6 Terra, OpenAI1,140 ms. Writes the JSON answer token by token.
  6. Claude Sonnet 5.5, Anthropic1,290 ms. Writes the JSON answer token by token.
  7. Gemini 3.1 Pro, Google2,240 ms. Thinks before it answers, then writes the JSON.
  8. Claude Opus 5.5, Anthropic2,480 ms. Thinks before it answers, then writes the JSON.
  9. GPT-5.6 Sol, OpenAI2,910 ms. Thinks before it answers, then writes the JSON.
  10. Claude Fable 5.1, Anthropic4,060 ms. Thinks before it answers, then writes the JSON.

milliseconds, log scale, lower is better

19 ms

DecisionNode-⁠1.0 Flash, end to end with the network. On our side it answers in about 5 ms.

The quickest general LLM here took 540 ms (Gemini 3.8 Flash); the slowest took 4,060 ms. That is 28× to 214× longer than Flash.

Counts
Median wall-clock time per request, sent one at a time from one client, network round trip included.
Not shown
TypeSafe Jev. A speed comparison against other decision APIs comes at launch, timed from the same client.

Median time to answer, how it was run: October 2026, preliminary Support routing and refund, 2,000 requests, one client, LLMs prompted for JSON.

Cost per 1,000 decisions

The tokens each API billed for one decision, at its list price, times 1,000. Log scale.

$0.001$0.01$0.10$1$10$100
  1. DecisionNode-1.0 Flash, Bynn Intelligence, ours$0.0038. 180 input tokens at $0.021 per million; output is free.
  2. DecisionNode-1.0, Bynn Intelligence, ours$0.0076. 180 input tokens at $0.042 per million; output is free.
  3. TypeSafe Jev 1.13, TypeSafe AI$0.0076. 180 input tokens at $0.042 per million; output is free.
  4. GPT-5.6 Luna, OpenAI$0.057. 420 input and 38 output tokens at $0.10 and $0.40 per million.
  5. Gemini 3.8 Flash, Google$0.23. 420 input and 40 output tokens at $0.30 and $2.50 per million.
  6. GPT-5.6 Terra, OpenAI$0.46. 420 input and 38 output tokens at $0.80 and $3.20 per million.
  7. Claude Sonnet 5.5, Anthropic$1.89. 420 input and 42 output tokens at $3 and $15 per million.
  8. Gemini 3.1 Pro, Google$4.44. 420 input and 300 output tokens at $2 and $12 per million.
  9. GPT-5.6 Sol, OpenAI$8.50. 420 input and 320 output tokens at $5 and $20 per million.
  10. Claude Opus 5.5, Anthropic$9.10. 420 input and 280 output tokens at $5 and $25 per million.
  11. Claude Fable 5.1, Anthropic$33.30. 420 input and 360 output tokens at $15 and $75 per million.

US dollars per 1,000 decisions, log scale, lower is better

$0.0038

DecisionNode-⁠1.0 Flash: 180 input tokens per decision, and output is free on both our models.

The cheapest general LLM here costs $0.057 (GPT-5.6 Luna). The most expensive costs $33.30 (Claude Fable 5.1), because its thinking tokens are billed as output.

Counts
Input and output tokens from each API's usage field, times its list price in October 2026.
Tokens
An LLM prompt runs about 420 tokens because it carries the JSON shape; the typed request runs about 180.

Cost per 1,000 decisions, how it was run: October 2026, preliminary Support routing and refund, 2,000 requests at list price, LLMs prompted for JSON.

Accuracy on the 15 suites

Mean accuracy over the same 15 text suites as the rest of this page, with 95% intervals.

0.700.750.800.85
  1. Claude Fable 5.1, Anthropic0.826. 95% interval 0.818 to 0.834.
  2. DecisionNode-1.0, Bynn Intelligence, ours0.817. 95% interval 0.809 to 0.825.
  3. Claude Opus 5.5, Anthropic0.812. 95% interval 0.804 to 0.820.
  4. GPT-5.6 Sol, OpenAI0.809. 95% interval 0.801 to 0.817.
  5. TypeSafe Jev 1.13, TypeSafe AI0.803. 95% interval 0.795 to 0.811.
  6. Gemini 3.1 Pro, Google0.797. 95% interval 0.788 to 0.806.
  7. Claude Sonnet 5.5, Anthropic0.789. 95% interval 0.780 to 0.798.
  8. DecisionNode-1.0 Flash, Bynn Intelligence, ours0.780. 95% interval 0.771 to 0.789.
  9. GPT-5.6 Terra, OpenAI0.768. 95% interval 0.759 to 0.777.
  10. Gemini 3.8 Flash, Google0.741. 95% interval 0.732 to 0.750.
  11. GPT-5.6 Luna, OpenAI0.712. 95% interval 0.703 to 0.721.

accuracy, 15 text suites, higher is better

0.817

DecisionNode-⁠1.0. Flash scores 0.780.

Claude Fable 5.1 scores higher, 0.826, and Claude Opus 5.5 (0.812) and GPT-5.6 Sol (0.809) come close. They take 2,480 ms to 4,060 ms and $8.50 to $33.30 per 1,000 decisions.

Counts
Accuracy per suite averaged over the text suites; intervals by a paired bootstrap.
Prompt
General LLMs get each item as a prompt asking for the JSON shape; an invalid answer counts as wrong.
Items
25,134 per system, in place of the 2,000 requests above.

Accuracy on the 15 suites, how it was run: October 2026, preliminary 15 text suites, 25,134 items per system, 95% intervals by paired bootstrap.

Pick DecisionNode when

The answer is a choice, a score or a Truth probability your code acts on: routing, gating, moderation, checks inside a workflow or an agent loop. It has to land inside a latency budget, the same way every time, at a price that lets you ask on every request.

Pick a general LLM when

You need written text: a reply, a summary, an explanation, a plan. Or the task is open-ended, with no fixed set of answers. Many builds use both: DecisionNode decides, and the LLM writes only when the decision says so.

The general-LLM figures and TypeSafe Jev's repeat match are preliminary and will be replaced by our full measured run before launch. Our own zero invalid answers and 100% repeat match hold by construction.

DecisionNode is independent and is not affiliated with, sponsored by or endorsed by OpenAI, Anthropic, Google or TypeSafe AI. Their names and model names belong to their owners and identify the systems measured.

Answered while a chat model is still typing

A general LLM writes its answer token by token: 540 to 4,060 ms for a short decision, timed end to end from one client. DecisionNode reads the input once and answers in one pass: 19 ms end to end on Flash, 5 ms of it on our side.

28× to 214×faster than a general LLM on the same decision, end to end, on Flash

1 ms10 ms100 ms1,000 ms
  1. DecisionNode-1.0 Flash19 ms end to end, 5 ms on our side
  2. DecisionNode-1.024 ms end to end, 9 ms on our side, preliminary
  3. General LLMs, 8 systems540 to 4,060 ms end to end

milliseconds end to end, log scale

End to end from one client, network included, a short decision; 8 general LLMs from OpenAI, Anthropic and Google. A speed comparison against other decision APIs comes at launch, timed from the same client. The Flash server time is measured; the other figures are preliminary, October 2026.

Accuracy is half of it. Calibration is the other half.

Every Choice and Score answer carries a confidence, and every Truth answer is the probability that a statement is true. Calibrated means those numbers mean what they say: of all the answers given at 0.9, about nine in ten are right.

That is what lets your code act alone. Set a threshold once and every answer above it is acted on instantly; below it, your code takes the safe path or re-checks with the full model. Calibration results per question type will be published with the full run.

How confidence works

  1. right
  2. right
  3. right
  4. right
  5. right
  6. right
  7. right
  8. right
  9. right
  10. wrong
Ten answers given at 0.90. Calibrated means about nine are right.an illustration of the definition, not a measurement

How we measured

15 text suites, 25,134 items. Accuracy with 95% intervals by a paired bootstrap. Jev was called through its official API; open models ran on our GPUs as their authors serve them. October 2026.

The suites

Held-out decisions636
Typed decisions on tasks kept out of development.
Public decision set117
A small public set of choice, score and truth questions.
Transfer tasks730
Decisions phrased the way other decision APIs phrase them.
Web tasks1,106
Pick the right element or action on a real web page.
Knowledge (MMLU-Pro)1,479
Hard multiple-choice questions that need world knowledge.
Short classifications178
One-line inputs sorted into a few labels.
Open decision tasks5,011
A broad public mix of typed decision tasks.
Typed decisions955
Mixed choice, score and truth questions in one request.
New task families3,629
Task families the model never saw while it was built.
Familiar task families4,100
New items from task families the model knows.
Large option sets1,677
Choices with dozens of options to pick from.
Long inputs1,147
Long threads and documents with the answer deep inside.
Multilingual1,812
The same decisions asked in many languages.
Rules and policies1,759
Apply a written policy to a case and decide.
Prompt injection798
Inputs that try to talk the model into the wrong answer.
Receipts and documents600, image
Totals, dates and currencies read from document images.
Photos800, image
Objects, conditions and layout judged from photos.
  • Each suite is scored as accuracy: the share of items where the answer matches the reference. Intervals are 95%, from a paired bootstrap over the same items for every system.
  • TypeSafe Jev 1.13 was called through its official API with its documented request shape. Open models ran on our GPUs, served the way their authors publish them.
  • Strands Decider 2B: its published training data contains items of two of our test sets, so its scores on those suites may read high.

Information last reviewed: October 2026.

Every system here changes over time, ours included. Scores describe the versions measured on the date above; check each provider's current documentation before you choose.

DecisionNode is not affiliated with TypeSafe AI, Cloudflare, OpenAI, Anthropic, Google or the other model publishers. Their names identify the systems measured.

Preliminary results, October 2026

Questions about the numbers

  • Why show the suites you lose?+

    Because a mean hides the shape. You should know that DecisionNode trails on knowledge-heavy questions and on prompt injection before you build on it, and that it leads on policies, web tasks and large option sets.

  • Are these numbers final?+

    No. They are preliminary results, october 2026, and the page says so wherever they appear. They will be replaced by the full measured run before launch.

  • Why is Jev missing from the image suites?+

    TypeSafe Jev reads text only, according to its public documentation, so there is nothing to score. It is shown as not supported, not as zero.

  • Can I check the numbers on my own data?+

    Yes. The playground runs any request against both models with no login, and every answer is deterministic, so the same request gives the same answer every time you run it.

  • How fast is it next to the others?+

    Flash answers a short request in about 5 ms on our side and 19 ms end to end. The general LLMs, timed from the same client on the same decision, took 540 to 4,060 ms, so Flash was 28× to 214× faster. A speed comparison against other decision APIs comes at launch, timed from the same client.

DecisionNodeDecisionNde

The decision model, and the inference API that serves it.

one endpoint: POST api.decisionnode.com/v1/decidePOST /v1/decide

start building

Your first decision in minutes.

Make a key, paste one curl, branch on a typed answer. Prepaid, no sales call.

  • api all systems normalapi normal
  • output is always freeoutput always free
  • same request, same answersame request, same answer

DecisionNode is built and run by Bynn Intelligence, Inc.

We train the model and serve it on our own GPUs.

hello@bynn.com

product

  • Model
  • Inference
  • Benchmarks
  • Pricing

developers

  • Docs
  • Quickstart
  • Examples
  • Playground
  • Console

company

  • Contact
  • Security
  • Report abuse
  • Dedicated

legal

  • Terms
  • Acceptable use
  • Privacy
  • Data processing
  • Defence addendum
  • Cookies

© 2026 Bynn Intelligence, Inc.

prices in USD per 1M input tokens