- Runs on
- Every
/v1/decidecall, every batch line, and the opening of a session and each of its frames - Categories
harm(choosing or targeting people to be hurt) andself_harm(help to hurt oneself)- Default for every key
- Flag both categories; refuse a request whose
self_harmprobability is above 0.9 - Where you see it
- The
x-decisionnode-safetyheader on every decision response, and on request asafetyobject in the body - A refusal
403witherror_typerefusal_errorand thecategory; not billed- Cost
- The check itself is never billed and never counts toward
usageor a limit
What it checks#
The check answers two questions about every request, as calibrated probabilities from 0 to 1. It reads what your request asks the model to decide: the questions, their options and their criteria, together with the state.
| Category | Covers | Not in it |
|---|---|---|
harm | Choosing, ranking or targeting people to be physically hurt or killed, or choosing whom to harm by a protected attribute such as race, religion, sex or disability | Harmful content in general, risk detection, content moderation, fraud checks, hiring, credit and the triage of one patient |
self_harm | Helping a person to hurt or kill themselves, such as choosing a method, a place or a dose | Asking whether a message shows a risk of self-harm, and routing a person to help |
Answer always, flag always#
- Every request is answered unless it is refused under your key's policy or the default block. The answers do not depend on your policy: two keys with different settings get the same answers for the same request.
- Every response says what the check found, in the
x-decisionnode-safetyheader, a403included. - The response body stays the same shape. Without the opt-in header below, a
200body holds onlymodel,answersandusage, so clients written for a Jev-shaped API read it unchanged.
The safety header#
x-decisionnode-safety is a structured header (RFC 8941 dictionary): each category with its probability to three decimals, a category above your key's threshold marked ;flagged, and the category that caused a refusal marked ;blocked. When the check did not run, the header says unchecked with a reason code.
| Value | Meaning |
|---|---|
harm=0.031, self_harm=0.004 | Checked; nothing above your thresholds |
harm=0.912;flagged, self_harm=0.004 | Checked; harm is above your key's flag threshold. The request was answered |
harm=0.004, self_harm=0.953;blocked | Checked and refused for self_harm (on the 403) |
unchecked;reason="no-model-question" | Nothing for the model to decide: every question has a single possible answer, such as a choice with one option |
unchecked;reason="does-not-fit" | The check's own reading of the request is longer than the model's limit |
unchecked;reason="no-answer" | The check gave no usable answer |
Read the probabilities by name: more categories may be added, and the order is not fixed.
def parse_safety(header: str) -> dict | None:
"""'harm=0.912;flagged, self_harm=0.004' -> {"harm": (0.912, {"flagged"}), ...}"""
if not header or header.startswith("unchecked"):
return None # the check did not run; the reason is in the header
result = {}
for item in header.split(","):
name, _, rest = item.strip().partition("=")
value, *marks = rest.split(";")
result[name] = (float(value), set(marks))
return result
def audit_if_flagged(response) -> None:
"""hold_for_audit is your own function."""
safety = parse_safety(response.headers.get("x-decisionnode-safety", ""))
if safety and "flagged" in safety["harm"][1]:
hold_for_audit(response.headers["x-request-id"])The safety body#
Send the request header x-decisionnode-safety-body: true (1 and yes work too) and a 200 body carries one more top-level field, safety, with the probabilities to four decimals. It is a request header, not a body field, so the same request works on every path and with clients that refuse unknown fields.
curl https://api.decisionnode.com/v1/decide \
-H "Authorization: Bearer $DECISIONNODE_API_KEY" \
-H "Content-Type: application/json" \
-H "x-decisionnode-safety-body: true" \
-d '{
"model": "decisionnode-latest",
"state": "Customer: I was charged twice and nobody has replied for 3 days.",
"questions": {
"refund": {
"type": "truth",
"instructions": "Refund this automatically?"
}
}
}'{
"model": "decisionnode-1.0",
"answers": {
"refund": { "type": "truth", "truth": 0.94 }
},
"usage": { "input_tokens": 31, "output_tokens": 0 },
"safety": {
"checked": true,
"harm": 0.0012,
"self_harm": 0.0006,
"flagged": []
}
}The safety object
checkedbooleantruewhen the check ran. When it isfalse, the object holds onlyreason.harmnumber- The probability that the request asks to choose, rank or target people to be harmed, 0 to 1.
self_harmnumber- The probability that the request seeks help to harm oneself, 0 to 1.
flaggedstring[]- The categories above your key's thresholds. Empty when none are.
reasonstring- Only when
checkedisfalse:no-model-question,does-not-fitorno-answer, as in the header.
Session replies always carry this object, one per frame. Batch result lines carry the header's value in their headers.
Your key's policy#
Each key has a policy: for each category an action and a threshold. Flag marks the category in the header and the body when its probability is above the threshold, and still answers. Block refuses the request with a 403 above the threshold. The policy applies to live calls, batches and sessions alike.
| Setting | Rule |
|---|---|
| The default | Flag harm and self_harm; refuse self_harm above 0.9. This applies to every key |
harm | Flag or block, at a threshold you choose |
self_harm | Flag or block, at a threshold you choose. A block can only be stricter than the default, at 0.9 or below; it can never be loosened or switched off |
harm off | Only on the named keys of a signed Defence Contract Addendum, set by us. No other key can switch a category off |
A policy only tightens: where we have set a block on your key, your own setting applies on top of it and our block stays. To set or change your key's policy, write to hello@bynn.com from your account's contact address with the key's name and the settings you want; settings are made on written request and confirmed back to you in writing. No request field can change the policy.
When a request is refused#
A refused request gets 403 instead of answers. The body names the category, x-decisionnode-refusal repeats it, the safety header marks it ;blocked, and x-decisionnode-billable-tokens says what the refusal bills: 0. The same request is refused again, so never retry it.
{
"detail": {
"error_type": "refusal_error",
"message": "This request was refused under the usage policy: it asks for help to harm oneself. If someone may be in danger, contact local emergency services or a crisis line.",
"category": "self_harm"
}
}Branch on error_type and category, not on the message: the message wording may change. A self_harm refusal's message points to emergency services or a crisis line; if your product talks to people, show them your own route to help as well.
import requests
def handle_refusal(response: requests.Response, user) -> bool:
"""True when the safety check refused the request, after acting on it.
log_refusal and show_crisis_resources are your own functions."""
if response.status_code != 403:
return False
detail = response.json().get("detail", {})
if detail.get("error_type") != "refusal_error":
return False
# refused under the usage policy; not billed, and refused again if resent
log_refusal(detail["category"], response.headers["x-request-id"])
if detail["category"] == "self_harm":
show_crisis_resources(user) # your product's route to human help
return TrueA 403 with error_type refusal_error is always this refusal: a missing, unknown or revoked key gets 401, and a paused workspace gets 403 with workspace_suspended and no category (see Errors). A 5xx or a 529 is never a refusal.
When the check cannot run#
- A key that only flags gets its answers as usual, with
uncheckedand the reason in the header. - A key that blocks a category is never answered unchecked: the request gets
500witherror_typeapi_error. It is safe to retry, and the same request usually succeeds. - A request with nothing for the model to decide (every question has one possible answer) is answered with the header marked
uncheckedfor the reasonno-model-question, whatever the policy.
Batches and sessions#
- Batch jobs. Every line is checked like a live call, under the same default block and your key's policy. A refused line has status
403and the live body in its result, and is not billed. See Get batch results. - Sessions. The request that opens a session and every frame are checked. Every reply carries the
safetyobject. A refused frame gets a frame error withrefusal_errorand never joins the window, and the socket stays open for the next frame. Unlike a refused call, a refused frame is billed: the model read it within the session. See Stream frames.
What is kept#
Records of flagged and refused requests hold the probabilities, the action taken, the key's id, the request id and the token counts. They never hold the text of a request: no state, question, option or image. See Data and privacy.