Two new question types answer where: point gives one spot in the image, box draws a rectangle around the thing. You get coordinates you can draw, crop or hand to a machine, and every answer says how likely the thing is there at all and how likely the place is right.
No other decision API returns a place. OpenAI's Decisions API answers predicate, choice and score questions only, and has no coordinates (checked on 8 October 2026).
Try it in the playground: the belt request below, with the point and the box drawn on the frame. Or open a build: conveyor pick, panel scratch, person at the fence, landing marker.
| Type | Goal | What comes back |
|---|---|---|
Point point | Where is it? | One spot in the image, as x and y, with the probability that the thing is there at all (present) and that the spot is on it (confidence). |
Box box | Where is it, and how far does it reach? | A rectangle around the thing, as its top-left corner (x1, y1) and bottom-right corner (x2, y2), with present and confidence. |
Request#
Ask for a place in plain words, next to an image: "Point to the yellow package closest to the robot arm", "Draw a box around the damaged package". Point and box questions mix freely with the other types in one request, and several of them on the same image share one read of it.
{
"images": [
{
"id": "belt",
"url": "https://files.example.com/line-3/cam-2/0917.jpg"
}
],
"questions": {
"pick": {
"type": "point",
"instructions": "Point to the yellow package closest to the robot arm."
},
"damage": {
"type": "box",
"instructions": "Draw a box around the damaged package."
},
"any_damage": {
"type": "truth",
"instructions": "Is any package on the belt damaged?"
}
}
}Question fields
type"point" | "box"required- A point for one spot, a box for a rectangle around the thing.
instructionsstring | object | arrayrequired- What to find: a short phrase ("the yellow package"), a description with conditions ("the person standing nearest the door"), or a question ("Where is the crack?"). The state and the other images are read as context, so the instructions may refer to them, an image by its
idin backticks or by its number ("picture 1"). imagestring- The
idof the image to look in, a string, never a number. Required when the request carries two or more images; optional with exactly one. See Referring to images. units"pixels" | "normalized"default"pixels"pixels: pixels of the image as you sent it, after its EXIF orientation.normalized: fractions of the width and the height, from 0 to 1.criteriaany- Not used. Ignored like an unknown field inside a question.
- An image is required. A request with a point or box question and no image is a
422on that question. imagenames an image of the request. An unknown id, or noimagein a request with two or more images, is a422on that question'simage.- Questions stay independent. A point or box answer does not depend on the other questions in the request, and several point and box questions on one image share one read of it.
- Video frames: to place something on a frame of a video, send that frame as an image.
Referring to images#
With several images, the image field names the one to look in, by its id. The instructions can refer to the other images by id, in backticks, or by number, counted from 1 in the order of the images array: here the box looks in dock_3 for the package shown in picture 1. Prefer ids to numbers: an id stays right when images are added or reordered. More on Images.
{
"images": [
{
"id": "dock_1",
"url": "https://files.example.com/dock-2/cam-1/0912.jpg"
},
{
"id": "dock_2",
"url": "https://files.example.com/dock-2/cam-2/0912.jpg"
},
{
"id": "dock_3",
"url": "https://files.example.com/dock-2/cam-3/0912.jpg"
}
],
"questions": {
"package": {
"type": "box",
"instructions": "Draw a box around the package from picture 1.",
"image": "dock_3"
}
}
}Response#
{
"answers": {
"pick": {
"type": "point",
"x": 812,
"y": 455,
"present": 0.97,
"confidence": 0.91,
"image": "belt",
"width": 1920,
"height": 1080
},
"damage": {
"type": "box",
"x1": 300,
"y1": 120,
"x2": 520,
"y2": 410,
"present": 0.95,
"confidence": 0.88,
"image": "belt",
"width": 1920,
"height": 1080
},
"any_damage": { "type": "truth", "truth": 0.96 }
}
}Answer fields
x, ynumber- Point: the spot, measured from the top-left corner of the image,
xto the right andydown. Whole pixels by default; withunits: "normalized", fractions with 4 decimals. x1, y1, x2, y2number- Box: the top-left corner (
x1,y1) and the bottom-right corner (x2,y2), in the same units.x1 <= x2andy1 <= y2. presentnumber- The probability that what you asked for is in the image at all. "Not there" is an answer: under a low
presentthe place is only the best guess for something that is probably absent. confidencenumber- The probability that the place is right, if the thing is there. For a point: that it lies on the thing. For a box: that it overlaps the true box by at least half (an intersection over union of 0.5 or more). Calibrated like every probability the API returns.
image, width, heightstring, number, number- Which image the answer refers to and its size in pixels, so you can draw or convert the answer without another lookup.
Every field is always there: the shape never changes with the answer, so code stays simple. present and confidence are rounded to 2 decimals like every probability in the body.
a = result["answers"]["pick"]
if a["present"] > 0.8 and a["confidence"] > 0.7:
robot.pick(a["x"], a["y"])
else:
robot.hold() # unsure: hold, and ask again on the next frameHolding and asking again on the next frame keeps the loop running on its own, the default for a line or a robot. Where a person is on hand, send the unsure image to them instead.
When several things match ("the yellow package" with three on the belt), the answer is the one that matches best. Say which one you mean in the instructions: "the leftmost yellow package".
What instructions can say#
Instructions can carry conditions ("the package with a torn label", "the person who entered after closing") and refer to the state and to other images in the request ("the item that matches the order in the state", "the part that differs from reference"). A fixed list of detector classes cannot.
| Use | Example question | Type |
|---|---|---|
| Production lines and sorting | "Point to the next yellow package to pick." | point |
| Quality control | "Draw a box around the scratch on the panel." | box |
| Surveillance | "Draw a box around the person climbing the fence." | box |
| Traffic | "Draw a box around the car that stopped in the crossing." | box |
| Warehouses | "Point to the empty slot on the second shelf." | point |
| Drones and robots | "Point to the landing marker." | point |
How DecisionNode answers#
DecisionNode reads the image once. A point or box question then runs as a few short readings on that one read: first whether the thing is there, then where it is across the width, then down the height at that spot (for a box, its size around that centre). Each reading is a calibrated probability, so the answer comes with honest present and confidence values instead of a guess written as text. No text is generated and nothing needs to be parsed.
Batch jobs and sessions#
A batch line takes point and box questions as the live call does. In a session they apply to a frame's image.
Price#
Billed like any other question: the tokens of its instructions, plus the image once per request at the image and video rate. No fee per point or box, and output is free. Proposed, like every price on the site. See Pricing and billing.