DecisionNode reads video as video. Consecutive frames are paired, each pair carries its timestamp, and the model reads them as one clip with order and timing, not as a pile of separate pictures. That is how it can tell a car that turned left from one that came back, or a door that was forced from one that was opened. It is also cheaper: a frame costs about half the tokens of the same picture sent as an image.
Why frames, not video files#
- Cameras already make frames. Edge boxes, production lines and camera networks produce frames; send the ones you already have.
- Small and fast to send. A few small JPEGs are a few MB, where a video file of the same scene can be 100 MB.
- Simple and safe. The API takes images it already knows how to check, with the times you give them.
Video files and video URLs are not taken: send frames. The dn CLI and the local MCP server cut a local video file into frames on your own machine (below).
Request#
Add a videos array next to the state, beside any images. Every question sees every video, every image and the state, and answers through the same question types: here a truth, a number and a score about one dock camera.
{
"state": "Warehouse camera 4, loading dock, shift B.",
"videos": [
{
"id": "dock",
"fps": 2,
"frames": [
{ "t": 30, "media_type": "image/jpeg", "data": "<base64>" },
{ "t": 30.5, "media_type": "image/jpeg", "data": "<base64>" }
]
}
],
"questions": {
"fall": {
"type": "truth",
"instructions": "Does a person fall in the video?"
},
"forklifts": {
"type": "number",
"instructions": "How many forklifts pass the door?",
"min": 0,
"max": 50
},
"risk": {
"type": "score",
"instructions": "How dangerous is the behaviour shown?",
"criteria": ["safe", "minor", "serious", "critical"]
}
}
}Video object
idstringrequired- Your name for the video, unique within the request across
imagesandvideos. Questions refer to the video by this id, in backticks (dock), or by its number ("video 1", "the first video"), counted from 1 in the order of thevideosarray and separately from the images. See Referring to images and videos. framesarrayrequired- The frames, in time order: each a frame object, below.
fpsnumberdefault2- A cap on how many frames a second are read, above 0 and at most 4. A frame closer than
1 / fpsseconds to the frame read before it is skipped, never moved. Send"fps": 4to have frames up to 4 a second read. startnumber- The start of the window read, in seconds on the frames' own times. A frame is read when
start <= t < end. Default: every frame sent. endnumber- The end of the window read, after
start, on the same clock.
Frame object
tnumberrequired- The frame's time in seconds, at least 0 and larger than the
tof the frame before it. Seconds from the start of the clip read best. media_typestringrequiredimage/jpeg,image/pngorimage/webp.datastringrequired- The frame's bytes, base64 encoded, with no
data:prefix.
Response#
The normal decision response, one answer per question, with three additions that only a request with videos carries:
| Field | Meaning |
|---|---|
usage.video_frames | The frames read, over all videos |
usage.video_seconds | The seconds of video read, summed over the windows: 20 frames half a second apart are 10 seconds |
warnings | Present when the server read fewer frames or a smaller size than asked, to stay within the request's budget: one sentence per video, naming the rate or the size it read |
Limits#
| Limit | Value |
|---|---|
| Videos per request | Up to 4 |
| Frames per request | Up to 128, counted over all the request's videos: 64 seconds of one camera at 2 a second, or 16 seconds each of 4 cameras |
| Frame types | JPEG, PNG or WebP, each checked as an image is |
| Frame size | The server picks a small size by default, about 480 x 256 for 16:9, so requests stay fast. A frame sent larger is scaled down |
| Tokens | About 64 tokens a frame at the default size, counted in usage.input_tokens and the context limit |
Video tokens are billed with image tokens, at $0.09 per million on DecisionNode-1.0 (proposed), with no fee per frame. See Pricing and billing.
How many frames a use case needs#
| Use | Frames per request | Example |
|---|---|---|
| Sorting one item on a line | 1 to 2 | "Send yellow packages to lane 2" |
| Incident check | About 8 (1 a second) | Fire, smoke, a break-in, a fall |
| Counting and flow | About 20 (2 a second over 10 s) | Cars passing, people entering |
| Longer review | Up to 128 | A short scene summed up in one decision |
At the default size, the 20 frames of a counting request are about 1,280 tokens.
Turning a video file into frames#
Two of our tools do it on your own machine with ffmpeg (ffmpeg and ffprobe on your PATH). Both take frames by the API's own rule at the rate and window you choose (--fps, default 2, at most 4; --start and --end in seconds from the video's first frame; at most --max-frames of them, 32 when left out and up to 128), write each as a JPEG of at most 640 pixels on the longer side, and send the videos field:
- The
dnCLI:--video PATHor--video ID=PATHondn truth,dn choice,dn scoreanddn number, up to 4 videos. - The MCP server, run locally:
decide_videotakes a video as a file path,{"id": "dock", "path": "~/, beside videos sent as frames. The hosted server takes frames only.clips/ dock.mp4", "fps": 2, "start": 30, "end": 45}
dn truth "Does a person fall in the video?" --video dock=dock.mp4 --fps 2 --start 30 --end 90 --max-frames 64The file and its sound never leave your machine; only the frames are sent.
What to ask#
- Traffic: how many cars passed (Number), which way a vehicle turned (Choice), whether a driver went the wrong way (Truth), how congested the junction is (Score).
- People in an area: how many people entered (Number), how crowded the space is (Score).
- Incidents: whether someone forced a door, whether smoke or fire appeared, whether a person fell (Truth), how serious it is (Score).
- Production lines: whether the conveyor stopped, whether a jam formed, how many products went by (Truth, Number).
- From above: monitoring a site, a field or a yard from drone frames, with the same questions.
To find where something is on a frame, send that frame as an image with a point or a box question.
Every answer is typed and calibrated, as on text and images, so your code acts on it the same way. See Number for counts and Confidence for thresholds.