| Model | Model id | Input / 1M | Batch / 1M | Output |
|---|---|---|---|---|
| DecisionNode-1.0 | decisionnode-latest | $0.042 | $0.021 | Free |
| DecisionNode-1.0 Flash | decisionnode-flash-latest | $0.021 (provisional) | $0.0105 (provisional) | Free |
| Dedicated | Reserved capacity on our own GPUs, custom limits: talk to us | |||
How a request is billed#
cost = usage.input_tokens
× price per token
- usage.input_tokens
- the state, images and every question, as returned in the response
- price per token
- the model's price per million, divided by 1,000,000
usage.output_tokens is always 0. There are no seats, no minimums and no charge per question: the price depends only on what you send. Requests refused with 400, 401, 402, 422, 429 or 529 are not billed, so you pay only for requests that return answers. A request refused by the safety check (403) is not charged either (pending confirmation).
Batch jobs#
Work that can wait runs as a batch job at half the price per input token, on either model. Output stays free, and a batch answer has the same shape as a real-time one. Batch jobs take every question type, number questions included.
- Batch results come with a multi-hour delay. There is no set time at which a batch runs or finishes.
- Poll the batch job for its results. Nothing is pushed to you.
- Use batches for backfills, nightly scoring and re-checks of past decisions. Anything your software is waiting on belongs in a real-time request.
Number questions#
A number question costs what a choice with the same number of options costs: each value of its grid counts like one option. A count on 0 to 50 is billed like a choice of 51 options, and the default grid of 0 to 255 like one of 256. A tight grid costs less. Images are priced as on every call.
Sessions coming soon#
A session bills its context once, when it opens: the instructions, the state and the questions, at the live price per input token. After that, each frame is billed for its own tokens and the questions it asks, at the same live price. The context is never billed again per frame, and output stays free.
session = opening tokens × price
+ sum(frame tokens × price)
- opening tokens
- the instructions, the state and the questions, once
- frame tokens
- each frame's own tokens plus the questions it asks
Worked examples#
Each workload asks three questions per request. Tokens is usage.input_tokens for one request, as the response reports it.
| Workload | Tokens | Requests a month | DecisionNode-1.0 | DecisionNode-1.0 Flash |
|---|---|---|---|---|
| Support tickets | 62 | 1,000,000 | $2.60 | $1.30 |
| Comment moderation | 94 | 30,000,000 | $118.44 | $59.22 |
| Receipt photo checks | 1,236 | 200,000 | $10.38 | $5.19 |
Prepaid balance#
- Add credit in the console under Billing. Requests draw it down token by token.
- Turn on auto reload to top up when the balance drops below an amount you choose.
- Every top-up gets an invoice in the console.
- When the balance reaches zero, requests return
402Out of credit until you add credit, so a runaway job cannot spend money you did not put in. A refused request is not billed.
Watching spend#
The console's Usage page charts spend, tokens and requests by model and by day. In code, read usage.input_tokens from each response; it is exactly what was billed.