Where the check goes#
Put the check between the agent deciding to call a tool and the tool running. The state is the proposed call plus the context that matters; the questions are your policy, written as Truth and Choice questions.
{
"model": "decisionnode-flash-latest",
"state": {
"user_request": "Clean up the old staging data",
"proposed_call": {
"tool": "run_sql",
"args": {
"database": "production",
"sql": "DELETE FROM orders WHERE created_at < '2026-01-01'"
}
}
},
"questions": {
"safe": {
"type": "truth",
"instructions": "Is it safe to run this call right now?",
"criteria": {
"true": "read-only, reversible, or clearly what the user asked for",
"false": "destructive, irreversible, or outside what the user asked for"
}
},
"in_scope": {
"type": "truth",
"instructions": "Does the call do what the user asked, and nothing more?"
},
"risk": {
"type": "choice",
"instructions": "What is the main risk of this call?",
"criteria": {
"none": "no meaningful risk",
"data_loss": "deletes or overwrites data",
"money": "moves money or changes billing",
"privacy": "exposes personal data",
"other": "another kind of risk"
}
}
}
}const BAR = { safe: 0.95, inScope: 0.9 };
export async function guardedCall(call: ToolCall, context: Context) {
let answers;
try {
answers = await decide(gateRequest(call, context));
} catch {
return block(call, { reason: "guardrail unavailable" }); // fail closed: the agent re-plans
}
const { safe, in_scope, risk } = answers;
if (safe.truth >= BAR.safe && in_scope.truth >= BAR.inScope) {
return runTool(call);
}
return block(call, { reason: risk.choice, safe: safe.truth, inScope: in_scope.truth });
}Why it works as a gate#
- Fast enough for every step: Flash answers short requests in milliseconds, so the check fits in front of every tool call.
- Deterministic: the same proposed call gets the same verdict, so you can test your policy with fixed cases in CI.
- Calibrated: a bar of 0.95 means what it says. Pin
decisionnode-1.0ordecisionnode-1.0-flashonce the bar is tuned.