What calibrated means#
Take every Truth answer the model gave near 0.9. If about nine in ten of those statements are actually true, the probabilities are calibrated. That is what makes a threshold meaningful: at 0.9 you accept about one wrong answer in ten among the ones that pass, and you can reason about that before you ship.
Where confidence lives in each type#
| Type | Field to gate on | Range and meaning |
|---|---|---|
| Choice | confidence | 0 when all options are equally likely, 1 when one option has all the probability |
| Score | confidence, or a sum of probabilities | How tightly the levels sit around the most likely one; tail sums answer "at least level n" |
| Truth | truth | The probability that the statement is true; distance from 0.50 is how sure it is |
| Number | confidence, or a sum of probabilities | The probability of number; add its neighbours for the chance the true value is within one |
Choice confidence#
confidence = (K · p_max − 1) / (K − 1)
- K
- the number of options
- p_max
- the probability of the chosen option
Score confidence#
confidence = 1 − E[|i − m|] / D(K)
D(K) = floor(K · K / 4) / K
- K
- the number of levels
- m
- the most likely level, the lowest one on a tie
- E[|i − m|]
- the expected distance from m: sum(p_i · |i − m|)
Both are computed from the unrounded probabilities, then rounded to 2 decimals. A score's confidence is 1 when all the probability sits on one level, and probability next to the most likely level costs little. With two levels the two formulas agree. The Score page works one through.
Choosing a threshold#
- 1.Start at 0.7 for choice and number confidence and 0.8 for Truth gates.
- 2.Run a few hundred real inputs whose outcome you already know, such as resolved tickets or settled disputes.
- 3.For each threshold, measure two numbers: how many inputs pass, and how many of those are wrong.
- 4.Pick the threshold where the error rate is one your product can live with. Above it your code acts instantly; below it, it takes the safe action or re-checks with the full model.
- 5.Re-check after you change the questions or move to a new model version.