E17–E28 · GROUP 2 OF 7
The confidence number, tested
This group tested what Jev's confidence number is good for: telling right answers from wrong ones, warning when an answer might flip, and deciding when a gate should let an answer through on its own.
Every system built on Jev has to decide, at some point, whether to trust an answer without a person checking it, and confidence is the one number sold for exactly that job. Nobody had measured whether it actually predicts correctness, whether it says anything the model's plain top probability does not already say, or whether a cutoff picked on one batch of answers holds up on the next batch. Twelve tests ran against real logged answers to find out, several of them by replaying stored data rather than asking new questions, which is why some of them cost nothing at all.
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.