E17–E28 · REPORT, AS WRITTEN
Report — group 2 (E17–E28, confidence): is confidence a second axis or a restatement of max(p)?
12 experiments, 5,875 records and 8,215,046 input tokens.
C
Confirmed
7
F
Falsified
4
M
Mixed
1
This group tested what Jev's confidence number is good for: telling right answers from wrong ones, warning when an answer might flip, and deciding when a gate should let an answer through on its own.
Reading the record…
Experiments in this group
Exp
Verdict
Question
E17
C
Confirmed
Does a high confidence score mean an answer is more likely correct?
E18
C
Confirmed
Is confidence a repeat of the model's top probability, or new information?
E19
F
Falsified
Out of five ways to measure how sure an answer is, which works best?
E20
M
Mixed
Where should a confidence cutoff be set, and does it pay off?
E21
C
Confirmed
Is confidence a better warning sign than adding a yes/no check question?
E22
F
Falsified
If a question is reworded without changing its meaning, does confidence hold steady?
E23
C
Confirmed
Does low confidence warn you that an answer might flip on a repeat?
E24
C
Confirmed
Can how a question is written push confidence up or down?
E25
C
Confirmed
Does confidence mean the same thing with 3 answer options as with 40?
E26
F
Falsified
Can a confidence cutoff be picked from old answers instead of new tests?
E27
F
Falsified
If you ask the same judgment three ways, do their certainty scores agree?
E28
C
Confirmed
Does a confidence-based pause button actually change what an automated loop does?
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.