E79–E92 · REPORT, AS WRITTEN
Report — experiment group 6 (E79–E92): composition, loops, replications
14 experiments, 15,512 records and 10,373,277 input tokens.
C
Confirmed
6
F
Falsified
8
This group tested whether doing well on single steps predicts doing well over a whole task, and re-ran several widely cited results from other projects to see whether they hold up.
Reading the record…
Experiments in this group
Exp
Verdict
Question
E79
F
Falsified
Does doing well on single steps predict doing well on a whole task?
E80
F
Falsified
Do mistakes in a loop pile up the way independent errors would?
E81
C
Confirmed
Does judgement get worse the longer a step-by-step task runs?
E82
F
Falsified
Which number better flags a bad step, confidence or the top answer's odds?
E83
F
Falsified
Does asking the same question many times in one request get the same answer?
E84
F
Falsified
Does asking the exact same question twice get the exact same answer?
E85
C
Confirmed
Is asking everything in one request as good as asking in two steps?
E86
C
Confirmed
Does accuracy hold up when the input text gets very long?
E87
F
Falsified
Does giving the model the whole map really break its navigation?
E88
F
Falsified
Does structuring the question's criteria actually improve answers?
E89
C
Confirmed
Does a text-based version of a game avoid locking onto one action?
E90
F
Falsified
Does picking the right tool for a task beat simple keyword matching?
E91
C
Confirmed
Does the order in which options are listed change the model's answer?
E92
C
Confirmed
Is the model's confidence just its top answer's probability in disguise?
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.