E63–E78 · REPORT, AS WRITTEN
Report — experiment group 5 (E63–E78): score, curate, localise
16 experiments, 4,588 records and 7,106,654 input tokens.
C
Confirmed
7
F
Falsified
3
B
Blocked
6
This group tested how to build a good rating question, whether a curation decision should discard its distribution, and how well Jev can point at one piece of text inside a longer document.
Reading the record…
Experiments in this group
Exp
Verdict
Question
E63
F
Falsified
Does a 5-level rating beat multiple-choice for judging wine quality?
E64
C
Confirmed
How many rating levels give the best accuracy: 2, 3, 5, 7, or 10?
E65
C
Confirmed
Does the rating's returned legend text ever drift from what was submitted?
E66
F
Falsified
Does the shape of a rating's probabilities predict better than its mean alone?
E67
C
Confirmed
How much does a rating change on its own, with no change to the input?
E68
C
Confirmed
Does splitting a rating into parts help, or is it just fitted weights?
E69
C
Confirmed
Does deciding by a rating's mean differ from deciding by its nearest level?
E70
B
Blocked
Does a keep/drop decision improve when the model sees more surrounding text?
E71
B
Blocked
Can a simple code check replace a yes/no question for spotting sycophancy?
E72
B
Blocked
Does one yes/no question miss circular reasoning buried in a long trace?
E73
F
Falsified
Does putting many rows in one request change the answers you get?
E74
B
Blocked
At a fixed budget, does ranking beat a simple threshold for keeping rows?
E75
B
Blocked
When the right text span is missing from the options, does the model say so?
E76
C
Confirmed
Does accuracy fall as the list of choices grows toward the 255-option limit?
E77
C
Confirmed
Does asking everything at once match a two-step narrow-then-read approach?
E78
B
Blocked
When nothing on a list matches, does the escape-hatch option catch it?
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.