THE RESEARCH RECORD
Claims that could lose, run against real systems.
Every experiment here started as a claim that could lose: a plain statement of what was expected, plus the exact result that would prove it wrong, written down before any call was made. Two programmes so far.
Jev
108 experiments · 29 falsified
PROGRAMME
Jev: 108 experiments on a System One model
108 experiments against the real Jev API, every number read straight from the run that produced it.
108
experiments
29
falsified
38,194
API responses logged
45.9M
input tokens
E47
Does a plain-English permission policy correctly gate real actions?
C
Confirmed
C
Confirmed
The claim held up: what we expected is what we measured.
55
F
Falsified
The claim did not hold up: what we measured was different from, or the opposite of, what we expected.
29
M
Mixed
Part of the claim held up and part did not: different pieces of the same hypothesis came out different ways.
10
I
Inconclusive
The experiment ran, but the result could not settle the claim either way.
1
B
Blocked
The experiment could not run: the real input it needed did not exist yet, and the programme does not allow invented data to fill the gap.
13
Routing
one request a turn
PROGRAMME · IN PROGRESS
Semantic routing: 36 typed questions in place of a phrase list
One request answers 36 typed questions across 27 sources; the golden rounds measure whether it opens the right ones.
Nothing here is counted yet. This programme's round results are not published, so the page carries its shape and no scoreboard.
How to read this record
counted · reported
Counted, or only reported
A counted figure is read straight from the file the run itself produced, by a build step that refuses to emit if the number moves. Every figure in the Jev programme is counted, and that is the only reason this record is worth reading.
A reported figure exists only in a written note. Where one appears here it is marked as reported and linked to the note it came from, and it is never added to, averaged with, or charted against a counted one.
Words used across the record
Every term either programme uses has a plain meaning, and the search palette finds any of them. Press the term anywhere it appears to read its meaning in place.
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.