A public research record: no sign-in, no account, nothing written by hand — every number on this page is read straight from the run that produced it.
START WITH
108
experiments
29
falsified
38,194
API responses
45.9M
input tokens
$1.93
spent
EVERY EXPERIMENT, E01–E108
E47
Does a plain-English permission policy correctly gate real actions?
C
Confirmed
C
Confirmed
The claim held up: what we expected is what we measured.
55
F
Falsified
The claim did not hold up: what we measured was different from, or the opposite of, what we expected.
29
M
Mixed
Part of the claim held up and part did not: different pieces of the same hypothesis came out different ways.
10
I
Inconclusive
The experiment ran, but the result could not settle the claim either way.
1
B
Blocked
The experiment could not run: the real input it needed did not exist yet, and the programme does not allow invented data to fill the gap.
13
SEVEN GROUPS
E01–E16
Limits and error messages
670 records
642,098 tokens
E17–E28
The confidence number, tested
5,875 records
8,215,046 tokens
E29–E46
Choosing where a request goes
2,988 records
12,365,109 tokens
E47–E62
Gate and verify
3,873 records
2,440,607 tokens
E63–E78
Scoring, curating, and finding text
4,588 records
7,106,654 tokens
E79–E92
Multi-step tasks and rechecking claims
15,512 records
10,373,277 tokens
E93–E108
Headlines, claims, and sources
4,688 records
4,718,489 tokens
Bar length is the group's experiment count against the largest group; segments are its verdicts. Records are the API responses the run logged and kept.
TWO PROGRAMMES
PROGRAMME · PUBLISHED
Jev: 108 experiments on a System One model
108 experiments against the real Jev API, every number read straight from the run that produced it.
PROGRAMME · IN PROGRESS
Semantic routing: 36 typed questions in place of a phrase list
One request answers 36 typed questions across 27 sources; the golden rounds measure whether it opens the right ones.
No counted scoreboard yet: no golden round has ever had its scored output committed, so this page carries shape, not numbers.
Every experiment in one table, the words the record uses, and a palette (⌘K) that finds any experiment, group or word by name.
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.