A public research programme with no sign-in and no account. Every number on this page is read from the results file of the run that produced it.
START WITH
108
experiments
29
falsified
38,194
API responses
45.9M
input tokens
$1.93
spent
EVERY EXPERIMENT, E01–E108
E47
Does a prose permission policy gate proposed actions at a usable false-positive rate?
C
Confirmed
C
Confirmed
The claim held: the measurement matched the expectation.
54
F
Falsified
The claim failed: the measurement differed from or reversed the expectation.
29
M
Mixed
Part of the claim held and part failed: components of one hypothesis yielded opposite outcomes.
10
I
Inconclusive
The experiment ran, and the result settles the claim in neither direction.
2
B
Blocked
The experiment did not run: the input it required did not exist, and the programme admits no invented data in its place.
13
SEVEN GROUPS
E01–E16
Limits and error messages
670 records
642,098 tokens
E17–E28
Confidence as a gate
5,875 records
8,215,046 tokens
E29–E46
Selection and routing
2,988 records
12,365,109 tokens
E47–E62
Gate and verify
3,873 records
2,440,607 tokens
E63–E78
Scoring, curating and localising
4,588 records
7,106,654 tokens
E79–E92
Loops and replications
15,512 records
10,373,277 tokens
E93–E108
Headlines, claims, and sources
4,688 records
4,718,489 tokens
Bar length is the group's experiment count against the largest group; segments are its verdicts. Records are the API responses the run logged and kept.
TWO PROGRAMMES
PROGRAMME · PUBLISHED
Jev: 108 experiments on a System One model
108 experiments against the real Jev API, every number read straight from the run that produced it.
PROGRAMME · IN PROGRESS
Semantic routing: 36 typed questions in place of a phrase list
One request answers 36 typed questions across 27 sources; the golden rounds measure whether it opens the right ones.
No golden round has had its scored output committed. This page carries the shape of a turn and no counted scoreboard.
Every experiment in one table, the words the record uses, and a palette (⌘K) that finds any experiment, group or word by name.
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-20.