JEV · RESULTS LEDGER
Every experiment and what it settled
108 experiments in 7 groups, answered against 38,194 records, each one an API response the run logged and kept.
Verdicts by group
Confirmed (C)
Falsified (F)
Mixed (M)
Inconclusive (I)
Blocked (B)
E01–E16
Limits and error messages
16
Confirmed 12 of 16
Falsified 3 of 16
Mixed 1 of 16
E17–E28
The confidence number, tested
12
Confirmed 7 of 12
Falsified 4 of 12
Mixed 1 of 12
E29–E46
Choosing where a request goes
18
Confirmed 5 of 18
Falsified 5 of 18
Mixed 1 of 18
Inconclusive 1 of 18
Blocked 6 of 18
E47–E62
Gate and verify
16
Confirmed 7 of 16
Falsified 5 of 16
Mixed 4 of 16
E63–E78
Scoring, curating, and finding text
16
Confirmed 7 of 16
Falsified 3 of 16
Blocked 6 of 16
E79–E92
Multi-step tasks and rechecking claims
14
Confirmed 6 of 14
Falsified 8 of 14
E93–E108
Headlines, claims, and sources
16
Confirmed 11 of 16
Falsified 1 of 16
Mixed 3 of 16
Blocked 1 of 16
Bar length is the group's experiment count against the largest group; segments are its verdicts. Solarized's green and red sit at the same lightness, so every verdict carries its letter and word too.
GROUP
VERDICT
108 of 108 shown.
Exp
Verdict
Question
Records
C
Confirmed
How many levels can a Score question have before Jev refuses it?
11
C
Confirmed
How many options can a Choice question have before Jev refuses it?
9
C
Confirmed
Does changing the order of a question's options change the answer?
6
C
Confirmed
Can other questions in the same request change an answer?
25
F
Falsified
Does Jev's confidence number always run below the top probability?
6
C
Confirmed
Can you calculate Jev's confidence number yourself from the probabilities?
6
M
Mixed
Does Jev's response say exactly which model version answered?
12
C
Confirmed
What error codes does Jev actually return for a bad request?
11
C
Confirmed
Do you have to include instructions text on every question?
3
C
Confirmed
What in a request actually drives its token cost?
12
C
Confirmed
What happens if a request goes over the token budget?
8
F
Falsified
Does asking Jev the exact same question twice give the exact same answer?
20
F
Falsified
Do Jev's answers always follow their own basic math rules?
5
C
Confirmed
Does renaming a question or its answer keys change the result?
6
C
Confirmed
How many requests at once before Jev starts rejecting them?
480
C
Confirmed
Does asking more questions in one request save money without slowing it down?
50
C
Confirmed
Does a high confidence score mean an answer is more likely correct?
642
C
Confirmed
Is confidence a repeat of the model's top probability, or new information?
963
F
Falsified
Out of five ways to measure how sure an answer is, which works best?
400
M
Mixed
Where should a confidence cutoff be set, and does it pay off?
0
C
Confirmed
Is confidence a better warning sign than adding a yes/no check question?
300
F
Falsified
If a question is reworded without changing its meaning, does confidence hold steady?
720
C
Confirmed
Does low confidence warn you that an answer might flip on a repeat?
888
C
Confirmed
Can how a question is written push confidence up or down?
360
C
Confirmed
Does confidence mean the same thing with 3 answer options as with 40?
1,200
F
Falsified
Can a confidence cutoff be picked from old answers instead of new tests?
0
F
Falsified
If you ask the same judgment three ways, do their certainty scores agree?
102
C
Confirmed
Does a confidence-based pause button actually change what an automated loop does?
300
F
Falsified
Can one Jev question correctly sort the router's own list of trigger words?
454
I
Inconclusive
Does the router send real conversation turns to the right destination?
144
M
Mixed
How many destination choices can one request hold before accuracy drops?
720
C
Confirmed
Does writing a full description for each destination beat giving just its name?
192
F
Falsified
Which part of a destination's description helps accuracy the most?
192
B
Blocked
Does a separate check catch missing-answer requests better than a none-of-these option?
0
B
Blocked
When a request needs two destinations at once, do separate checks beat one ranked pick?
0
B
Blocked
How similar can two destinations be before the model confuses them?
0
F
Falsified
Does the order destinations are listed in change the answer?
286
C
Confirmed
Does the model give a different answer to the exact same question twice?
160
B
Blocked
Does one convincing wrong option change answers on requests that already have a right one?
0
B
Blocked
When a request needs two destinations, does the model pick both or just one?
0
C
Confirmed
How does Jev compare to a top AI model and a plain similarity search?
48
B
Blocked
Is it cheaper to pick a destination in one step or narrow it down in two?
0
F
Falsified
Which confidence number best tells you when to trust the model?
0
C
Confirmed
Do the router's five hand-set trust thresholds hold up against real data?
0
C
Confirmed
Is the model less sure when picking a destination than a descriptive label?
96
F
Falsified
Does bundling many questions into one request save money without changing answers?
696
C
Confirmed
Does a plain-English permission policy correctly gate real actions?
403
F
Falsified
Do disguised commands fool the gate into allowing bad and blocking good?
350
M
Mixed
Does the injection filter block real attacks without blocking safe text?
280
F
Falsified
Does the irreversibility gate track consequences or just scary verbs?
241
M
Mixed
Can wording changes or a self-written excuse move a gate's verdict?
360
C
Confirmed
When the input is broken or missing, does the gate fail safe?
240
C
Confirmed
Can a cheap first check filter most rows before the full check?
759
M
Mixed
Does a five-level quality score have a working middle grade?
100
C
Confirmed
Can the model tell a supported claim from an unsupported one?
100
F
Falsified
Does the format check catch real violations without flagging normal data?
180
C
Confirmed
Can the reviewer spot planted bugs, including ones needing context?
160
F
Falsified
Does verifying a claim without its real source produce false confidence?
200
C
Confirmed
Which number should a safety gate actually threshold on?
0
F
Falsified
Does the gate tell a quoted mention apart from a real command?
120
M
Mixed
Can the gate tell when it's missing the information it needs?
180
C
Confirmed
Can a gate recognize a denied goal being retried in disguise?
200
F
Falsified
Does a 5-level rating beat multiple-choice for judging wine quality?
200
C
Confirmed
How many rating levels give the best accuracy: 2, 3, 5, 7, or 10?
200
C
Confirmed
Does the rating's returned legend text ever drift from what was submitted?
160
F
Falsified
Does the shape of a rating's probabilities predict better than its mean alone?
200
C
Confirmed
How much does a rating change on its own, with no change to the input?
1,185
C
Confirmed
Does splitting a rating into parts help, or is it just fitted weights?
330
C
Confirmed
Does deciding by a rating's mean differ from deciding by its nearest level?
0
B
Blocked
Does a keep/drop decision improve when the model sees more surrounding text?
0
B
Blocked
Can a simple code check replace a yes/no question for spotting sycophancy?
0
B
Blocked
Does one yes/no question miss circular reasoning buried in a long trace?
0
F
Falsified
Does putting many rows in one request change the answers you get?
990
B
Blocked
At a fixed budget, does ranking beat a simple threshold for keeping rows?
0
B
Blocked
When the right text span is missing from the options, does the model say so?
0
C
Confirmed
Does accuracy fall as the list of choices grows toward the 255-option limit?
603
C
Confirmed
Does asking everything at once match a two-step narrow-then-read approach?
720
B
Blocked
When nothing on a list matches, does the escape-hatch option catch it?
0
F
Falsified
Does doing well on single steps predict doing well on a whole task?
3,312
F
Falsified
Do mistakes in a loop pile up the way independent errors would?
1,499
C
Confirmed
Does judgement get worse the longer a step-by-step task runs?
489
F
Falsified
Which number better flags a bad step, confidence or the top answer's odds?
150
F
Falsified
Does asking the same question many times in one request get the same answer?
20
F
Falsified
Does asking the exact same question twice get the exact same answer?
61
C
Confirmed
Is asking everything in one request as good as asking in two steps?
302
C
Confirmed
Does accuracy hold up when the input text gets very long?
60
F
Falsified
Does giving the model the whole map really break its navigation?
8,550
F
Falsified
Does structuring the question's criteria actually improve answers?
300
C
Confirmed
Does a text-based version of a game avoid locking onto one action?
228
F
Falsified
Does picking the right tool for a task beat simple keyword matching?
210
C
Confirmed
Does the order in which options are listed change the model's answer?
211
C
Confirmed
Is the model's confidence just its top answer's probability in disguise?
120
C
Confirmed
Does splitting 'unrelated' from 'neutral' change the model's answer?
14
B
Blocked
Do two human readers agree with each other before the model gets graded?
0
C
Confirmed
Does the order options are listed in change the model's answer?
14
F
Falsified
Does the model spot which headline style actually won a real click test?
1,049
C
Confirmed
Does checking several branches instead of one change a headline's category?
2,409
C
Confirmed
Does naming the target inside the question work as well as pointing to it in data?
28
C
Confirmed
Can the model tell when a headline claims more than the article says?
10
C
Confirmed
Does the overstatement score change with how much article text the model sees?
408
C
Confirmed
Are tone, urgency, and conflict-framing really three separate things?
14
M
Mixed
Does a three-way support-or-contradict check work on real news and abstracts?
10
M
Mixed
Does a 'mixed' option get used correctly, or as a hedge?
128
C
Confirmed
Can the model name a study's design and judge if it can prove a cause?
24
C
Confirmed
Does the model rate a randomized trial as stronger evidence than an observational one?
179
M
Mixed
Can the model tell if a trial was superseded, or a paper was retracted?
253
C
Confirmed
When the true source is missing, does the model say so instead of guessing?
12
C
Confirmed
Does the model agree with a real medical review's own included-study list?
136
Records are the API responses logged for that experiment.
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.