E29–E46 · DESIGN FILE, AS WRITTEN
Select and Route
18 experiments, 2,988 records and 12,365,109 input tokens.
C
Confirmed
5
F
Falsified
5
M
Mixed
1
I
Inconclusive
1
B
Blocked
6
We tested whether Jev can pick the right destination for a request as reliably as the hand-built matching code it would replace.
Reading the record…
Experiments in this group
Exp
Verdict
Question
E29
F
Falsified
Can one Jev question correctly sort the router's own list of trigger words?
E30
I
Inconclusive
Does the router send real conversation turns to the right destination?
E31
M
Mixed
How many destination choices can one request hold before accuracy drops?
E32
C
Confirmed
Does writing a full description for each destination beat giving just its name?
E33
F
Falsified
Which part of a destination's description helps accuracy the most?
E34
B
Blocked
Does a separate check catch missing-answer requests better than a none-of-these option?
E35
B
Blocked
When a request needs two destinations at once, do separate checks beat one ranked pick?
E36
B
Blocked
How similar can two destinations be before the model confuses them?
E37
F
Falsified
Does the order destinations are listed in change the answer?
E38
C
Confirmed
Does the model give a different answer to the exact same question twice?
E39
B
Blocked
Does one convincing wrong option change answers on requests that already have a right one?
E40
B
Blocked
When a request needs two destinations, does the model pick both or just one?
E41
C
Confirmed
How does Jev compare to a top AI model and a plain similarity search?
E42
B
Blocked
Is it cheaper to pick a destination in one step or narrow it down in two?
E43
F
Falsified
Which confidence number best tells you when to trust the model?
E44
C
Confirmed
Do the router's five hand-set trust thresholds hold up against real data?
E45
C
Confirmed
Is the model less sure when picking a destination than a descriptive label?
E46
F
Falsified
Does bundling many questions into one request save money without changing answers?
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.