E63–E78 · GROUP 5 OF 7
Scoring, curating, and finding text
This group tested how to build a good rating question, whether a curation decision should discard its distribution, and how well Jev can point at one piece of text inside a longer document.
Anyone building a real pipeline with Jev has to pick between a rating, a multiple-choice question, or a plain yes/no question, and then decide what to do with the numbers that come back. Getting that choice wrong either wastes requests on a design that adds nothing, or throws away a signal a curation or search pipeline needed. Several of these tests also try to run a real curation pipeline end to end and could not, which is itself a finding about what such a pipeline needs before it can be trusted.
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.