E93–E108 · GROUP 7 OF 7
Headlines, claims, and sources
Whether Jev can read a short headline or claim and correctly judge its stance, its packaging, its overstatement, and whether the evidence behind it actually holds up.
Headlines and claims are short, and short text gives a model far less to work with than a full article. Before anyone builds a tool that scores headlines for bias or clickbait, or checks whether a study can actually back up a claim, we needed to know where the model reads real signal in the text and where it reads nothing at all. Several of these checks also had a source of truth we did not write ourselves, such as a real click-test result or a Cochrane review's own decisions, which is rare and worth testing carefully.
108 experiments, 38,194 API responses logged, 45,861,280 input tokens, $1.9262 at $0.042 per million.
Compiled from dependency-docs-reference/jev/notes/12-experiments on 2026-09-19.