The question comes before the experiment
TASTE asks whether models choose the research proposals that experienced AI safety researchers prefer. Its authors assembled 92 pairs and report that tested models trailed their estimated human baseline. The target is agreement about promising work, not the eventual results of doing it. [1]
- Measured: researcher preference
- Still open: eventual discovery
Who decides which proposal is better?
The paper describes researchers rating model-generated proposals, discussing disagreements in pairs and revising judgments. Further filtering produces the benchmark. The labels therefore reflect an organized judgment process, rather than an answer available from a calculator. [2]
Why a score needs its context
The overview cautions that the limited set of pairs produces wide uncertainty around individual model scores. It also explains that its human baseline is an estimate. A league table stripped of those details would invite stronger conclusions than this experiment supports. [1]
What this means for using AI advice
Our interpretation: an appealing recommendation and a well-chosen research direction are different things. If AI suggests what to investigate next, ask it to expose assumptions, alternatives and a way to test the idea. This study alone does not establish how any particular advice will turn out.
Go a little deeper
Optional reading · about 1 more minute
Could agreement be the wrong target?
A thought experiment: every reviewer prefers a familiar approach, but an unconventional one later succeeds. Agreement would reward the familiar choice at proposal time. This hypothetical illustrates why predicting preferences and predicting discoveries should not be treated as interchangeable.
What was actually reviewed?
We checked the authors’ overview and the paper’s opening methodology. Both come from the same team. Our account does not turn that into independent validation, and it makes no claim that the benchmark measures every kind of research judgment.
Original sources
Attributed synthesis, not original reporting. Examples labeled hypothetical or illustrative are explanatory. Reviewing a source does not independently validate its findings.
- Anthropic: TASTE research overview ↗
August 28, 2026. Benchmark construction, evaluation and uncertainty sections read September 12. Developer-authored report.
- TASTE paper: Baig, Joren and Benton ↗
Abstract and introduction plus sections 2.1–2.2 reviewed September 12, 2026. Same research team as the overview; not independent corroboration or a full appendix audit.
