INFRA Signal 142
Princeton shadow evaluation finds AI agents lack research judgment, slowing self-improvement forecasts
A Princeton-led study using a new "shadow evaluation" method found that AI agents could handle engineering tasks but failed to produce research-quality papers when tested on open-ended questions from unpublished NeurIPS 2026 submissions.
The finding challenges forecasts of rapid AI recursive self-improvement, suggesting that automating AI research requires judgment and creativity that current agents lack. Engineers relying on AI for autonomous research workflows should expect current models to handle narrow, checkable tasks well but struggle with exploratory or open-ended work.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Researchers proposed "shadow evaluation," testing AI on research questions from unpublished papers to prevent memorization of answers.
Claude Opus 4.8 agents given six days and $3,000 in API credits produced two papers that were both rejected by the original authors.
Agents excelled at engineering tasks but lacked creativity, judgment, and ability to backtrack from failing approaches.
THE CLUSTER
↗