AI Signal 46
Proposed monitoring scheme simulates blocked actions instead of blocking monitors
Illustration only Photo by Vishnu Mohanan on Unsplash
An argument proposes replacing blocking monitors with a scheme that simulates blocked actions during AI evaluation, as the absence of blocking monitors during OAI's cyber evaluations was viewed positively.
For engineers building AI evaluation infrastructure, this proposes shifting from monitors that prevent actions to simulation-based monitoring that evaluates what would happen. Removing blocking monitors entirely is acknowledged as infeasible, so the alternative matters for how safety evaluations get implemented.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Blocking monitors were not active during OAI's cyber evaluations, and their absence was viewed positively.
Labs ideally would stop using blocking monitors until models pose takeover risk, but this is considered infeasible.
A different monitoring scheme should simulate blocked actions rather than preventing them outright.
THE CLUSTER