ELSEIF
Your brief EB
232 stories from 207 feeds 1245 clusters Refreshed 31 minutes ago next pull 21:39

AI Signal 46

Proposed monitoring scheme simulates blocked actions instead of blocking monitors

Illustration only Photo by Vishnu Mohanan on Unsplash

An argument proposes replacing blocking monitors with a scheme that simulates blocked actions during AI evaluation, as the absence of blocking monitors during OAI's cyber evaluations was viewed positively.

WHY IT MATTERS

For engineers building AI evaluation infrastructure, this proposes shifting from monitors that prevent actions to simulation-based monitoring that evaluates what would happen. Removing blocking monitors entirely is acknowledged as infeasible, so the alternative matters for how safety evaluations get implemented.

Written by elseif from the cluster below · every claim links back to a source

The three things worth knowing

01

Blocking monitors were not active during OAI's cyber evaluations, and their absence was viewed positively.

02

Labs ideally would stop using blocking monitors until models pose takeover risk, but this is considered infeasible.

03

A different monitoring scheme should simulate blocked actions rather than preventing them outright.

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
Lesswrong Blocking Monitors are Bad Open ↗