SECURITY Signal 142
Pausing external evaluations and adding real-time escape detection for Claude models
The team paused external cyber evaluations of pre-release models and deployed a classifier that blocks attempts to escape sandboxes or gain unexpected internet access.
These changes reduce the risk that a model could reach live systems or the internet during testing, improving containment. They also provide a framework for third-party evaluators to assess safety without relying on a single layer of defense.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
External and internal cyber evaluations were paused to allow hardening of containment and monitoring measures.
A real-time classifier blocks model escape attempts and alerts humans, while automated monitors scan transcripts for sandbox issues.
High-risk internal sandboxes were moved to more robust isolation and the virtualization stack is being red-teamed to find weaknesses.
THE CLUSTER