WEB Signal 123
Obfuscated Adversarial Training produces activations that evade detectors yet still cause harmful outputs
Obfuscated activations are internal model states adversarially optimized to appear benign, evading activation-based detectors while still leading to harmful behavior.
Activation-based safety detectors can be bypassed by adversarially optimized internal states, so monitoring activations alone is insufficient. Obfuscated Adversarial Training aims to train the model itself, but the result is that harmful behavior persists despite appearing benign. Engineers building safety mechanisms should consider that obfuscated activations undermine detection-based approaches.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Obfuscated activations are internal model states adversarially optimized to appear benign.
These activations evade activation-based detectors while still producing harmful outputs.
Obfuscated Adversarial Training (Bailey et al., 2024) trains the model rather than the monitor.
THE CLUSTER
↗