AI Signal 397
Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same provider can make it output the traces in plaintext (Will Knight/Wired)
Researchers demonstrated that encrypted reasoning traces from frontier AI models can be extracted in plaintext by feeding them to a weaker model from the same provider.
This technique undermines the assumption that a model's internal reasoning remains opaque, which matters for anyone relying on reasoning-trace secrecy as a security or IP boundary. The finding affects multiple major model families, Claude, GPT, and Gemini, suggesting a structural vulnerability rather than an isolated flaw.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Researchers extracted reasoning traces from Claude, GPT, and Gemini by passing a frontier model's encrypted traces to a weaker model from the same provider.
The weaker model output the reasoning traces in plaintext, bypassing the intended obfuscation of the frontier model's internal reasoning.
The vulnerability spans multiple providers, indicating a shared architectural weakness in how reasoning traces are protected.
THE CLUSTER
↗