PLATFORMS Signal 391
AI failed to properly patch software flaws 74% of the time, 1Password's study warns
1Password's Off-By-1-Labs found that LLMs produced usable security patches only 26% of the time, with over half of attempts either failing to fix the vulnerability or introducing new bugs.
Engineers relying on AI-generated patches for security vulnerabilities are getting working fixes less than a third of the time, and more than half the time the patches either don't work or introduce new defects. Automated patching pipelines that skip human review are currently unsafe for production use, and teams should focus AI tooling on vulnerability discovery and triage rather than remediation.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
LLMs generated suitable patches only 26% of the time across 6,080 attempts targeting six recent vulnerabilities.
Over half (53.9%) of AI-generated patches either failed to fix the vulnerability, introduced new bugs, or both.
1Password released its FLAWED tooling on GitHub for researchers to evaluate AI-generated patch quality.
THE CLUSTER
↗