PERFORMANCE Signal 448
New benchmark tests coding agents on large-scale refactoring; best model resolves 41.2%
A new benchmark evaluates AI coding agents on large-scale refactoring tasks, where the best-performing model achieves only a 41.2% resolve rate.
Most existing coding agent benchmarks avoid large-scale refactoring, leaving a blind spot in how well these tools handle complex, multi-file changes. The low top score indicates current agents are not yet reliable for significant codebase restructuring work.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
The benchmark specifically targets large-scale refactoring, an area most coding agent benchmarks skip.
The best-performing model achieved only a 41.2% resolve rate on the benchmark.
The low resolution rate suggests current AI coding agents struggle substantially with complex refactoring tasks.
THE CLUSTER
↗