PERFORMANCE Signal 143
New Real-SWE benchmark tests AI agents on licensed enterprise codebases
Illustration only Photo by Alexandre Debiève on Unsplash
The Real-SWE benchmark evaluates frontier AI models on tasks sourced from licensed private enterprise codebases, testing their ability to navigate company-specific complexity and business consequences.
Existing benchmarks rely on expert-generated or synthetic tasks that lack the complexity of actual production environments. Real-SWE forces agents to navigate proprietary systems, business rules, and infrastructure tooling, providing a measure of how well models handle the verbatim tasks enterprise engineers face. The initial results show a resolution rate of 38.8% for the top model, highlighting significant gaps in current capabilities.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
The benchmark uses tasks licensed from real companies rather than synthetic or expert-generated problems.
Evaluations test model-and-harness combinations, requiring agents to work across code, infrastructure, and business tools.
The top-performing model, Fable 5.1 using Claude Code, resolved 38.8% of tasks.
THE CLUSTER