ELSEIF
Your brief EB
249 stories from 71 feeds 36 clusters Refreshed 1 minute ago next pull 18:05

TECH Signal 566

What's the largest software project AI can complete on its own?

WHY IT MATTERS

Conventional software-engineering benchmarks cap AI spending at roughly $1–10 per task, which may understate what models can do end-to-end; MirrorCode instead spent $2,600 on a single run that had the AI working for 19 days without human help, repositioning compute allowance as a binding constraint on measuring autonomous coding capability. The practical consequence is that claims about AI coding limits depend heavily on how much wall-clock and dollar budget the evaluator is willing to grant. Only one feed is carrying this so far, so the framing is early and the empirical results are not yet corroborated across coverage.

Written by elseif from the cluster below · every claim links back to a source

The three things worth knowing

01

Epoch AI released a study called MirrorCode aimed at measuring the upper bound of autonomous AI software development.

02

Its methodology deliberately raises the inference budget, with one cited run costing $2,600 and running for 19 days of uninterrupted AI work, in contrast to the roughly $1–10 ceiling common in existing code benchmarks.

03

The piece frames 'scale-aware' evaluation — how much compute a model is allowed — as itself a variable that shapes conclusions about AI coding reach.

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
Hacker News What's the largest software project AI can complete on its own? Open ↗