WEB Signal 56
Modelling variation in the METR Uplift Study
elseif has not written about this yet · Lesswrong describes it this way
Disclaimer: Not associated with METR.SummaryI reanalysed the METR uplift study under a meta-analytic model to estimate the range of plausible effects that we may see in repeated trials. If the average population effect varied as much as medical and economic Randomized Control Trials (RCTs), we would see a 95% credible interval of results from -55% to +193% even without study-specific effects. The mean population effect was between 10-20% with heavy tails. The probability of a slowdown is between 66% and 79%. The study was therefore compatible with both large speedups and slowdowns, even with no further AI progress.We should be careful about drawing inferences on the sign or magnitude of AI-assistance when study variation is high, but this may explain the radically different effectiveness heard from self-reports. It would be good to find more ways of squeezing more information out of older data, since it is unlikely for us to gain more before the point of no return (if that isn't already passed).AI speedups and slowdownsThe METR RCT remains the highest quality evidence that we have for developer productivity improvements using AI tools. The original METR uplift RCT was published in
THE CLUSTER
↗