ELSEIF
Your brief EB
203 stories from 105 feeds 339 clusters Refreshed 11 minutes ago next pull 04:07

TECH Signal 468

Self-reported big-pickle run tops all Mini-SWE-Agent entries on SWE Atlas Codebase QnA at 50.8%

A free stealth model called big-pickle, accessed through OpenCode Zen, scored 50.8% (63/124) on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold, outscoring every entry in that scaffold class on the official leaderboard.

WHY IT MATTERS

This is a single-trial, self-reported evaluation that Scale did not verify, and the model's identity is unconfirmed. The result suggests a free model can compete with paid frontier models on codebase question-answering tasks, but the caveats, single trial, unknown model identity, and potential data exposure, mean the number should not be treated as leaderboard-equivalent without independent reproduction.

Written by elseif from the cluster below · every claim links back to a source

The three things worth knowing

01

big-pickle resolved 63 of 124 tasks (50.8%) using the mini-swe-agent scaffold, the highest score in that scaffold class on the official SWE Atlas QnA leaderboard.

02

The run used a single trial per task rather than the official 3-trial protocol, yielding a standard error of approximately ±4.5 points.

03

The model's identity is unconfirmed; leaked API signatures suggest DeepSeek infrastructure, and OpenCode states prompts during the free period may be used to improve the model.

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
github.com via Hacker News Big Pickle on SWE Atlas – Codebase QnA Open ↗