TECH Signal 468
Self-reported big-pickle run tops all Mini-SWE-Agent entries on SWE Atlas Codebase QnA at 50.8%
A free stealth model called big-pickle, accessed through OpenCode Zen, scored 50.8% (63/124) on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold, outscoring every entry in that scaffold class on the official leaderboard.
This is a single-trial, self-reported evaluation that Scale did not verify, and the model's identity is unconfirmed. The result suggests a free model can compete with paid frontier models on codebase question-answering tasks, but the caveats, single trial, unknown model identity, and potential data exposure, mean the number should not be treated as leaderboard-equivalent without independent reproduction.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
big-pickle resolved 63 of 124 tasks (50.8%) using the mini-swe-agent scaffold, the highest score in that scaffold class on the official SWE Atlas QnA leaderboard.
The run used a single trial per task rather than the official 3-trial protocol, yielding a standard error of approximately ±4.5 points.
The model's identity is unconfirmed; leaked API signatures suggest DeepSeek infrastructure, and OpenCode states prompts during the free period may be used to improve the model.
THE CLUSTER