ELSEIF
Your brief EB
468 stories from 219 feeds 1268 clusters Refreshed 59 minutes ago next pull 10:42

TECH Signal 133

Cognition’s SWE-2 reportedly tops Terminal-Bench 2.1 with 92.8 score at lower cost than Claude Fable 5.1

Cognition’s SWE-2, a 2.8T-parameter MoE model, achieves a self-reported 92.8 on Terminal-Bench 2.1 while claiming 64% lower cost than Claude Fable 5.1 for comparable FrontierCode performance.

WHY IT MATTERS

SWE-2 demonstrates that multi-trillion-parameter models can deliver competitive agentic coding performance at significantly reduced operational cost. However, its proprietary weights and lack of independent benchmark validation limit immediate adoption for engineers outside Cognition’s ecosystem. The gap in long-horizon tasks (Terminal-Bench 4.0) highlights where further development is needed.

Written by elseif from the cluster below · every claim links back to a source

The three things worth knowing

01

SWE-2 uses a 2.8T-parameter MoE architecture with 104B active parameters per token, post-trained from Kimi K3 with reinforcement learning at scale.

02

Self-reported benchmarks show parity with Claude Fable 5.1 on FrontierCode 1.1 (50.0 vs 50.9) at a claimed 64% lower cost, but lag significantly on Terminal-Bench 4.0 (27.3 vs 55.8).

03

No local inference or per-token API is available; SWE-2 is accessible only through Cognition’s Devin Desktop, CLI, and web platforms.

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
tokenstead.ai via Hacker News Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1 Open ↗