ELSEIF
Your brief EB
314 stories from 101 feeds 304 clusters Refreshed 13 minutes ago next pull 12:36

TECH Signal 494

Opus 5 feels worse to work with than Opus 4.7, 4.8, and Fable

Illustration only Photo by Peter Ivey-Hansen on Unsplash

Engineers find Opus 5 requires more babysitting and makes assumptions without checking, unlike Opus 4.7, 4.8, and Fable.

WHY IT MATTERS

This increased need for oversight slows down development and reduces trust in the model's autonomy. Teams may need to allocate extra time for verification or revert to earlier versions for smoother workflow.

Written by elseif from the cluster below · every claim links back to a source

The three things worth knowing

01

Opus 5 is more capable on benchmarks but feels worse to work with than Opus 4.7, 4.8, and Fable.

02

Opus 5 tends to make assumptions and update plans without asking for clarification.

03

Earlier versions stop and ask questions when intent is unclear, reducing the need for careful babysitting.

THE READ

What the cluster adds up to.

ORIGINAL ANALYSIS

Engineers report that Opus 5 feels worse to work with than Opus 4.7, 4.8, and Fable despite being more capable on benchmarks. The model often makes assumptions and updates plans without asking for clarification. This behavior contrasts with earlier versions that stop and ask questions when intent is unclear.

Adopting Opus 5 therefore requires additional oversight, which engineers describe as careful babysitting. Teams must spend extra time verifying outputs and correcting unintended assumptions. This extra verification can increase development cycles and reduce the perceived autonomy of the coding agent. In practice, the cost of this oversight may offset the performance gains seen on benchmarks.

The model’s tendency to guess works poorly in real-world coding scenarios where requirements are ambiguous or incomplete. Without clear intent, Opus 5 may produce code that does not align with business logic or constraints. Consequently, the agent stops being useful when the task cannot be reduced to a self-contained benchmark.

The author links this behavior to two pressures at Anthropic and frontier labs: the drive toward self-improving AI that could bootstrap to AGI/ASI and the pressure to score highly on benchmarks. Benchmark optimization favors models that make bold, usually-correct assumptions in the face of ambiguity. Such training penalizes models that seek clarification, which is exactly what engineers value in a coding agent. Thus, the training incentives create a mismatch between benchmark performance and everyday usability.

Written by elseif from the cluster below · checked for specifics the sources never contained

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
mun-logadan.github.io via Hacker News Why does Opus 5 feel worse to work with? Open ↗