TECH Signal 494
Opus 5 feels worse to work with than Opus 4.7, 4.8, and Fable
Illustration only Photo by Peter Ivey-Hansen on Unsplash
Engineers find Opus 5 requires more babysitting and makes assumptions without checking, unlike Opus 4.7, 4.8, and Fable.
This increased need for oversight slows down development and reduces trust in the model's autonomy. Teams may need to allocate extra time for verification or revert to earlier versions for smoother workflow.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Opus 5 is more capable on benchmarks but feels worse to work with than Opus 4.7, 4.8, and Fable.
Opus 5 tends to make assumptions and update plans without asking for clarification.
Earlier versions stop and ask questions when intent is unclear, reducing the need for careful babysitting.
THE READ
What the cluster adds up to.
Engineers report that Opus 5 feels worse to work with than Opus 4.7, 4.8, and Fable despite being more capable on benchmarks. The model often makes assumptions and updates plans without asking for clarification. This behavior contrasts with earlier versions that stop and ask questions when intent is unclear.
Adopting Opus 5 therefore requires additional oversight, which engineers describe as careful babysitting. Teams must spend extra time verifying outputs and correcting unintended assumptions. This extra verification can increase development cycles and reduce the perceived autonomy of the coding agent. In practice, the cost of this oversight may offset the performance gains seen on benchmarks.
The model’s tendency to guess works poorly in real-world coding scenarios where requirements are ambiguous or incomplete. Without clear intent, Opus 5 may produce code that does not align with business logic or constraints. Consequently, the agent stops being useful when the task cannot be reduced to a self-contained benchmark.
The author links this behavior to two pressures at Anthropic and frontier labs: the drive toward self-improving AI that could bootstrap to AGI/ASI and the pressure to score highly on benchmarks. Benchmark optimization favors models that make bold, usually-correct assumptions in the face of ambiguity. Such training penalizes models that seek clarification, which is exactly what engineers value in a coding agent. Thus, the training incentives create a mismatch between benchmark performance and everyday usability.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER