AI Signal 447
Claude Code reportedly returns empty thinking blocks while billing for full reasoning tokens
Anthropic’s Claude Code API is returning blank or truncated thinking summaries despite charging for all generated reasoning tokens
Engineers relying on Claude’s thinking summaries for debugging or transparency are getting incomplete outputs while still incurring full token costs. The discrepancy between billed tokens and delivered content complicates cost management and trust in the API’s behavior. If unresolved, this could undermine confidence in Anthropic’s billing transparency and model reliability
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Developers report empty or truncated thinking blocks in Claude Opus 4.8 and Sonnet 5 despite explicit requests for summaries
Anthropic’s documentation confirms customers are billed for all thinking tokens generated, even if the output is redacted or missing
The issue affects both API responses and integrations like Claude Code for VS Code, with no clear timeline for resolution
THE READ
What the cluster adds up to.
Claude Code’s thinking blocks, a feature designed to show the model’s intermediate reasoning, are failing to return usable summaries. Reports indicate the API delivers empty strings or truncated outputs for specific models, even when users explicitly request summarized thinking. This breaks a key workflow for developers who rely on these summaries to validate or debug model behavior. The problem appears intermittent but has persisted across multiple versions and integrations, suggesting a systemic issue rather than a transient glitch.
The financial impact of this bug is non-trivial. Anthropic’s billing policy charges for all thinking tokens generated, regardless of whether the output is delivered or redacted. This means developers are paying for reasoning they cannot see or use, with no apparent adjustment for truncated summaries. The company’s documentation explicitly states that collapsed or missing thinking blocks do not reduce costs, leaving users with no recourse for recovering lost value. For teams operating under tight token budgets, this could inflate expenses without corresponding improvements in output quality.
The root cause remains unclear, but the pattern of reports points to potential flaws in how thinking summaries are streamed or truncated. Some developers speculate the issue stems from network tuning or request termination logic, particularly for long-running sessions. Anthropic’s acknowledgment of the problem suggests it is under investigation, but the lack of a fix or workaround leaves users in a bind. Until resolved, engineers may need to disable thinking entirely to avoid unpredictable costs, sacrificing a feature that was intended to improve model performance and transparency.
This incident highlights broader risks in relying on opaque AI billing models. When token costs are decoupled from delivered output, users have little visibility into whether they are paying for useful work or wasted computation. The discrepancy also raises questions about how Anthropic handles edge cases in its API, whether truncation is a bug or an intentional design choice. For now, developers must weigh the benefits of thinking blocks against the risk of silent failures and unadjusted bills, a trade-off that undermines the feature’s intended value.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER