Models

Opus 5.5, Fable 5.1, GPT-6 Astra, DeepSeek V4.1 and Grok 4.7: pricing and speed

October 5, 2026 · 5 min read

The illuminated skyline of Los Angeles at night under a deep blue sky
Henning Witzel / Unsplash

Opus 5.5, Fable 5.1, GPT-6 Astra, DeepSeek V4.1 and Grok 4.7 were released in early September 2026. The announcements came from Anthropic, OpenAI and other developers within a few days. The article explains how to interpret the numbers that companies publish, where the figures differ and how to calculate the cost of solving a task rather than the cost per token. First, the author separates the numbers that producers report from independent measurements. Company tables are biased; they can choose tests, reasoning levels, and prompt templates to favor their own results. For example, Terminal-Bench 4.0 yields 37.6 % for Grok 4.7 according to xAI and 33 % according to Artificial Analysis, showing that raw tables cannot be compared directly. Independent aggregators such as Artificial Analysis run the same ten‑test suite for all models and also measure output speed, time to first token, output tokens and task price. Their index averages ten tests and adds separate speed metrics; the index is limited because it is an average temperature across a hospital. A third category consists of direct side‑by‑side comparisons performed by blogs and journalists who pick one or two tasks and measure time, dollars and answer quality. These comparisons are closest to real‑world use but involve only a handful of tasks, so they are not statistically robust. The risk is that dozens of rewrites appear on aggregators within a day; index differences for a single model can reach 15 points and published prices may be missing, as with GPT‑6.1 Sol where OpenAI does not list a price. The only reliable way to resolve such gaps is to consult primary sources: official announcements, pricing pages and model listings on Artificial Analysis. A separate section discusses reasoning levels, which range from low to max; a single model can behave differently on different levels. For instance, Opus 5.5 low versus max shows a 16‑point index gap and an eleven‑fold price difference, so the highest measured level should be used when comparing peak performance.

The second step converts the published price per word into the price of solving a task. The price per word is the list price, but the actual payment is for a completed task; the number of input words varies widely across models, and reasoning tokens are billed at the output price. Artificial Analysis calculates the average cost of a task at a given reasoning level, producing a non‑obvious pattern. The most expensive task belongs to Sonnet 5.5 at $7.67 on the highest reasoning level, followed by Fable 5.1 at $7.63 and Opus 5.5 at $5.98. Sonnet’s token price is half that of Opus and a quarter of Fable’s because it generates far more output tokens, 420 million versus a median of 81 million, while Opus produces 260 million and Fable 190 million. GPT‑6 Astra costs $3.26 per task at a high token price, half the cost of Opus, while GPT‑6.1 Sol processes 67 million tokens per run, scores 52 on the index and charges $0.72 per task, offering the best quality‑to‑price ratio among models scoring above 50. Grok 4.7 is five times cheaper than Astra on list price, but its task price is $3.74 because it outputs 81 000 tokens per task versus 36 000 for Grok 4.6. At the low‑price end, GPT‑6 Luna scores 38 points for $0.07, DeepSeek V4.1 Flash scores 39 points for $0.27 and DeepSeek V4 Pro scores 36 points for $0.67. The senior DeepSeek model is more expensive than Flash yet delivers a worse result, making it a less obvious choice after a summer price increase. The author concludes that the decisive metric is the cost of solving a task, not the list price per token; at identical token prices the final cost can differ by a factor of two. The third step aligns the Artificial Analysis index with industry benchmarks. As of 3 October, the top scores are: Claude Opus 5.5-58 points (first among 224 models), Claude Sonnet 5.5-56, Claude Fable 5.1-53, GPT‑6 Astra, 53, GPT‑6.1 Sol, 52, Grok 4.7-46, DeepSeek V4.1 Flash, 39, GPT‑6 Luna, 38, DeepSeek V4 Pro, 36. The top five spots are dominated by two U.S.

companies, with a six‑point gap that exceeds the difference between low and medium reasoning levels for a single model. Additional tests include Terminal‑Bench 4.0, which evaluates agents in a terminal environment; Sonnet 5.5 scores 70.6 %, Opus 5.5 66.4 %, GPT‑6 Astra 57.9 %, Fable 5.1 55.8 %, Grok 4.7 37.6 % and Grok 4.6 20.3 %. Results for GPT‑6.1 Sol, Luna and DeepSeek have not been published. OSWorld reports Opus 5.5 at 81.8 % and Sonnet 5.5 at 80.1 %. In AutomationBench, OpenAI data shows GPT‑6.1 Sol surpassing Opus 5.5 by 2.2 points. In the programming‑focused Terminal‑Bench, Astra leads; in the Harvey benchmark, Grok 4.7 achieves 19.6 % versus 6.7 % for Fable 5.1. Each model occupies a distinct niche, and the published findings converge on a few patterns: long‑form programming tasks favor Claude, the best quality‑to‑price ratio belongs to GPT‑6.1 Sol, Grok 4.7 shows improvement but still lags behind flagship models, and DeepSeek remains the strongest open‑source option, though independent measurements fall short of its own claims. To avoid most discrepancies, the article recommends calculating task cost from token consumption, always checking the reasoning level, and verifying figures against primary sources. Limitations remain: the aggregated index averages niche models, direct comparisons rely on tiny sample sets and some models lack published independent measurements at the required reasoning level. A full version of the analysis will also cover printing speed, time to first token, fast modes, document length handling and performance of models hosted in Russia.