ELSEIF
Your brief EB
331 stories from 95 feeds 230 clusters Refreshed 4 minutes ago next pull 01:21

PLATFORMS Signal 235

DeepSeek overtakes Google on volume, cost per token falls 13.6%

DeepSeek surpassed Google to become the second-largest AI lab by token volume on the AI Gateway while the average price per token fell 13.6% due to increased use of cheaper models.

WHY IT MATTERS

Engineers building cost-sensitive inference pipelines can now rely on DeepSeek’s V4 Flash for higher volume at lower per-token cost, shifting budget allocation. The rise of open-weight models like Kimi K3 and GLM 5.2 gives teams evaluating agent workloads new cheap alternatives that match the token density of frontier closed models. However, the spend share of open-weight models remains low, so cost savings are limited to workloads that fit their token-heavy, latency-tolerant profiles.

Written by elseif from the cluster below · every claim links back to a source

The three things worth knowing

01

DeepSeek’s token volume exceeded Google’s by more than double, making it the second-largest lab on the AI Gateway in July 2026.

02

The average price paid per token fell 13.6% as overall token volume grew 59% while spend rose only 37%, driven by cheaper model selections.

03

Open-weight models such as Kimi K3 and GLM 5.2 now provide agent-level token usage at low cost, shifting workloads from closed-weight frontier labs.

THE READ

What the cluster adds up to.

ORIGINAL ANALYSIS

DeepSeek moved from under 1% of gateway token volume in April to a quarter of total volume by July, surpassing Google’s 11% share. This shift was driven by consumer-assistant workloads, where Google’s share fell by more than half and most of the lost volume went to DeepSeek. DeepSeek’s V4 Flash model alone accounted for nearly a fifth of all tokens processed on the gateway in July.

Overall token volume on the AI Gateway rose 59% in July while spend increased only 37%, causing the average price per token to drop 13.6%. The decline came from customers routing more inference to cheaper models, exemplified by the shift to GPT-5-Nano for OpenAI traffic. Even if the June model mix had been held constant, the average price would have held essentially flat instead of declining. This trend lowered the cost basis for high-volume inference pipelines.

Open-weight models increased their share of gateway token volume from 11% in April to 36% by July, while their spend share rose from under four cents to nearly nine cents per dollar. Kimi K3, released mid-July, quickly scaled to heavy agent workloads, reaching token usage comparable to Claude Opus 4.8 by month end. GLM 5.2 showed a similar trajectory, and together they delivered agent-level token density at a fraction of the price of frontier closed models. As a result, teams building long-horizon agents now have a low-cost alternative that was previously unavailable at scale.

Despite the volume growth, open-weight models still accounted for less than nine cents of every gateway dollar, indicating that spend remains dominated by frontier labs. Google’s share of token spend fell as it lost volume to DeepSeek, but little of that spend migrated to DeepSeek’s cheap models. Consequently, cost savings from adopting the newest open-weight models may be limited to workloads that fit their specific token-heavy, latency-tolerant profiles.

Written by elseif from the cluster below · checked for specifics the sources never contained

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
Vercel DeepSeek overtakes Google on volume, cost per token falls 13.6% Open ↗