INFRA Signal 384
Kog is going deeper to squeeze more inference out of GPUs
Kog claims software optimisations can extract far higher inference speeds from existing AMD and Nvidia GPUs without hardware changes.
If proven at scale, the approach could reduce inference costs and latency for enterprises already invested in GPU infrastructure. It also challenges the assumption that GPUs are inherently inefficient for agentic workflows, potentially delaying the need for specialised AI chips.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Kog demonstrated 3,000 tokens per second on a 2B-parameter model using AMD MI300X and Nvidia H200 GPUs.
The startup targets latency-sensitive use cases like software engineering and game generation, where faster inference directly impacts revenue.
Low-level GPU optimisations require months of reverse-engineering per chip, limiting initial support to a handful of datacenter GPUs.
THE READ
What the cluster adds up to.
Kog’s core claim is that conventional GPUs can deliver inference speeds previously thought to require purpose-built chips. The demo of 3,000 tokens per second on a small model suggests that memory bandwidth, not compute, is the bottleneck, and that software can unlock it. This contradicts the narrative that GPUs are fundamentally unsuited for agentic workflows, where low-latency, single-request decoding is critical. If the approach scales to larger models, it could extend the lifespan of existing GPU fleets, reducing capital expenditure for enterprises already invested in Nvidia or AMD hardware.
The trade-off is development time. Kog’s method relies on deep reverse-engineering of each GPU architecture, taking weeks or months per chip. This hands-on approach limits the startup’s ability to support a wide range of hardware quickly, especially consumer-grade GPUs. The team’s background in offensive cybersecurity and solid-state physics informs this strategy, treating GPUs as systems to be hacked rather than black boxes to be abstracted. While effective, this methodology is labour-intensive and may struggle to keep pace with the rapid release cycles of new GPUs from Nvidia, AMD, and Intel.
Early traction suggests demand for faster inference in latency-sensitive domains. Kog’s CEO cited 200 business leads, with software engineering and prompt-driven app generation as primary use cases. These workflows often involve long wait times for results, where even a 10x speedup could meaningfully improve productivity or revenue. However, the startup’s focus on larger models, rather than fine-tuning small ones, reflects a bet on where enterprise demand is headed. This aligns with trends like Anthropic’s Fast Mode, which charges a premium for lower latency, but it also means Kog must prove its optimisations work at scale before securing further funding.
Competition in this space is emerging, with other startups like ZML pursuing hardware-agnostic software optimisations. Kog differentiates itself by targeting deeper, chip-specific optimisations, similar to Stanford’s Hazy Research. This could yield higher performance gains but also limits flexibility. The startup’s European roots may provide tailwinds, as the region seeks to build sovereign AI capabilities, but its long-term success hinges on demonstrating 10x speedups on major LLMs, a milestone slated for September. Until then, the approach remains unproven for the models most enterprises rely on.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗