INFRA Signal 419
French startup Kog claims software optimisation of existing GPUs can cut single-request AI inference latency
Kog proposes to accelerate AI inference by optimising software on existing GPUs rather than deploying new accelerator chips, targeting lower latency for single-request workloads.
Lower inference latency can shorten iteration cycles for developers building coding agents or generative tools, directly affecting productivity. A software-only approach avoids the capital expense and lead time of buying new chips, offering a faster path to performance gains. However, the benefit depends on how well the optimisation transfers across different GPU architectures and larger models.
Written by elseif from the cluster below · every claim links back to a sourceThe three things worth knowing
Kog demonstrated 3,000 tokens per second on a single request using its Laneformer 2B model on AMD MI300X and Nvidia H200 GPUs.
The company claims the possibility of 30-times faster large-language-model inference but has only shown this on a small, purpose-built model.
With only 11 engineers, Kog can support a limited set of GPUs and models, requiring weeks of manual work for each new hardware target.
THE READ
What the cluster adds up to.
Kog asserts that the next leap in AI inference will come from understanding existing GPUs more deeply rather than replacing them with specialised chips. The startup’s pitch concentrates on single-request decoding, aiming to make one demanding user’s response arrive faster. It demonstrated this on conventional data-center hardware, including AMD MI300X and Nvidia H200 GPUs. The demonstration used its Laneformer 2B model, achieving 3,000 tokens per second for a single request.
Adopting Kog’s method requires low-level GPU engineering, assembly code, and reverse-engineering of each target accelerator. CEO Gaël Delalleau said the team may spend weeks or months studying a new GPU. With only 11 employees, the company can support a limited selection of chips and models for the foreseeable future. This manual effort contrasts with the quicker deployment of a software update once the optimisation is written.
The approach’s differentiation is also its limitation: optimisation tuned to one model and one processor may not generalise as architectures change. Moving the technique to larger models that customers actually want introduces greater memory use and more complex decoding behaviour. Kog has not yet shown repeatable gains on widely used models under realistic concurrency. Until such evidence appears, the claimed 30-times faster inference remains an attention-grabbing ceiling rather than a reliable planning figure.
Kog hopes to embed its optimisation into agent-based pipelines that could adapt the system across more hardware, but this is still a future ambition. Institutional backing from Scaleway, Bpifrance and the French Tech 2030 programme gives the effort a European sovereignty angle. Software that makes several vendors’ GPUs more productive could lessen dependence on any single accelerator roadmap. The decisive test will be performance on production-scale models with realistic workloads, which will determine whether the latency benefit translates into commercial traction.
Written by elseif from the cluster below · checked for specifics the sources never containedTHE CLUSTER
↗