In section Startups & Technology

Kog Aims to Unlock Massive GPU Speeds with Low-Level Optimization

While the market fixates on purpose-built AI chips like those from Cerebras, French startup Kog is taking a different path. By treating standard datacenter GPUs as hardware to be reverse-engineered, CEO Gaël Delalleau believes he can squeeze 30x faster inference out of existing infrastructure through deep software optimization.

Kog Aims to Unlock Massive GPU Speeds with Low-Level Optimization

The company gained traction in May after a demonstration on AMD MI300X and NVIDIA H200 hardware proved that extreme single-request decoding is possible on commodity silicon. For enterprise users currently bottlenecked by the high latency of professional AI tools, the promise of the Kog Inference Engine (KIE) offers a way to bypass expensive hardware upgrades. Delalleau reports that early previews have already generated over 200 business leads, particularly from developers frustrated by the slow speeds of current LLM coding assistants.

Delalleau, a veteran of DEF CON’s CTF tournaments, brings a cybersecurity mindset to the challenge. His team treats GPUs not as black boxes, but as systems to be deconstructed down to assembly language and binary code. This hands-on approach requires months of engineering research for every new chip, limiting the startup's current scope to a team of 11. While the firm initially showcased its speed with a 2-billion parameter open-source model, the true test arrives in September, when the startup plans to demonstrate a 10x speed boost on a major, large-scale LLM. Success in that milestone will serve as the primary catalyst for their upcoming Series A funding round.

Share:on TelegramXFacebook

Subscribe to our newsletter

Once a week — the best stories from our editors, no ads or push notifications. Delivered Sunday morning.

Comments (0)

Leave a comment

No comments yet. Be the first!