French AI startup Kog claims massive speed gains on standard GPUs
Kog says its software can boost inference speed on existing datacenter GPUs, promising up to 30-fold acceleration for large language models.
Kog, a French AI startup, showcased a software engine that extracts far more performance from conventional datacenter GPUs such as AMD’s MI300X and Nvidia’s H200, achieving 3,000 tokens per second on a 2-billion-parameter model. The demo sparked enthusiasm and generated roughly 200 concrete business inquiries, according to CEO Gaël Delalleau, who sees software-centric use cases as the initial market. While the current proof-of-concept used a small open-source model, the firm intends to apply the same optimization to larger language models, promising up to a 30-fold boost in inference speed.
Delalleau, a former white-hat hacker with a physics background, emphasizes deep low-level GPU research, though the small 11-person team limits the number of chips they can target at once. Design partners in gaming and app generation stand to benefit from faster outputs, and the startup hopes a successful 10-times speed demo in September will secure a Series A round. Backers include Scaleway, Bpifrance and French Tech 2030, positioning Kog within Europe’s push for AI sovereignty.
Why it matters
If proven, the tech could cut AI inference costs and speed up applications without new hardware purchases.
In this story
