French Startup Kog Accelerates AI Model Performance on Standard GPUs by 30x

Photo: TechCrunch
Quick answer
French startup Kog is developing software to accelerate LLM inference on standard GPUs by up to 30 times. The technology optimizes memory and computational resource usage, reducing costs and improving performance for…
French startup Kog has entered the market with an ambitious goal: to prove that standard GPUs used in data centers can handle large language model (LLM) inference far more efficiently than previously thought. In May, the company showcased its technology, achieving speeds of 3,000 tokens per second using a small Laneformer 2B model with 2 billion parameters. Kog’s founder and CEO, Gaël Dellalau, asserts that similar optimizations can be applied to larger models, potentially transforming how AI infrastructure is deployed.
Kog focuses on software optimization to maximize the potential of existing hardware. According to Dellalau, many companies face delays when working with models like Anthropic’s Claude, where waiting for results can stretch into hours. The startup has already attracted interest from over 200 potential clients, including game developers and applications where generation speed directly impacts revenue. Kog’s near-term plans include accelerating large model performance by 10x, a key milestone for securing Series A investment.
Kog’s approach is rooted in deep GPU architecture analysis and low-level optimization. The startup’s founder, who previously worked in cybersecurity and participated in reverse-engineering competitions, has applied this expertise to software development. However, this method requires significant time investment: adapting the technology for each new chip takes weeks or even months. In the long term, Kog plans to automate the process using agent-based systems to expand model and hardware support.
The market already includes competitors like French startup ZML, which offers hardware-agnostic software for inference acceleration. However, Kog positions itself closer to scientific research, akin to projects from Stanford’s Hazy Research lab. Support from European investors, including Bpifrance and Scaleway, underscores the project’s strategic importance for the region’s technological independence.
Common questions
- What is inference in the context of AI?
- Inference is the process of using a trained AI model to generate predictions or responses based on input data. For LLMs, this involves text generation.
- Why is accelerating GPU inference important for businesses?
- Faster inference reduces latency in AI applications, which is critical for real-time use cases. It lowers operational costs and enhances the competitiveness of AI-driven solutions.
- Which GPUs does Kog support?
- Kog is currently testing its technology on AMD MI300X and Nvidia H200 GPUs. Support for additional chips is planned for the future.
Dzen feed: /feed/dzen.xml · RSS: /feed.xml