The race for faster AI inference continues, and the market warmly welcomed Cerebras and its proprietary chip in its IPO debut in May. But French startup Kog is betting it can squeeze even more power out of traditional GPUs.
The startup made the front page of Hacker News in May with a technology preview aimed at proving that “very fast single request decoding is possible on standard data center GPUs that enterprises already own,” such as the AMD MI300X GPU and NVIDIA H200 GPU used in the demo.
While some were disappointed to hear that this wouldn’t extend to laptop GPUs, others saw the potential. At a time when inference speed and cost are critical bottlenecks, Kog’s promise to unlock new capabilities on existing hardware through software optimization has garnered attention from more than just onlookers. “We had 200 concrete business leads,” CEO Gaël Delaleau told TechCrunch.
Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran users of Claude Code are well aware that you may have to wait for hours before getting results. Anthropic themselves understand that speed is worth money, and they charge many times the price for Claude’s Fast mode.
Kog wants to target customers who typically rely on AI workflows for specialized tasks and are therefore discouraged by these delays. But the startup also has design partners that allow users to generate games and apps with prompts, and Delaroux said the faster results, thanks to the Kog Inference Engine (KIE), means more revenue.
The company recognizes that this market is not yet fully mature. As Kog observed demand, he learned that potential customers weren’t ready to fine-tune a smaller model. “That’s why we’ve been committed to accelerating the development of our larger models since launch to meet the demand we’ve seen.”
This forces Kog to take a huge leap forward to deliver on its promise of “30x faster LLM inference.” The demo showed impressive results of 3,000 tokens per request per second (TPS). However, Laneformer 2B, which is now open source, used a proprietary, smaller model with only about 2 billion parameters.
Contrary to skeptics, Delalleau believes the same approach will work just as well in an LLM. The size of the LLM can be a challenge for inference chips. “GPUs have a bright future,” he said. For Kog’s CEO, the idea that Kog is not very good at decoding is a misconception. Newer GPUs have more and more memory bandwidth to unlock.
Kog isn’t alone in thinking that software optimizations can allow GPUs to do more than what’s written on the package. Another French company, ZML, has released hardware-agnostic software that bypasses Nvidia’s CUDA and supports fast inference between competing chips. But Delalleau said Kog is similar to Hazy Research, a lab at Stanford University, with a deeper level of focus on GPU acceleration.
Delalor is not an academic himself, and his first startup, TechCrunch50 2009 alumnus Stribe, has no connection to his new startup, other than his former co-founder turned VC Kamel Zeroual. His firm, Varsity VC, co-led Kog’s seed round. But the startup’s deep-level focus stems from his unique background.
After studying solid state physics at the Ecole Polytechnique in France, he went on to work in the field of offensive cybersecurity (also known as white hat hacking). According to Delaroe, this shaped the mindset he now encourages teams to adopt. On the science side, “the idea is to understand the laws of physics and the laws of the GPU and make the most of them.”
When it comes to hacking, the four-time finalist in DEF CON’s CTF tournament said he was taught to “reverse engineer things at a very low level, all the way down to assembly language and binary code, to understand how it works and try to use it to accomplish goals that weren’t necessarily designed for.”
The disadvantage of this approach is that it is very hands-on and time-consuming. “With each new GPU, we spend weeks, and sometimes months, digging deep into the details and performing GPU engineering studies on that hardware.”The team is 11 people, which limits the number of chips Kog can handle, at least for the time being.
In the long term, Kog hopes to incorporate its methodology into agent-based pipelines to support more chips and models. This could be a sovereign tailwind for the startup, which is already backed by Scaleway and supported by France’s Bpifrance and French Tech 2030 programs, as Europe looks to build its own capabilities in these two areas.
But for now, Kog needs to prove to the world that its approach works for LLMs. This is also the key to securing more funding. “Once we implement our first major model at 10x speed, which I think will be in September, we’ll start to demonstrate customer traction and be able to raise our Series A from there,” Delaroux said.
If you buy through links in our articles, we may earn a small commission. This does not affect editorial independence.
