Skip to content

Eleven People Claim the Chips Companies Already Bought Can Do Thirty Times More: the Demo Worked, but With a Model Ten Times Smaller Than the Promise

1 min read
Share
Eleven People Claim the Chips Companies Already Bought Can Do Thirty Times More: the Demo Worked, but With a Model Ten Times Smaller Than the Promise

While the entire AI market chases new chips, one Frenchman with a team of eleven claims the money has already been spent - it is just that nobody knows how to use it. Kog, a startup from France, is building software that squeezes drastically faster processing out of the graphics processors companies have already bought and installed.

In May the firm made the front page of Hacker News with a technical demonstration meant to prove one thing: that "extremely fast single-request decoding is possible on the standard data-centre GPUs firms already own". The demo ran on AMD MI300X and Nvidia H200. The result was 3,000 tokens per second per request.

The reaction was bigger than expected. "We had 200 concrete business enquiries," says CEO Gaël Delalleau. For a startup with eleven employees, that is more work than it can take on.

There is just one problem, and Delalleau does not hide from it. The demo ran on a model of around 2 billion parameters - Laneformer 2B, which has since been released as open source. The company's promise is "30 times faster processing of large language models". Between a small purpose-built model and a genuinely large language model sits a leap that has not been made yet. Sceptics say it will not work. Delalleau says GPUs have a bright future and that the belief they are unsuited to decoding has become a misconception - new chips have ever more memory bandwidth, simply waiting for somebody to unlock it.

The way they work is manual to the point of pain. For every new chip the firm spends weeks, sometimes months, digging through the details of that specific hardware. With a team of eleven, that means a small number of supported chips. Delalleau is not a researcher by training - he studied solid-state physics at École Polytechnique, then worked in offensive cybersecurity and was a four-time finalist in the CTF competition at DEFCON. That experience, he says, taught him "to break things down to a very low level - down to assembly language and binary code - to understand how they work and use them for a purpose they were not designed for".

Kog is not alone in this. ZML, also French, has built software that bypasses Nvidia's CUDA. Behind Kog stand Bpifrance, the French Tech 2030 programme and Scaleway - meaning the French state is inside, which is not an accident. Europe wants its own capacity in both chips and models, and anyone promising that existing hardware can do more makes that job easier.

The next date is September. "As soon as we run the first large model at ten times the speed, we will be able to show real customers and raise a Series A off the back of it," Delalleau says. Until then, Kog is a promise with a good story behind it - eleven people, one founder and a two-billion-parameter demo against an entire market spending billions on new chips. If it works, it means part of those billions never needed to be spent at all.