Skip to content

A French Firm Promised Thirtyfold Speedup Without a New Chip: The Demo Ran on a Model a Hundred Times Smaller Than the Promise

1 min read
Share
A French Firm Promised Thirtyfold Speedup Without a New Chip: The Demo Ran on a Model a Hundred Times Smaller Than the Promise

The whole artificial intelligence race of recent years has come down to one thing: buy more chips. The French firm Kog starts from the opposite assumption - that the graphics processors companies already own are running far below capacity, and that the difference is not in the hardware but in how deeply someone was willing to understand it.

In May the firm reached the front page of Hacker News with a technical write-up meant to prove that „extremely fast single-query decoding is possible on the standard data-centre GPUs enterprises already have”. The demonstration ran on AMD MI300X and Nvidia H200 - cards already sitting in server rooms, no new orders and no waiting.

The reaction was immediate. „We had 200 concrete business contacts,” chief executive Gaël Delalleau said. Not buyers - contacts. But for an eleven-person firm with a single founder, that is enough to establish that the problem it is attacking is real.

Whoever waits, pays

The reason is prosaic. Inference speed has become a bottleneck with a price tag. Users of tools like Claude Code know what it means to wait hours for a result, and Anthropic itself charges a multiple for its faster mode. In other words: speed is already sold as a separate product, and Kog wants to sell it on existing hardware.

Here is also the biggest hole in the story, and Delalleau does not hide it. The promise reads „30 times faster execution of large language models”. The demonstration showed 3,000 tokens per second per query - but with a purpose-built small model of around two billion parameters, Laneformer 2B, since released as open source. Between a small model and a genuinely large one there is a distance nobody has yet crossed with the promise intact.

Delalleau insists he will manage it. „GPUs have a bright future,” he says, explaining that new cards keep gaining memory bandwidth that is simply waiting to be unlocked. The test is coming soon and he set the date himself: he expects the first large model at tenfold speed in September, and the next funding round depends on it.

A physicist who learned to break systems

The firm's approach is explained by its founder's background. Delalleau studied solid-state physics at the École Polytechnique in France, then worked in offensive cybersecurity - a four-time finalist at DEFCON's CTF competition. From science, he says, comes „that way of thinking about understanding the laws of physics and the laws of the GPU in order to extract the maximum from them”. From hacking - the ability „to break things down at a very low level, down to assembly language and binary code, to understand how it works and use it for a purpose it was not intended for”.

The price of that approach is time. „For every new GPU we will spend several weeks or even months really getting into the details,” he admits. With eleven people, that means a small number of supported chips for the foreseeable future.

Kog is not alone in this idea - another French firm, ZML, released software that bypasses Nvidia's CUDA platform to support fast inference on competing chips. And that is no coincidence. Europe has no leading-edge chip fabs and will not have them soon, so the only remaining position is software: if you cannot manufacture the hardware, at least learn to extract the maximum from someone else's. Kog is backed by Bpifrance and the French Tech 2030 programme - state money behind technical sovereignty.

That is a strategy a Balkan reader ought to recognise. When you lack the capital to buy the best, what remains is the skill of making what you already bought perform beyond its spec sheet. The difference is that in France that skill gets state backing and a September deadline.