Friday, August 14, 2026 SourcesAbout🌓
🇺🇸 US ▾
BREAKING
Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials Podcast: One country musician is calling for other artists to oppose assault rifles Bidets save you money and reduce waste — we tested the best options out there 50+ products to make your life easier and our planet cleaner Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love A head-to-toe guide of how men should dress this spring, and where they should shop 42 of the most useful travel products you can buy on Amazon The 7 best high-yield savings accounts of April 2023 Taxes are due tomorrow. Here's how to file for an extension Dominion still has pending lawsuits against election deniers such as Rudy Giuliani and Sidney Powell Russia is 'going backwards' in equipment and deploying post WWII-era tanks, according to Western officials Podcast: One country musician is calling for other artists to oppose assault rifles Bidets save you money and reduce waste — we tested the best options out there 50+ products to make your life easier and our planet cleaner Mother's Day is around the corner. Here are 50+ thoughtful gifts she'll love A head-to-toe guide of how men should dress this spring, and where they should shop 42 of the most useful travel products you can buy on Amazon The 7 best high-yield savings accounts of April 2023 Taxes are due tomorrow. Here's how to file for an extension
Technology

Kog is going deeper to squeeze more inference out of GPUs

TechCrunch ·
Kog is going deeper to squeeze more inference out of GPUs

The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there’s a lot more power to be squeezed out of conventional GPUs.

The startup hit the front page of Hacker News in May with a tech preview aimed at proving that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own” — such as the AMD MI300X and NVIDIA H200 GPUs it used for its demo.

Some were disappointed to hear this didn’t extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kog’s promise to unlock new capabilities on existing hardware with software optimization attracted more than onlookers . “We had 200 tangible business leads,” CEO Gaël Delalleau told TechCrunch.

Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and charges a price multiple for Claude’s Fast Mode.

Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said.

The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune small models. “And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen.”

This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, the now open-sourced Laneformer 2B .

Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. “GPUs have a bright future,” he said. For Kog’s CEO, the idea that they aren’t well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked.

Kog isn’t alone in thinking that software optimization can help GPUs do more than it says on the box. ZML , also from France, released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research , with an even deeper-level focus on GPU acceleration.

Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alum Stribe , has nothing to do with his new one — other than his former cofounder turned VC Kamel Zeroual, whose firm Varsity VC co-led Kog’s seed round.

Read the full article on TechCrunch ›

5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on techcrunch.com — the content belongs to TechCrunch.

More from TechCrunch

See all ›

More in Technology

See all ›