On July 20, Infinity.inc announced a $15 million seed round at a $100 million post-money valuation, but the real news is not the capital, it is that the company is already generating millions of dollars in annual recurring revenue from a single commercial partner, d-Matrix, within its first year. Infinity has solved a problem the entire chip industry has been trying to dodge: how to make inference software that works on any accelerator, not just Nvidia's GPUs running CUDA. The company's flagship agent, Ignition, writes production-grade kernel code automatically, testing and debugging as it goes, compressing a process that historically consumed months into single-digit days. When Infinity's agents got their first access to d-Matrix's Corsair chip, they hit 92% of its theoretical peak performance in 10 hours and had three frontier models running end-to-end in 10 days. That is not a lab benchmark. That is a commercial deployment on a chip no developer had targeted before.
The scale of the opportunity is hiding in plain sight: inference is on track to represent two-thirds of all AI compute spending in 2026, but Nvidia controls an estimated 80% of data-center AI accelerator share. That concentration looks like dominance until you realize it is dominance almost entirely through software lock-in, not hardware superiority. AMD, Qualcomm, and AWS all build chips with performance that meets or exceeds Nvidia's specs. d-Matrix's Corsair chip does real work. The problem is that writing optimized inference stacks for these alternatives has always been a custom engineering contract, hire a team, build for months, pray it hits production-grade latency. Infinity's bet is that this process is automatable. The company's seed investors include Touring Capital, Principal VC, and chip executives, plus researchers from OpenAI and Anthropic. The specificity of those backers signals something: they have tested this on actual silicon.
The proof point is not a press release about future plans. On a Qwen3-8B model, Infinity's agents boosted inference throughput from roughly 1,400 tokens per second to more than 20,000 tokens per second, a 14x gain in a single day, outperforming the widely used vLLM framework by more than 34%. That result was not achieved by hand-tuning or throwing more hardware at the problem. It was achieved by Infinity's autonomous agent learning a new chip's hardware topology, querying the chip's constraints in real time, and generating code that the chip had been theoretically capable of executing all along but no human developer had yet written. The implication cuts both ways: it means the hundreds of billions of dollars in non-Nvidia accelerator silicon scattered across data centers worldwide has been dramatically underutilized, and it means anyone who can automate the software layer between hardware and inference models wins control over which chips get deployed next.
AMD is not sitting still. On July 17, the company announced that FastFlowLM has joined AMD to advance AI inference, an explicit move to add inference optimization to AMD's software stack. That decision was made after AMD watched Infinity's approach work live with d-Matrix. The competitive shape is now clear: the hardware makers have spent years building chips that match or beat Nvidia on raw compute density. What they were missing was the automated software layer to extract that performance. Infinity has now proved that layer is buildable and profitable. For Nvidia, this is the inflection point where dominance through network effects (more models target CUDA because more hardware runs CUDA) starts to weaken. It is not an existential threat to Nvidia's GPU business, the company's fundamental advantage in training and high-end inference remains intact. But for inference deployment on edge accelerators, custom silicon, and hyperscaler chips, the software moat just cracked.
The markers to watch are straightforward. First, whether Infinity's revenue ramp continues beyond d-Matrix and whether other chip makers sign deployment agreements with the company, Qualcomm and AWS are the obvious next names. Second, how much engineering effort Nvidia invests in keeping CUDA ahead of Infinity's universal inference layer, because if CUDA needs constant upgrades to stay competitive, that is a sign the software lock-in is slipping. Third, the actual deployment velocity: how fast do companies with non-Nvidia chips currently sitting idle in data centers begin bringing them back online now that the software barrier is removable. The inference market is worth two-thirds of all AI compute spending this year. Infinity has just handed every chip maker except Nvidia a usable tool to capture a piece of it.
