AI infrastructure startup Infinity announced a $15 million seed round at a $100 million valuation on July 20, 2026, led by Touring Capital and Principal VC, with participation from individual researchers at OpenAI and Anthropic. The company is building a CUDA-alternative kernel stack that lets AI models run efficiently on non-Nvidia chips, including SRAM-based accelerators, phone SoCs and systolic arrays.
Chipping away at Nvidia's software moat
Nvidia's lead in AI is grounded not only in Blackwell and Rubin silicon but in CUDA, the general-purpose compute layer that PyTorch and TensorFlow are built on top of. Most application-level AI startups do not have the kernel engineering resources to port workloads to alternative accelerators, which keeps them locked into Nvidia's stack even when the underlying hardware is available at lower cost per token. Infinity's pitch is that its agent, Ignition, can automate that porting work end-to-end.
Ignition — an AI agent for kernel code
Ignition writes the low-level inference code, tests it against target hardware, measures throughput, and rewrites the code in a feedback loop until performance plateaus. Because it treats kernel authoring as a search problem rather than a hand-craft task, it adapts to different chip architectures without proprietary tuning. In an early case study with AI-chip challenger D-Matrix, Infinity says Ignition compressed what would have been a months-long porting effort down to hours to days.
A pay-for-performance model
Founder Jeremy Nixon, a former Google Brain researcher who created the AGI House hacker community, told TechCrunch the company charges customers based on measured performance gains and cost savings rather than an upfront licence, aligning revenue with realised throughput improvements in tokens per second. Infinity is in active discussions with several other chip and cloud providers beyond D-Matrix.
Where the money sits in AI infrastructure
The round lands as investors continue to pour capital into the layer between models and silicon, including inference clouds, compiler stacks and agent frameworks. It is a comparatively small check by 2026 standards — humanoid robotics and foundation-model labs are absorbing rounds an order of magnitude larger — but the strategic logic is clear: every hyperscaler and every foundry now has an interest in commoditising the software layer that keeps a single GPU vendor in a dominant position.
Reporting based on coverage from TechCrunch.
