Silicon Valley inference specialist d-Matrix and NVIDIA on July 8, 2026 disclosed a strategic technology partnership that pairs NVIDIA's Blackwell GPUs with d-Matrix's Corsair digital in-memory compute (DIMC) accelerators in a single hybrid rack, targeting the ballooning cost of running frontier generative AI models at scale.
From Rivals To Rack-Scale Partners
The collaboration marks a rare rapprochement between the world's dominant AI silicon vendor and one of the highest-profile inference startups aiming to disrupt it. d-Matrix, valued at $2 billion after closing a $275 million Series C in November 2025 led by BullhoundCapital, Triatomic Capital and Temasek, said Corsair will be co-scheduled with Blackwell GPUs to serve disaggregated inference pipelines that split the prefill and decode stages of large language model serving.
In independent benchmarks with Gimlet Labs earlier this year, a Corsair-plus-GPU pipeline delivered up to 10x the throughput and roughly 3x the cost efficiency of a GPU-only pipeline running Llama 70B, while cutting energy per token by 3 to 5 times.
Aiming At Inference Unit Economics
The partnership is aimed squarely at hyperscalers and neo-cloud operators grappling with runaway inference bills as chat assistants, agentic workflows and retrieval-augmented generation pipelines drive token demand far past training. "Inference is becoming the dominant cost in production AI," Triatomic Capital's Jeff Huber said when d-Matrix closed its Series C, calling out the startup's digital in-memory compute architecture as "purpose-built for low-latency, high-throughput inference workloads that matter most."
d-Matrix CEO Sid Sheth has framed the company's mission as fixing a structural AI problem: "When trained models needed to run continuously at scale, the infrastructure wouldn't be ready. We've spent the last six years building the solution: a fundamentally new architecture that enables AI to operate everywhere, all the time."
What Ships In The Hybrid Rack
The rack combines Blackwell HGX modules for the compute-bound prefill phase with racks of d-Matrix Corsair PCIe cards for the memory-bound decode phase, glued together by d-Matrix's JetStream I/O accelerator and the open standards-based SquadRack reference architecture co-designed with Arista, Broadcom and Supermicro. d-Matrix's Aviator software stack handles model partitioning and scheduling across the two accelerator classes, drawing on the company's Corsair AI inference platform, which entered full production in June for hyperscaler shipments.
Signal On A Fragmenting AI Silicon Market
The move follows a similar NVIDIA collaboration with SambaNova and echoes the company's stance that heterogeneous accelerators, not a single flagship GPU, will define next-generation inference factories. It also lands the same week as broader inference-economics news, from Together AI's $800 million Series C to Anthropic's Reflect dashboard quantifying real-world Claude consumption.
d-Matrix now counts M12 (Microsoft's Venture Fund), the Qatar Investment Authority and EDBI among its backers, and lists Toronto, Sydney, Bangalore and Belgrade among its global offices. NVIDIA's partnership, industry watchers say, moves the startup from GPU challenger to GPU complement — a pragmatic hedge as inference workloads outgrow every architecture family in isolation.
Reporting based on coverage from d-Matrix, GuruFocus, Value the Markets and AI Weekly.
