Apple is quietly designing its first fully in-house AI server and is in early talks with NVIDIA over the chipmaker's NVLink Fusion interconnect technology, according to a September 17, 2026 report from The Information. The system is targeted for 2029 and will pair Apple's next-generation M8 Ultra silicon with NVIDIA scale-up and scale-out networking.
Two and four M8 Ultra SoCs per node
The current design contemplates two configurations — a 2x M8 Ultra module and a heftier 4x M8 Ultra system — built on an extension of Apple's Mac Pro architecture. The chips would be the successor to the M4 Ultra and represent Apple's most aggressive silicon step-up in the M-series lineage. Apple is reportedly weighing NVIDIA-provided switches, chiplets that add NVLink connectivity and a bundled software stack rather than building its own high-speed fabric.
Why NVLink Fusion, why now
NVLink Fusion is NVIDIA's plan for wiring third-party CPUs and accelerators into its rack-scale AI systems. The company has already lined up MediaTek as a launch partner and is using the same fabric to bolt AWS Vera CPUs into NVIDIA racks. Apple, which has largely stayed out of NVIDIA's ecosystem in recent years, would gain a proven scale-up path without designing its own interconnect from scratch.
Servers to power a bigger Private Cloud Compute
The Information's sources tie the 2029 server to Apple's Private Cloud Compute strategy, the encrypted server fabric that already handles overflow Apple Intelligence inference. Building custom hardware at that scale would let Apple pair the privacy story with more competitive economics as the company continues to fold generative-AI features into iPhone, iPad and Mac, following recent moves like the iPhone Duo Foldable reveal.
Not a done deal
Apple and NVIDIA have not confirmed the report, and The Information cautioned that Apple's evaluation of NVLink Fusion "is not formalized". Even so, the leak reframes Apple's AI-server ambitions from the Mac-derived Baltra prototype era into something that could plausibly compete on capacity — and, by extension, on token cost — with hyperscaler rivals by decade's end.
Reporting based on coverage from The Information, Tom's Hardware and AppleMagazine.
