Santa Clara chip startup d-Matrix, presenting jointly with Meta at Hot Chips 2026, showed off Raptor — a generative-AI inference accelerator that abandons HBM entirely and instead bonds a TSMC 4nm compute die face-to-face on top of a custom-designed DRAM die at a 36-micron pitch. The result is 32 GB of on-package memory reaching roughly 100 TB/s per card, at a fraction of HBM4's energy cost.
3D-Stacked DRAM Instead Of HBM
Raptor's headline number is memory bandwidth density. d-Matrix reports the vertical interface to the DRAM die runs at about 0.37 pJ/bit, versus roughly 2.4 pJ/bit for pushing data into an HBM4 base die — an order-of-magnitude improvement that d-Matrix's slides put at roughly 32.6 GB/s per mm² of silicon, or 20x HBM4's density. Pushing 100 TB/s through HBM would cost about 1.92 kW for I/O alone, a power budget today's packages simply do not have.
Engineered With Meta For Frontier Decode
The Meta co-presentation is telling. Generative-inference decode is memory-bandwidth-bound, and a 64-user, 1M-context workload can require roughly 935 GB of KV cache — a capacity and bandwidth problem at once. d-Matrix says a 72-card Raptor rack fits a frontier model such as Kimi K3 at 1M context in 4-bit weights and 8-bit KV cache, sustaining roughly 1,000 output tokens per second per user on a 3-trillion-parameter class model. To keep 3D DRAM stable at a 105°C junction, the company introduced pinless stream flipping DBI, thermal-aware refresh, and a bank-chaining scheme that lets 72 spare banks absorb faults anywhere on the die.
Where Raptor Sits In A Crowded Field
Raptor lands in the same Hot Chips week that Cerebras previewed CS-6 with its own wafer-scale 3D-DRAM stack, that Intel detailed Crescent Island, and that OpenAI's Jalapeño ASIC got its architectural unpacking. It also builds on d-Matrix's earlier Corsair inference platform now in full production, and the Nvidia hybrid-rack partnership announced in July. If 3D DRAM proves manufacturable at scale, the industry's HBM-versus-SRAM debate gets a third answer — and Meta, sharing the podium, is signalling it wants that answer to succeed.
Reporting based on coverage from ServeTheHome, Tom's Hardware and igor'sLAB.
