NVIDIA founder and CEO Jensen Huang used the company's Computex Taipei 2026 keynote on June 1 to launch Nemotron 3 Ultra, the largest model in the Nemotron family and the most intelligent open-weights model ever published by a U.S. lab. The 550-billion-parameter network ships with open weights and a permissive license aimed at enterprise developers building agentic and reasoning workloads on top of NVIDIA hardware.
Architecture and capabilities
Nemotron 3 Ultra is a hybrid Mamba-Transformer mixture-of-experts model with 55 billion active parameters per token, supporting a 1-million-token context window. NVIDIA reports throughput exceeding 300 output tokens per second on a pre-release DeepInfra endpoint, with strong reasoning performance on math, code and tool-use benchmarks. The architecture lets a single agent ingest entire codebases or hundreds of research documents without retrieval gymnastics.
Three-tier family
Ultra sits at the top of a Nano-Super-Ultra hierarchy. The Nano variant targets low-power edge deployment, while the Super model, launched in March 2026 with 120 billion parameters, is positioned for mid-range enterprise inference. Together, the three give NVIDIA a Llama-style stack that maps cleanly onto its Blackwell and upcoming Rubin platforms.
Foundry and supply chain
Production silicon for the GPUs running Nemotron continues to roll off advanced nodes at leading semiconductor foundries including TSMC, with system assembly handled by NVIDIA's Taiwan-based ODM partners including Foxconn. NVIDIA also reiterated its physical-AI partnership road map with Cosmos and GR00T extensions for robotics.
Why this matters
Nemotron 3 Ultra closes much of the open-weights performance gap with proprietary Western and Chinese models, but it still trails the very best closed systems on some reasoning benchmarks. By making the model openly downloadable, NVIDIA is betting that enterprise customers will increasingly want to host frontier-class inference on their own clusters — and on its silicon. The release also tightens NVIDIA's hold on the agentic AI stack alongside its Cosmos physical-AI and GR00T humanoid platforms.
Reporting based on coverage from NVIDIA's Computex keynote, Artificial Analysis, Decrypt and ExplainX.
