Fei-Fei Li's World Labs Debuts Atlas Omni World Model For 3D Generation And Robotics Sim

World Labs' new Atlas model natively handles text, images, video and 3D, beating specialist reconstruction and camera-control baselines on both tasks.

Fei-Fei Li's World Labs Debuts Atlas Omni World Model For 3D Generation And Robotics Sim

World Labs, the spatial-intelligence startup founded by Stanford AI pioneer Fei-Fei Li, has released Atlas, a next-generation "omni world model" that generates, reconstructs and simulates 3D environments from text, images, video and depth inputs. The San Francisco company unveiled Atlas in early access on September 1, 2026, positioning it as the foundation for its Marble creation platform and for future robotics simulation workflows.

What Atlas actually does

Atlas is a multimodal autoregressive diffusion transformer pre-trained from scratch on text, images, video and 3D depth maps. World Labs says it can produce pixel-perfect, camera-controlled imagery and video at up to 1440p and one minute in length, plus explicit 3D outputs such as point clouds and Gaussian splats. From as few as one input image, Atlas can extrapolate an entire scene from any camera angle while maintaining 3D consistency across every generated view.

Beating specialist models on their own turf

On sparse-view 3D reconstruction benchmarks, Atlas posted a 25.3 mean absolute-relative pointmap error, outperforming five specialist baselines including Pi3X, VGGT-Ω 1B and Depth Anything 3, according to the paper published alongside the release. Third-party human raters preferred Atlas over rival video models such as Google DeepMind's Gemini Omni Flash, MiniMax H3 and FLUX 3 in 75% to 94% of camera-control trials. World Labs argues the omni architecture, which grounds every image in a shared 3D "spatial context," is what lets one model beat specialists at both generation and reconstruction.

World Labs Atlas world model interactive scene illustration

Robotics is the payoff

The strongest commercial angle is Real-to-Sim-to-Real for robotics. Atlas can reconstruct a warehouse, retail floor or home from a handful of phone-video frames, then serve a simulated robot the RGB and depth feeds it would see from its own body cameras as it moves through the scene. That directly targets the training-data bottleneck now driving billions into physical-AI investments such as Nscale's $3.5B Figure deal and Kinetic's robotics training-data marketplace.

Early access and what comes next

Atlas is not a public product. World Labs is admitting select partners through a request form and says the model will power future versions of Marble, its browser-based world-generation product. The company has not disclosed pricing, an API launch date or a public benchmark suite beyond the paper. Backed by roughly $230 million from Andreessen Horowitz, Radical Ventures and NEA at a reported $1 billion valuation last year, World Labs is also hiring across research and engineering as it pushes further into scaling world models toward the general spatial-intelligence stack Li has spent the last decade advocating.

Reporting based on coverage from World Labs, SiliconANGLE, Techstrong.ai and Startup Fortune.

Category: Machine Learning

Tags: AI world simulation embodied AI robotic simulation AI Foundation Models Generative AI

Related Articles