Fei-Fei Li's World Labs Unveils Atlas, a Multimodal World Model
World Labs, the startup founded by AI pioneer Fei-Fei Li, launched Atlas on September 1: a world model that reconstructs 3D scenes from a single photo and simulates space and time for creators and robots.

World Labs, the artificial-intelligence startup founded by Stanford professor and AI pioneer Fei-Fei Li, unveiled a new world model called Atlas on September 1. The company describes it as an omni model, built to natively process text, images, video and 3D data inside a single system, and says it can reconstruct a real-world 3D scene from as little as one photograph, generate minute-long high-resolution video with precise camera control, and produce the depth data robots need to train inside simulated environments. Fei-Fei Li called the launch a milestone for World Labs, saying Atlas opens the door to applications ranging from visual effects to robotics.
What Atlas can actually do
- Camera-controlled generation: from one to six reference images, Atlas generates video with pixel-level camera control, outputting up to one minute of footage at 1440p resolution.
- Spatial reconstruction: fed one to dozens, sometimes over a hundred, input images, Atlas rebuilds real-world scenes as novel-view video frames and as explicit 3D outputs, point clouds and 3D Gaussian splats, outperforming specialized state-of-the-art 3D-reconstruction models.
- Space-time simulation: from ordinary video shot on a handful of cell phones, Atlas can freeze and reframe a scene from impossible angles, and it feeds real-to-simulation pipelines that let robots train and test in generated environments.
- Image generation: Atlas produces images and 360-degree panoramas from text prompts, following complex instructions, rendering legible text and covering a wide range of visual styles.
Under the hood, Atlas is what World Labs calls a multimodal autoregressive diffusion transformer, trained from scratch rather than adapted from an existing video or language model. Every input, text, images, camera poses and 3D depth maps, is anchored to a position in a shared spatial context, and Atlas generates new output conditioned on everything already placed in that context; two unrelated photos, for instance, can be positioned in 3D space and stitched into one continuous, coherent world. That sets Atlas apart from Sora-style video generators, which steer camera movement through imprecise text prompts: Atlas takes camera trajectories and scene geometry as native inputs, giving it frame-by-frame control that text alone cannot provide. In evaluations on camera-controlled generation and reconstruction from sparse views, World Labs says Atlas outperformed both leading video models and the best specialized 3D-reconstruction systems, with its advantage growing on more complex camera paths, and that performance keeps improving as training compute scales up.
A bet on robots, not just visual effects
The launch continues a pattern World Labs set earlier this year. In June, Fei-Fei Li outlined a framework splitting world models into renderers, simulators and planners, arguing the simulator matters most because it carries the geometry, physics and dynamics that both rendering and action depend on. On July 21, World Labs acquired SceniX, a robot-simulation startup; a week later, on July 28, it unveiled a real-to-sim-to-real system built around the idea that robotics' biggest bottleneck is the lack of cheap, controllable training experience at scale. Atlas is the piece that connects those efforts: it can turn ordinary photos or video into a 3D space, then into the sensor views and variable simulated environments a robot needs to train in. Jim Fan, Nvidia's senior robotics research lead and a former student of Li's, described the launch as a major step for real-to-sim work in robotics.
World models generate, reconstruct, and simulate any possible world, so we can render imagined worlds for creators, simulate the real world in high fidelity, and help robots plan their actions.
World Labs says Atlas is rolling out gradually, starting with early access for select partners; other users can request access on the company's website. The company frames Atlas as the foundation for future versions of Marble, its earlier interactive-world product, and for the rest of its lineup. For AI labs, game and VFX studios, and robotics companies anywhere, the practical shift is a cheaper path to usable 3D data: instead of expensive capture rigs or months of specialized reconstruction work, a handful of ordinary photos or a phone video becomes a starting point for a simulated environment, a game level or a robot's training ground, provided the early results hold up once more of the field can put Atlas to the test.
Sources
- Atlas: A World Model for Spatial IntelligenceWorld Labs · September 1, 2026
- 李飞飞发布:全球首个多模态世界模型量子位 (QbitAI) · September 2, 2026
- 李飛飛推出新一代世界模型 Atlas,模擬真實世界和機器人運作TechNews 科技新報 · September 2, 2026



