Back to a16z Podcast

Fei Fei Li: The Race to Build World Models For AI

a16z Podcast

Full Title

Fei Fei Li: The Race to Build World Models For AI

Summary

This episode introduces Atlas, a new AI model that predicts new views of a scene, aiming to build world models capable of understanding and reasoning about physical space.

Atlas unifies 3D reconstruction and generation, moving beyond next-token or next-frame prediction towards a more comprehensive spatial intelligence.

Key Points

  • Atlas's core innovation is "new view prediction," a primitive for AI models that allows it to generate novel perspectives of a scene from a limited set of input views, unlike language models predicting the next token or video models predicting the next frame.
  • The model significantly reduces the data and computational requirements for 3D reconstruction, enabling complex scenes to be generated from as few as three cameras, eliminating the need for studios or green screens.
  • Atlas uniquely combines 3D reconstruction and generative capabilities within a single architecture, a feat previously separated into distinct fields within computer vision, by natively incorporating camera poses and 3D data as modalities.
  • Previous work, like the Marvel world model, focused on Gaussian Splats for 3D output, but Atlas's new view prediction primitive allows for more flexibility and direct generation of RGB frames and 3D data without this bottleneck.
  • The development of Atlas involved scaling up models and training them for longer, leading to significant improvements, and the team had strong conviction in the underlying hypotheses of scaling and next viewpoint prediction.
  • The ability to create spatially contextualized and grounded pixels is a major step towards true spatial intelligence, going beyond simple pixel generation.
  • Atlas's capabilities have immediate applications in creative industries, allowing for easier scene generation and manipulation, and are also crucial for robotics by streamlining the real-to-sim data pipeline.
  • Future directions for Atlas include incorporating richer dynamics, interaction, and bridging the gap between simulation and real-world robotics through more advanced planning.
  • The concept of "AI completeness" is discussed, with next token prediction being considered AI complete for language, and generative new view prediction for spatial intelligence holding similar potential.

Conclusion

Atlas represents a fundamental advancement in AI's ability to understand and generate spatial information by unifying reconstruction and generation through new view prediction.

The model's impact extends across various industries, from creative media to robotics, by significantly lowering the barrier to entry for complex 3D content creation and simulation.

The development of Atlas highlights the potential of "next viewpoint prediction" as a core primitive for spatial intelligence, potentially mirroring the impact of "next token prediction" on language models.

Discussion Topics

  • How will generative new view prediction fundamentally change how we interact with and create digital content?
  • What are the most exciting near-term applications of Atlas beyond creative industries, particularly in fields like robotics and urban planning?
  • Considering the concept of AI completeness, how might generative new view prediction become a foundational primitive for future AI systems aiming for true spatial understanding?

Key Terms

New View Prediction
The AI model's ability to generate novel perspectives or views of a scene based on existing input views.
Spatial Intelligence
The capability of an AI system to understand, reason about, and interact with physical space and its properties.
3D Reconstruction
The process of creating a three-dimensional model of an object or scene from two-dimensional images or other sensor data.
Generative AI
Artificial intelligence systems capable of creating new content, such as images, text, or in this case, novel views of a scene.
Modality
A type of data input or output, such as text, image, video, or 3D data.
Camera Pose
The position and orientation of a camera in three-dimensional space.
Gaussian Splatting
A rendering technique for representing 3D scenes using a collection of semi-transparent, ellipsoidal primitives (Gaussians).
Real-to-Sim
The process of translating real-world data and environments into a simulated environment for training AI models.
Sim-to-Real
The process of transferring a trained AI model from a simulated environment to operate effectively in the real world.
Neural Simulator
A simulation environment built using neural networks, capable of modeling physical interactions and responses based on learned data.
AI Completeness
The concept that a specific AI task or primitive, if solved in its most general form, could be used to solve any general intelligence task.
Turing Completeness
In computer science, the ability of a system to simulate any Turing machine, meaning it can perform any computation that a computer can.

Timeline

00:00:06

Atlas represents a significant step in generating spatially contextualized and grounded pixels.

00:00:15

Atlas's core primitive is new view prediction, a novel approach for AI models.

00:02:37

The functionality of Atlas is explained using the "bullet time" effect from The Matrix as a comparison.

00:03:35

Atlas's fundamental primitive is defined as new view prediction, differentiating it from next token or next frame prediction.

00:04:41

Atlas achieves spatially grounded meaning by associating each input image with a 3D camera pose, allowing for precise reconstruction.

00:05:54

Atlas is a new architecture that jointly handles generation and reconstruction, unlike previous specialized models.

00:06:35

Atlas is natively multimodal, working with text, images, videos, and camera poses as inputs, and 3D as a modality.

00:07:41

The unification of pixel generation and pixel reconstruction in Atlas is highlighted as a significant achievement.

00:08:35

Atlas is positioned as a significant step towards the broader goal of spatial intelligence, which involves generating, reasoning within, and interacting with space.

00:11:17

The previous Marvel world model generated 3D worlds as Gaussian Splats, while Atlas focuses on new view prediction as its core primitive.

00:12:31

The development of Atlas involved overcoming challenges in representation and scaling, requiring extensive experimentation.

00:13:19

The team had conviction that their formulation for Atlas would scale.

00:14:20

Dense reconstruction, historically an arduous process requiring many views, is contrasted with Atlas's ability to work with fewer inputs.

00:17:10

The Stanford demo illustrates Atlas's ability to reconstruct a large area from a limited number of ground-level images.

00:18:09

The need for generative capacity to fill in gaps in reconstruction is emphasized, as perfect coverage from input views is impossible.

00:18:47

The concept of a long context window for reconstruction and generation is linked to the value seen in LLMs.

00:19:33

Atlas overcomes the limitations of previous models like Marvel by handling a much larger number of input images for reconstruction.

00:20:23

Atlas functions as a rendering engine that can produce a dense fly-through from a reduced set of input views.

00:21:23

The 3D consistency of Atlas's generations is discussed, attributed to data and the scaling hypothesis.

00:21:33

The team had conviction that Atlas would work, though the speed and quality exceeded initial expectations.

00:22:55

The team believes they are at the beginning of Atlas's potential, with current limitations being training compute.

00:23:54

A pivotal moment occurred when the model successfully generated a fly-through under a table, solidifying the decision to build Atlas.

00:24:44

The use cases for Atlas are explored, extending from creative applications to industrial and design purposes.

00:29:24

Atlas is crucial for robotics by streamlining the real-to-sim data generation process, which is a major bottleneck.

00:30:58

The primary challenge in robotics is data collection, and Atlas significantly aids in creating realistic and randomized simulation environments.

00:32:12

Atlas's multimodal nature is key for future robotics applications, including integrating dynamic data and bridging action planning with its outputs.

00:32:55

Training robotics policies differs from other AI applications because they are agents acting in the real world, requiring exposure to a wide range of scenarios, which simulation helps provide.

00:34:43

The learned simulator concept for robotics is explored, with the idea that the simulator itself could become the "player."

00:35:24

The need for more dynamics in world models, as noted by an expert, is discussed.

00:36:00

Atlas already fundamentally supports dynamics, with elements like water waves and moving cars appearing in demonstrations.

00:36:54

Exposing the model to dynamics during training, even for static output goals, is found to be beneficial.

00:37:57

The question of whether future developments will lead to "4D video" is posed.

00:38:15

The horizontal primitive nature of Atlas is discussed, suggesting its potential to enable an entire industry.

00:38:26

The excitement around the multimodal aspect and adding control conditioning to models like Atlas is shared.

00:40:37

The concept of "AI completeness" is introduced, suggesting that generative new view prediction may be AI complete for spatial intelligence, analogous to next token prediction for LLMs.

00:42:54

New viewpoint prediction is framed as an evolutionary necessity for animals, making it a fundamental aspect of intelligence.

00:43:24

Congratulations are offered on the successful launch of Atlas.

Episode Details

Podcast
a16z Podcast
Episode
Fei Fei Li: The Race to Build World Models For AI
Published
September 4, 2026