Introduction of Atlas, a Next-Generation World Model for Spatial Intelligence
Atlas is a newly launched multimodal world model designed to generate, reconstruct, and simulate various environments using text, images, video, and 3D data. It aims to enhance spatial intelligence for applications in robotics, gaming, and creative industries by providing precise control over scene generation and camera movements.
Atlas is a next-generation world model developed by World Labs, capable of operating on multiple data types including text, images, video, and 3D.
It functions as a multimodal autoregressive diffusion transformer, combining inputs into a shared spatial context to generate coherent outputs.
Atlas can create new views from one or more reference images, allowing users to specify camera positions and angles for scene generation.
The model can reconstruct real-world spaces from sparse input images without requiring specialized capture equipment, outperforming existing 3D reconstruction models.
Atlas is designed to scale, with performance improving as more training compute is applied, and it can generate both 2D images and 3D outputs such as point clouds.
The model supports various applications, including video generation with controlled camera movements, simulation for robotics, and creative scene generation.
Atlas's architecture integrates advancements from both language and video models, enabling it to handle a wide range of tasks in world modeling.