Atlas: A World Model for Spatial Intelligence(worldlabs.ai)
261 points by johnsutor 8 days ago | 59 comments
tl;dr: World Labs unveiled Atlas, a multimodal autoregressive diffusion transformer that natively handles text, images, video, and 3D depth maps within a shared spatial context. It performs camera-controlled video generation (up to 1 minute at 1440p), sparse-view 3D reconstruction (outputting point clouds or Gaussian splats), space-time reframing, and robotics Real-to-Sim workflows, reportedly outperforming specialized models on camera-following and 3D reconstruction benchmarks. Atlas is entering early access with select partners and will power future versions of World Labs' Marble product.
HN Discussion:
  • Excitement about novel applications like game prototyping and 3D reconstruction from sparse images
  • ~Article underplays interesting uses like extracting semantic latent knowledge for robotics
  • Skepticism about consistency and limitations like hallucinations and frozen-time motion
  • Questions about practical capabilities like exact dimensions, characters, and splat output costs
  • Confusion over the vague and overused term 'world model'