WorldCrafter Introduces Implicit 3D-Aware Memory for Coherent Video World Exploration
WorldCrafter is a video world model designed to maintain spatial coherence as users navigate generated scenes over time. Rather than relying on explicit depth estimation or geometric warping, it uses a camera-queryable implicit 3D-aware memory: a jointly trained encoder and pose-conditioned readout system that converts past viewpoints into target-view-specific memory tokens. This allows the model to recover previously observed scene content when a camera revisits a location, while recent temporal context handles ongoing motion. The system accepts a single image or text prompt as input and supports minute-scale exploration across both static and dynamic scenes.
On a benchmark covering 145 scenes and 725 camera trajectories, WorldCrafter outperforms evaluated baselines on long-horizon revisit consistency, camera-control accuracy, and overall VBench score. Few-step distillation enables real-time streaming interaction. The model also supports multi-view image inputs for novel view synthesis, and reconstructed point clouds are offered as visual validation of spatial coherence across revisits.