LoGo Uses Spatially Localized Rewards to Fix 3D Inconsistency in Long-Horizon Video Generation
Researchers from Caltech and World Labs have introduced LoGo, a post-training method designed to address a persistent failure mode in camera-controlled video generation: 3D inconsistency. As video models generate longer sequences with complex camera trajectories, objects lose permanence, scene layouts shift on revisit, and artifacts appear. LoGo tackles this by blending spatially localized rewards — which assign fine-grained credit to specific regions of a scene via 3D reprojection and voxel-level error scoring — with a global reward that preserves overall video quality and camera-following fidelity.
The method introduces three components: reward localization (using depth unprojection and a voxel grid to score per-region consistency), local-global blending (to prevent reward hacking that degrades sharpness), and reward interleaving. The team also releases TrajectoryBench, a new benchmark for evaluating long-horizon, complex camera trajectories — a gap they identify in current evaluation frameworks. LoGo demonstrates consistent improvements across three base video models on both DL3DV and TrajectoryBench.