Publication: Memory Consistency in Video World Models: Evaluating and Optimizing for Long-Horizon Coherence
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Recent advances in diffusion-based video generation have enabled high quality, action-conditioned video synthesis, raising the possibility of using these models as world models for simulation and decision making. However, despite strong short-term visual fidelity, these models often fail to maintain consistency over long horizons. In particular, when revisiting previously observed viewpoints, generated scenes exhibit drift, hallucination, and structural inconsistencies.
In this work, we identify memory consistency as a fundamental failure mode of autoregressive video generation. To study this phenomenon, we introduce geometry-aware evaluation metrics based on DROID-SLAM reprojection error and spatial distance, which explicitly measure cross-temporal consistency. We show that these metrics capture failure modes that are not reflected by standard perceptual metrics such as PSNR, SSIM, and LPIPS, and that the two are only weakly correlated.
We further investigate methods for improving memory consistency through metric-guided optimization. At inference time, we apply best-of-n sampling to select samples with higher consistency scores. At training time, we adapt Group Relative Policy Optimization (GRPO) to directly optimize the model toward improved long-horizon coherence. While best-of-n provides modest gains, GRPO yields consistent improvements in memory consistency and perceptual quality.
Our results reveal a fundamental gap between optimizing for perceptual fidelity and achieving consistent world modeling behavior. This highlights memory consistency as a distinct and critical dimension of video generation quality, and underscores the need for architectures with explicit mechanisms for persistent spatial memory.