Search papers, labs, and topics across Lattice.
This paper introduces TetherMem, a novel query-aware spatiotemporal memory routing technique designed to enhance long-horizon autoregressive video generation by decoupling subject and scene queries. By employing region- and age-conditioned priors, TetherMem allows subject queries to maintain identity while enabling scene queries to adapt more dynamically to changes in background and viewpoint, addressing the issue of memory-anchored scene under-progression. In evaluations involving 2,400 blinded pairwise judgments, TetherMem outperformed eight baseline models, achieving the highest scores for overall quality and scene progression, thus demonstrating its effectiveness in generating coherent and dynamic video content.
TetherMem enables video generation models to dynamically adapt scenes while keeping subjects consistent, achieving a significant leap in overall video quality and scene progression.
Streaming autoregressive video models generate long videos chunk by chunk, using historical memory to maintain consistency. Existing methods typically expose subject and scene queries to history through similar policies. This stabilizes the subject, but can also lock backgrounds, viewpoints, and scene structure to previously generated states even when local motion continues. We call this failure memory-anchored scene under-progression; consistency and motion metrics alone can miss it. We introduce TetherMem, a training-free, query-aware spatiotemporal memory router for frozen video generators. TetherMem separates subject and scene queries and modulates historical access with region- and age-conditioned priors: subject queries retain identity-bearing history, while scene queries reduce reliance on subject history and stale backgrounds. Across 2,400 blinded pairwise judgments from 10 annotators, TetherMem achieves the highest estimated expected preference among eight streaming long-video baselines for overall quality (0.780) and scene progression (0.769). On complete 30-second videos, it sustains changes in background, viewpoint, and scene state while preserving subject recognizability and temporal continuity.