Search papers, labs, and topics across Lattice.
This paper introduces JoyAI-Echo-1.5, a unified audio-visual generation system designed for long-form narratives and interactive environments, featuring two specialized variants. The long-video variant utilizes composable cross-shot memory and speaker cues to maintain character consistency across various media, while the world-model variant enables flexible navigation through calibrated 6-DoF camera trajectories. Experimental results show that JoyAI-Echo-1.5 outperforms existing baselines in cross-shot consistency, visual quality, and speech fidelity, achieving top scores on relevant benchmarks.
JoyAI-Echo-1.5 not only enhances long-video generation but also sets a new benchmark for interactive storytelling with its advanced memory and geometric control mechanisms.
Video generation is progressing beyond isolated clips toward long-form narratives and interactive worlds, requiring models to preserve identities, follow user controls, and remain stable over extended rollouts. We present JoyAI-Echo-1.5, a unified audio-visual generation system with two purpose-built variants. The long-video variant introduces composable cross-shot memory that aggregates visual evidence across multiple prior shots and speaker cues derived from speech-filtered full-shot audio, enabling persistent character appearance and voice identity across flexible combinations of text, image, and memory conditioning. The world-model variant converts heterogeneous navigation inputs into calibrated metric 6-DoF camera trajectories and injects them through a geometry-aware conditioning pathway, enabling controller-agnostic interaction across flexible viewpoints. To support efficient long-horizon generation, we transform a bidirectional audio-visual backbone into a causal few-step generator using progressive teacher forcing and short- and long-horizon Self-Gradient Forcing on self-generated rollouts. Experiments demonstrate strong performance in both settings. JoyAI-Echo-1.5 achieves improvements over existing long-video baselines in cross-shot consistency, visual quality, text alignment, and speech fidelity. Its world-model variant ranks first on WBench, with an average score of 81.7, and achieves leading visual quality and long-horizon persistence on SANA-WM-Bench. Together, these results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds. Project page: https://echo-team-joy-future-academy-jd.github.io/Echo-1.5-Page/.