Search papers, labs, and topics across Lattice.
Engram-E2VID introduces a novel framework for reference-based event-to-video reconstruction that utilizes generative activation of appearance engrams to recover target RGB frames from a reference frame and event stream. By encoding the reference frame into token-space appearance engrams and transforming the event stream into a target-time motion-structure scaffold, the method effectively associates event-derived structures with relevant appearance information. The results demonstrate significant improvements in PSNR and LPIPS across three benchmarks, indicating enhanced reconstruction fidelity even over longer temporal intervals.
Target frame reconstruction from sparse event data achieves up to 3.29 dB improvement in PSNR, showcasing a breakthrough in video fidelity.
Reference-based event-to-video reconstruction aims to recover target RGB frames from a reference frame and the event stream captured over the reference-to-target interval. Although events provide fine-grained temporal cues, they encode sparse and asynchronous log-intensity changes rather than absolute appearance, making faithful reconstruction intrinsically challenging. The central challenge lies in associating event-derived target-time structures with relevant appearance information from the reference frame, especially under complex motion and long temporal intervals. In this work, we propose Engram-E2VID, a structure-guided framework that reconstructs target frames through the generative activation of appearance engrams. Specifically, the reference frame is encoded into token-space appearance engrams, while the event stream and reference context are transformed into a target-time motion-structure scaffold that captures motion boundaries and event-induced structural changes. Within a one-step diffusion backbone, scaffold-derived structural tokens progressively interact with and activate relevant appearance engrams across layers. This token-space association allows target structures to access reference appearance without relying on direct pixel-wise correspondence, while the diffusion prior complements uncertain or newly revealed regions. Across three benchmarks, Engram-E2VID improves PSNR by up to 3.29 dB and reduces LPIPS by up to 0.08 over the strongest same-input baseline, while degrading more slowly as the reconstruction interval increases.