Search papers, labs, and topics across Lattice.
This paper introduces MINARD, a novel pipeline for generating narrated videos that explain complex scientific figures by grounding the narration to specific regions of the figure and its corresponding paper. The approach addresses the gap in current video generation systems, which lack the capability to provide step-by-step, paper-grounded explanations that enhance understanding of intricate scientific content. Experimental results on the newly released FigTalk benchmark demonstrate that MINARD produces humanlike, accurate narrations that significantly outperform existing methods in both automatic and human evaluations.
MINARD generates humanlike, paper-faithful narrations for scientific figures, bridging a critical gap in understanding complex visual data.
Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capability missing from current video generation systems and benchmarks. To address this, we introduce paper-grounded figure-to-video generation: generating narrated, region-grounded walkthrough videos from a figure and its paper. We propose MINARD (Multimodal Interpretation of Narrated Architecture via Region Decomposition), a pipeline that generates paper-grounded narrations and sequentially grounds them to figure regions. We also release FigTalk, a benchmark with new sequential and component-level grounding metrics derived. On FigTalk, MINARD generates humanlike, paper-faithful narrations and outperforms narration-conditioned figure spatial grounding compared to existing approaches in both automatic and human evaluation