Search papers, labs, and topics across Lattice.
This paper introduces Visual Attribution Distillation (VAD), a novel counterfactual target-reconstruction algorithm designed to enhance multimodal on-policy distillation by isolating visual evidence in teacher corrections. By evaluating the impact of visual evidence on student-generated trajectories, VAD effectively distinguishes between corrections supported by visual signals and those influenced by linguistic priors. The results demonstrate that VAD significantly outperforms traditional distillation methods across multiple fine-grained visual benchmarks, particularly in scenarios where visual evidence contradicts incorrect predictions.
VAD reveals that isolating visual evidence can dramatically improve target reconstruction in multimodal learning, leading to more accurate student outputs.
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstruction algorithm that estimates the visually attributable part of a teacher correction. At each student-generated prefix, VAD evaluates the same fixed teacher with the relevant evidence present and removed. The corresponding change in centered log-probabilities defines ut, a signed proxy for the visual evidence direction that estimates how revealing the evidence supports or refutes candidate tokens. VAD projects the original correction onto this proxy to obtain an intervention-aligned component and a proxy-unexplained residual, then reconstructs a student-anchored target from the former. During training, this reconstructed target supplies the primary supervision signal, while the privileged teacher contributes a weak regularizer. Across six fine-grained visual benchmarks at 4B and 9B scales, VAD outperforms direct privileged-view distillation and visual-advantage weighting. Token- level and controlled-target analyses show that the proxy-aligned component is enriched in task-relevant visual corrections and yields stronger target shifts, especially when evidence refutes a mistaken answer. These results support counterfactual target reconstruction as an effective alternative to source-mixed supervision.