Search papers, labs, and topics across Lattice.
This paper introduces Failure-informed Image Self-Augmentation (FISA), a novel framework that enhances multimodal large language models (MLLMs) by generating augmented images from the model's own failure cases. By focusing on visually challenging yet answer-preserving modifications, FISA improves the model's performance on visual question answering tasks, demonstrating significant gains in both in-distribution and out-of-distribution scenarios. The approach also shows superior data efficiency compared to traditional image augmentation methods, validating its effectiveness through rigorous filtering strategies.
Augmenting images from a model's own failures can lead to substantial performance boosts in multimodal tasks, outperforming conventional augmentation techniques.
Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimodal data that are costly to annotate. Self-augmentation offers a promising alternative by enabling models to expand their own training data without external supervision. However, existing MLLM self-augmentation methods are largely text-centric, while image augmentation remains underexplored and typically relies on generic or handcrafted transformations that are weakly aligned with the model's actual incapability. We propose Failure-informed Image Self-Augmentation (\textbf{FISA}), a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases. Our method generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion. Experiments on visual question answering benchmarks show that the proposed method consistently improves performance across both in-distribution and out-of-distribution settings. Further experiments validate the compatibility of FISA with existing textual self-augmentation approaches, the superior data efficiency of the synthesized samples over generic image augmentation baselines, and the practical effectiveness of the proposed filtering strategy.