Search papers, labs, and topics across Lattice.
This paper introduces MAG, a framework that enhances few-shot in-context learning (ICL) in multi-modal large language models (MLLMs) by efficiently selecting demonstrations from unlabeled data. By formulating demonstration selection as a semi-supervised propagation problem on a multi-modal graph, MAG employs a two-stage strategy to identify high-impact unlabeled samples and optimize demonstration selection. Experimental results across eight multi-modal benchmarks reveal that MAG significantly outperforms existing methods in label-scarce scenarios, achieving notable improvements with minimal pseudo-labeling resources.
MAG leverages abundant unlabeled multi-modal data to boost few-shot learning performance, achieving significant gains in label-scarce environments.
Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the quality and coverage of the selected demonstrations. While unlabeled multi-modal data is abundant, it remains elusive how to exploit them for ICL. We propose MAG (MAnifold-Guided semi-supervised in-context demonstra- tion selection), an efficient framework that leverages unlabeled data to improve multi-modal ICL. MAG formulates demonstration selection as a semi-supervised propagation problem on a multi-modal graph and adopts a two-stage strategy: (i) relevance score propagation identifies a compact set of high-impact unlabeled samples for pseudo-labeling, reducing MLLM inference cost; (ii) multi-modal relevance is used to select the final demonstrations. We show that textual represen- tations are more effective for relevance propagation, while both visual and textual modalities are crucial for high-quality demonstration selection. Experiments on eight multi-modal benchmarks demonstrate that MAG consistently outperforms strong baselines in label-scarce regimes, achieving significant gains with a limited pseudo-labeling budget.