Search papers, labs, and topics across Lattice.
This paper investigates the effectiveness of implicit multimodal in-context learning (M-ICL) by introducing the Selection--Realization Hypothesis, which posits that demonstrations induce a compact family of internal changes that queries can select from. Through controlled experiments, the authors reveal that the efficacy of static task vectors is contingent on the degree of shared demonstration-induced changes across queries, with more complex interventions becoming necessary when query-specific structures are present. The findings not only validate the proposed hypothesis but also offer a framework for optimizing intervention complexity in practical applications, particularly in visual question answering (VQA) tasks.
Static task vectors can effectively capture demonstration-induced changes, but their success hinges on the shared structure across queries鈥攃omplexity is only needed when queries diverge significantly.
Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned transformations and attention routing. Despite their common goal, these methods differ substantially in how the intervention depends on the query and where it modifies the model, leaving unclear which additional complexity is necessary for a given task. We propose the Selection--Realization Hypothesis. It views demonstrations as inducing a compact family of internal changes from which the query selects, while the model's computation constrains how the selected change can be implemented. We evaluate this account using controlled multimodal tasks in which query dependence varies without changing the underlying task primitives or prompt format. By contrasting correct demonstrations with matched counterfactuals, we measure the structure of explicit M-ICL and test whether it predicts intervention behavior. We find that the success of a static task vector is closely tied to how much of the demonstration-induced change is shared across queries. Additional intervention complexity becomes useful when explicit M-ICL contains query-specific or distributed structure that a local additive shift cannot recover. These relationships extend to natural VQA benchmarks and support cost-aware method selection without access to test performance. Our results provide a unified empirical theory of when demonstrations can be compressed into a task vector and when a more expressive intervention is warranted.