Search papers, labs, and topics across Lattice.
This paper introduces an iterative proxy correction framework for robust incomplete multimodal sentiment analysis (MSA), addressing the limitations of existing one-shot proxy methods that can lead to unreliable sentiment predictions. By constructing a language-oriented proxy from non-language modalities and refining it through gated residual correction, the model effectively balances proxy compensation with trustworthy linguistic evidence. Extensive experiments on datasets such as MOSI, MOSEI, and SIMS reveal that this approach significantly outperforms competitive baselines, ensuring reliable sentiment analysis even with incomplete inputs.
Iterative proxy correction boosts sentiment analysis accuracy by refining initial proxies and adapting to incomplete multimodal inputs.
Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues. However, real-world multimodal inputs are often incomplete or corrupted, which can weaken cross-modal complementarity and introduce misleading information into downstream fusion. Existing proxy-based methods for incomplete MSA commonly rely on one-shot proxy construction to compensate for degraded language information, but the generated proxy may be coarse or unreliable at initialization. Prematurely injecting such a proxy into multimodal reasoning can propagate initial errors and compromise sentiment prediction. To address this limitation, we propose an iterative proxy correction framework for robust incomplete MSA. Our method constructs a language-oriented proxy from non-language modalities and progressively refines it under multimodal context through gated residual correction. The corrected proxy is then adaptively fused with the observed language representation according to an estimated language reliability score, allowing the model to balance proxy-based compensation and trustworthy linguistic evidence. In addition, we introduce a stage-wise latent correction objective that uses the complete language representation as a training-time semantic anchor to stabilize the proxy refinement trajectory. Extensive experiments on MOSI, MOSEI, and SIMS under diverse missing-modality settings demonstrate that the proposed framework consistently outperforms competitive baselines and achieves robust sentiment prediction under incomplete inputs.