Search papers, labs, and topics across Lattice.
This paper addresses the challenges of understanding idiomatic expressions in low-resource Southeast Asian languages by introducing Varnika, a multimodal idiom corpus that includes 3,533 idioms with diverse tonal representations. To enhance idiomatic comprehension, the authors propose a Hybrid Mixture-of-Experts (HybridMoE) framework that effectively integrates insights from multiple idiomatic experts while addressing expert sparsity through controlled hybridization and Idiomatic Property Signals. Empirical results show that HybridMoE yields a 5-6% performance improvement in advanced vision-language models, underscoring its effectiveness in capturing the nuances of figurative language and cultural context.
HybridMoE boosts idiomatic understanding in multilingual models by integrating insights from multiple experts, leading to significant performance gains in figurative language representation.
In the contemporary epoch of multilingual education, learning idioms provides a fascinating gateway towards creativity, cultural values, historical context, and diverse perspectives inherent to various linguistic traditions. This paper showcases the navigation of retaining figurative and cultural semantics in low-resource Southeast Asian languages such as Hindi, Bengali, and Thai, where culturally rich idioms pose significant obstacles for computational modeling and cross-linguistic transfer due to their deep metaphorical complexity. To tackle such complexity, we present Varnika, a reconstructed multimodal idiom corpus comprising 3,533 multilingual idioms, enriched with seven idiomatic tones aligned with both textual and visual representations. Additionally, to infer informative idiomatic understanding, we introduce a Hybrid Mixture-of-Experts (HybridMoE) framework that embeds multiple idiomatic expert opinions while mitigating expert sparsity by integrating outputs from both selected and unselected experts through controlled hybridization, further augmented with Idiomatic Property Signals via masked multimodal embeddings. To analyze the performance across multiple dimensions, we propose the IDIO-TONE and Idiomatic Validation Score, a three-stage evaluation pipeline measuring (i) literal translation fidelity, (ii) visual-semantic alignment, and (iii) idiomatic meaning retention. Empirical evaluations highlight that HybridMoE achieves 5--6\% performance gains across advanced vision language models, demonstrating improved representation of figurative language and culturally embedded meaning in multilingual multimodal settings