Search papers, labs, and topics across Lattice.
This paper introduces a multimodal approach to identifying and characterizing sexism in memes and short-form videos using a late-fusion pipeline that integrates gradient-boosted regression models with various feature types. The study reveals that while LLM-derived semantic cues significantly enhance the identification of sexism in memes, the performance of video analysis is heavily influenced by feature selection and cross-modal noise. Ultimately, the findings emphasize the importance of tailored feature engineering for static content and the necessity for improved temporal modeling in dynamic video contexts.
Targeted semantic feature engineering can dramatically enhance sexism detection in memes, but video analysis reveals a surprising sensitivity to feature dimensionality and noise.
We present the AILS-NTUA submission to the EXIST 2026 Lab at CLEF, addressing multimodal sexism identification and characterization in memes (Task 2) and short-form videos (Task 3). Our system follows a feature-engineered late-fusion pipeline built around gradient-boosted regression models and hierarchical post-processing. For memes, we combine visual, textual, demographic, biometric, and LLM-derived semantic indicators designed to capture high-level cues such as stereotyping, objectification, irony, and misogyny. For videos, we investigate the effect of feature selection, frame-based visual representations, OCR-based textual features, acoustic descriptors, and sensor-derived metadata. Development results show that focused LLM-derived semantic cues improve meme sexism identification, while video performance is highly sensitive to feature dimensionality and cross-modal noise. For videos, development results favor compact feature selection, but official test results show that this conclusion does not fully transfer to unseen data, where the unfiltered representation generalizes better. Overall, our findings highlight the usefulness of targeted semantic feature engineering for static memes and the need for more robust temporal modeling in noisy short-form video settings.