Sharif University of TechnologyApr 2, 2026路also Amirkabir University of Technology, Asal Mohammadjafari Mamaqani, Beheshti University, Kariminia Shahid +3
Despite advances in vision-language models, they still fail at rebus puzzles, highlighting a critical gap in cognitive visual reasoning that neither scaling nor in-context learning can fix.