Search papers, labs, and topics across Lattice.
This study investigates the attractiveness ratings of faces by Multimodal Large Language Models (MLLMs) in comparison to human judgments, utilizing a dataset of 2,513 participants and four commercial AI models. The results reveal that while MLLMs consistently assign higher attractiveness ratings and exhibit strong rank-order correlation with human judgments, they fail to match human ratings in absolute terms and show varied predictive cues for attractiveness across models. Notably, only face age aligns as a predictor of attractiveness for both humans and MLLMs, indicating a divergence in the criteria used by AI compared to human evaluators.
MLLMs may rank faces similarly to humans, but they systematically overrate attractiveness, raising questions about their reliability in aesthetic assessments.
Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular amongst users, companies, and aestheticians. This raises the question of whether these AI models can accurately reflect human judgments of attractiveness. In a pre- registered exploratory study, we compared the attractiveness ratings of 2,513 human participants to four widely used commercial AI models: Claude, Gemini, GPT, and Grok. Results showed that MLLMs systematically rate faces more favourably and within a narrower range than humans and, at the time of study, do not reproduce human ratings in absolute terms. However, MLLMs exhibit strong correlations with human attractiveness judgments, accurately tracking the rank-ordering of faces. MLLMs may judge faces by different cues than humans; only face age was a predictor of facial attractiveness in both humans and MLLMs, with inconsistent patterns across models for ethnicity and gender. AI models strongly agree with one another, except for Grok, which also showed the lowest agreement with humans. Our findings suggest that while they may be able to approximate rank-orderings of human attractiveness, current off-the-shelf commercial MLLMs systematically overrate the beauty of human faces.