Search papers, labs, and topics across Lattice.
This study audits six state-of-the-art vision-language models (VLMs) to evaluate their consistency in representing the affective qualities of untextured 3D objects using Kansei adjective pairs. The results reveal that while there is some convergence in how these models rank objects along affective dimensions, the agreement is partial and varies significantly across different object categories. Notably, the models' convergence is influenced more by the alignment of representational variation with the evaluated semantic direction than by overall shape variation, highlighting the limitations of VLMs in capturing human-like affective judgments.
VLMs show only partial agreement on the affective qualities of shapes, revealing significant inconsistencies that could impact generative design interfaces.
Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as"more elegant"or"more minimalist,"typically encoded by a vision-language model (VLM). A practical question is whether state-of-the-art VLMs represent objects consistently in terms of the same concept. We audit 6 VLMs by ranking untextured 3D objects along Kansei adjective pairs, where Kansei describes affective impressions of product form, with each axis defined as the difference between the text representations of its two poles. Geometric pairs serve as positive controls, and pairs of unrelated adjectives establish an empirical null. Across 10 categories of ShapeNet database, affective axes converge above the null (mean pairwise rank correlation 0.36 vs. 0.14) but below the geometric ceiling (0.44). The agreement between models is partial and highly uneven: on the three axes shared by all categories, mean convergence ranges from 0.21 for bookshelves to 0.51 for jars. Convergence depends primarily on whether a category's representational variation aligns with the semantic direction being evaluated, rather than simply on how much the objects vary in shape overall. Cross-model convergence does not imply agreement with human judgments. Based on our findings, we implement a UI prototype that shows how the audit can inform which Kansei descriptors to expose as controls for a given object class and which to withhold.