Search papers, labs, and topics across Lattice.
This study investigates the cross-cultural validity of taste-sound correspondences in AI-generated music by conducting a three-country experiment with participants from Argentina, Italy, and Japan. While preferences for fine-tuned music excerpts based on taste prompts were confirmed in Argentina and Italy, they were not replicated in Japan, highlighting significant cultural differences in how taste is attributed to sound. The findings reveal that apparent cross-cultural divergence is largely due to response-style biases, although a structural component of taste attribution remains consistent across cultures, suggesting that evaluations of generative music systems must account for these biases.
Taste-sound correspondences in AI-generated music vary significantly across cultures, with response-style biases obscuring deeper perceptual differences.
Sonic seasoning research has shown that listeners attribute systematic gustatory and emotional meaning to sound, and text-to-music generative artificial intelligence has recently been used to render gustatory prompts as musical stimuli. Whether the taste-sound correspondences acquired by such models hold beyond the cultural context in which they were validated remains untested. We extended a single-country study to a three-country online experiment conducted in Argentina, Italy, and Japan (N = 361). Participants first indicated their preference between base and fine-tuned MusicGen excerpts generated from four taste prompts (sweet, sour, bitter, salty), and then rated fine-tuned excerpts on twelve taste, emotion, and thermal descriptors. Preference for the fine-tuned model was confirmed in Argentina and Italy but not in Japan, and the salty prompt yielded the weakest correspondence in all three cohorts. Ratings differed substantially between countries, yet the main effect of country was no longer detectable once ratings had been standardized within participant, whereas the interactions characterizing the mapping of prompts onto descriptors remained essentially unchanged. Much of the apparent cross-cultural divergence is therefore attributable to differences in scale use; a structural component nevertheless persists. In addition an exploratory factor analysis indicated that the twelve descriptors were organized along different latent dimensions in each cohort. These results indicate that cross-cultural variation in AI-mediated sonic seasoning operates at two levels: the overall level at which taste is attributed to a given stimulus, and the relational structure of those attributions. Evaluations of generative music systems across populations should accordingly distinguish response-style bias from genuine perceptual reorganization.