Search papers, labs, and topics across Lattice.
This study conducts a cross-national audit of harmful content exposure on TikTok using sockpuppet accounts representing various age groups in France, Italy, and Sweden, collecting nearly 37,000 videos. By validating four multimodal language models against native-speaker labels, the research identifies that Gemini 2.5 Flash outperforms others in annotation efficiency and accuracy, revealing a significant increase in harmful content detected through keyword searches compared to passive scrolling. The findings highlight Italy as having the highest harm rates across all ages, underscoring the limitations of current safety filters in accurately capturing explicit harms.
Keyword searches on TikTok reveal up to 56% harmful content, significantly outpacing passive scrolling results and challenging existing moderation assumptions.
Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas (13, 16, 19, 40), collecting 36,971 videos from passive For-You-page scrolling and active sessions that scroll, search for harm keywords, and scroll again. To scale annotation, we validate four multimodal LLMs against native-speaker labels on a 300-video reference set. Gemini 2.5 Flash with eight sampled frames plus text performs best (aggregate kappa = 0.42), at half the per-call cost of native-video upload, and we apply it to a 10% sample for approximately \$50 in total API spend across both modalities. Keyword search returns 35-56% harmful content, a 1.5-7.5x increase over the scrolling baseline in ten of twelve country-age combinations; the spike is temporary and flattens the age differences observed in France and Sweden. Under passive scrolling, Italy has the highest harm rate at every age, with Italian age-19 reaching 48.6%. Overall, MLLM-based auditing offers a scalable approach for cross-national youth-safety audits, while provider safety filters (1.1% refusal rate) under-count the most explicit harms.