Search papers, labs, and topics across Lattice.
This research explores the efficacy of large language models (LLMs) in detecting scam phone calls in Turkish, a low-resource language, by introducing a novel multi-modal dataset comprising 100 aligned audio-transcript pairs. The study evaluates seven LLMs across three input conditions鈥攔aw audio, automatic speech-to-text transcripts, and human-refined transcripts鈥攔evealing that transcript-based inputs consistently outperform direct audio processing. These findings underscore the importance of developing culturally and linguistically inclusive AI safety measures, particularly in addressing real-world threats in underrepresented languages.
Transcript-based inputs for scam detection in Turkish outperform audio processing, highlighting a critical gap in AI safety for low-resource languages.
Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other high-resource languages. In low-resource settings such as Turkish, detection is especially difficult, as annotated data is scarce and technological defenses remain limited. This research investigates how large language models (LLMs) can support scam detection in Turkish by introducing the first public multi-modal dataset of 100 aligned audio-transcript pairs of scam and benign conversations. We evaluate seven LLMs spanning three model families: Gemini 2.5 (Flash, Flash-Lite, Pro), GPT-4o, and Qwen (Max, Plus, Turbo), under three input conditions: raw audio, automatic speech-to-text transcripts, and transcripts refined by a native speaker. Our results suggest that transcript-based inputs consistently outperform direct audio processing, while human-corrected and uncorrected transcripts perform comparably. By centering a low-resource language and real world threat, this work highlights the urgent need for culturally and linguistically inclusive AI safety research and more robust multi-modal systems for fraud prevention.