Search papers, labs, and topics across Lattice.
This study introduces a regression-based method for predicting speaker origin in Arabic dialects by modeling dialectal variation as a continuous geographic space instead of discrete categories. Utilizing a hierarchical neural architecture that integrates advanced encoder representations with phonotactic descriptors, the model achieves a median localization error of 481.2 km, while also demonstrating the dialect continuum through a permutation Mantel test. The findings highlight the effectiveness of continuous geographic modeling for dialect geolocation, revealing both its capabilities and areas for further improvement.
Continuous geographic modeling reveals that predicting Arabic dialect origins can achieve impressive accuracy while exposing significant room for improvement in zero-shot scenarios.
We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rather than discrete categories. Speaker origin is predicted as continuous latitude-longitude coordinates using a hierarchical neural architecture that fuses frame-level XLS-R-300M and Whisper-large-v3 encoder representations with phonotactic descriptors through a Transformer encoder and a learnable attention-pooled query. A spherical geodesic loss directly optimizes great-circle distance on Earth's surface, avoiding distortions inherent to planar coordinate regression. Under a leakage-free 5-fold GroupKFold protocol grouped by source recording, our model attains a pooled median localization error of 481.2 km. Auxiliary country and city heads reach 64.5% and 45.2% accuracy, respectively. A permutation Mantel test on the learned latent space provides quantitative support for the Arabic dialect continuum hypothesis. To probe true generalization, we further introduce a city-masking protocol in which two cities per fold are removed from training but retained in validation. Under this zero-shot regime, the mean error rises to 1173.3 km, a 1.32x degradation relative to seen cities. Our findings establish continuous geographic modeling as a principled framework for Arabic dialect geolocation and quantify both its strengths and the substantial headroom that remains.