Search papers, labs, and topics across Lattice.
This paper critiques the common use of normalized edit distance for evaluating lexicon quality in unsupervised word discovery, highlighting its bias towards larger clusters and its failure to account for the distribution of true classes. To address these issues, the authors introduce two novel metrics: one that adjusts for cluster size in assessing within-cluster consistency and another that evaluates the distribution of true words across clusters. Experiments reveal that these metrics provide a more accurate correlation with ground-truth distributions and demonstrate robustness against biases that typically affect lexicon evaluations.
Current lexicon evaluation methods are biased towards large clusters, but new metrics reveal a more accurate picture of lexicon quality.
Building a lexicon from discovered word-like units is a central goal in zero-resource speech processing. But do our evaluations provide a trustworthy indication of lexicon quality? A common metric, normalized edit distance, averages the phoneme edit distances between discovered units in each cluster. We show that this metric has an inherent bias toward the quality of large clusters, inhibiting fair evaluation. Moreover, it ignores how well true classes are distributed across clusters. Based on established theory in clustering literature, we propose two metrics that address these shortcomings: a modified metric that weighs cluster size when assessing within-cluster consistency, and an inverse metric that assesses how true words are spread across clusters. Through experiments on synthetic and real-world lexicons, we demonstrate that combined, these metrics are: (1) more closely correlated with how similar a lexicon is to the ground-truth distribution, and (2) more robust to biases that skew lexicon evaluations.