Search papers, labs, and topics across Lattice.
This paper critiques the use of BLEU-4 as a benchmark for sign language translation (SLT), arguing that it fails to accurately reflect sign language proficiency due to its reliance on spoken-language metrics. By evaluating six SLT models on the Phoenix-2014T and CSL-Daily datasets, the authors demonstrate that improvements in BLEU-4 do not correlate with enhanced sign understanding. They propose a novel evaluation protocol based on language-learning assessments, which is significantly more robust and aligns better with human evaluations, revealing that gloss-free systems perform similarly while gloss-supervised systems show a substantial advantage that BLEU-4 overlooks.
BLEU-4 may mislead SLT progress, masking the true performance differences between gloss-free and gloss-supervised systems.
BLEU-4 is the standard metric for evaluating sign language translation (SLT), but spoken-language metrics may not adequately reflect sign language proficiency. The multimodal, low-resource context of SLT allows models to exploit spurious correlations and spoken-language priors, rather than learning stronger sign representations. In this paper, we evaluate the relationship between spatio-temporal understanding and BLEU-4 across six SLT models on Phoenix-2014T and CSL-Daily, showing that gains in BLEU-4 are not on their own evidence of better sign language understanding. This work introduces an alternative inspired by language-learning assessment, using an open-weight-LLM QA protocol that measures salient content preservation. It aligns more closely with human rankings and is six to seven times more paraphrase-invariant than BLEU-4. Applied to SLT, this protocol targets content transfer, is more robust to train-test overlap, and gives a different picture of the field: the five gloss-free systems are largely within noise of one another on Phoenix-2014T, while the gloss-supervised system stands 9.3 points higher, a gap invisible to BLEU-4.