Search papers, labs, and topics across Lattice.
2
0
4
The central finding is that aggregate WER hides code switching behavior, and the best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric.
Audio language models can grasp the broad gist of stuttered child speech, but their reasoning completely collapses and leaks multi-speaker context as disfluency rates rise.