Search papers, labs, and topics across Lattice.
2
1
4
0
Bengali speakers face a staggering 67:1 training-token deficit compared to English, highlighting a critical inequity in AI language support.
Tokenization disparities can inflate costs and reduce effective context length for underserved languages, with Bengali and Yoruba facing tokenization premiums of up to 4.5 times that of English.