Search papers, labs, and topics across Lattice.
2
0
2
Adding language-specific cues in tokenization can significantly improve how multilingual models handle cross-lingual homographs, leading to better translation outcomes.
Subword tokenization just got a whole lot more efficient: ToaST slashes token counts by 11% and boosts language model performance by up to 7.6% compared to standard methods.