Search papers, labs, and topics across Lattice.
Harvard University, Massachusetts Institute of Technology
2
0
4
Subword tokenization just got a whole lot more efficient: ToaST slashes token counts by 11% and boosts language model performance by up to 7.6% compared to standard methods.