Search papers, labs, and topics across Lattice.
4
0
5
3
Training data influence shifts dramatically over the course of language model pretraining, with literature data dominating early and STEM data taking over later.
LLMs mirror human biases against non-native Japanese speakers, but with critical underestimations that could exacerbate inequities in high-stakes evaluations.
The first-ever SCPI classifier for Japanese text reveals significant challenges in detecting sensitive information, crucial for compliance with privacy regulations.
Optimizing multilingual training? Shapley values reveal the hidden cross-lingual transfer effects that current scaling laws miss, leading to better language mixture ratios.