Search papers, labs, and topics across Lattice.
3
0
4
Auditing Chinese web content reveals pervasive pollution that shifts over time, challenging the integrity of LLM training data.
Released tokenizer vocabularies can yield precise estimates of hidden corpus compositions, revealing insights that were previously obscured.
Existing unlearning methods can leak sensitive knowledge through multi-hop reasoning paths, exposing a critical vulnerability in LLMs.