Search papers, labs, and topics across Lattice.
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences
5
0
8
14
TokenWall slashes the attack success rate to 12.5% while ensuring a 97.4% pass rate for benign interactions, all with just 0.69 seconds of added latency.
Training LLMs on data detoxified with HSPD slashes toxicity by more than half, outperforming existing methods that only address toxicity during or after training.
Just one carefully crafted poisoned document can cripple an LLM's reasoning abilities in retrieval-augmented generation.
Neural retrievers' preference for LLM-generated text isn't an inherent flaw, but rather a learned bias from artifacts present in training data, offering a path to debiasing without architectural changes.
Prompt highlighting in LLMs gets a serious upgrade: PRISM-$\Delta$ steers models to focus on relevant text spans with better accuracy and fluency, even in long contexts.