Search papers, labs, and topics across Lattice.
Huazhong University of Science and Technology
4
0
6
6
Over 95% reduction in backdoor attack success rates with only 0.1% neuron intervention reveals a breakthrough in LLM security.
A staggering 28% of mid-sized models' answers are mere spurious guesses, exposing critical flaws in their reasoning capabilities.
Performance of large reasoning models drops significantly as logical complexity rises, revealing critical gaps in current evaluation benchmarks.
Safety in MoE LLMs isn't about routing harmful requests to "refusal experts"鈥攊t's surprisingly localized within specific experts, and you can break it without significantly changing the model's overall routing behavior.