Search papers, labs, and topics across Lattice.
3
0
6
0
Merging experts based on their phase roles can enhance MoE-VLM performance by up to 9.6%, challenging the effectiveness of traditional global aggregation methods.
REFLEX redefines MoE inference by aligning expert computation with the distinct refinement needs of tokens, achieving efficiency gains without compromising quality.
Backdoored LLM agents can stealthily leak your sensitive data via disguised tool calls, and the risk grows with each turn of interaction.