Search papers, labs, and topics across Lattice.
3
0
7
0
MoEGen achieves instance-specific adaptations without the storage burden of full LoRA experts, revolutionizing how we think about parameter-efficient fine-tuning.
Rule-based defenses may protect LLMs while preserving task performance, but many popular methods compromise usability and efficiency.
Tri-serve redefines energy efficiency in multimodal inference by addressing hidden power inefficiencies, achieving a 22% boost without latency trade-offs.