Search papers, labs, and topics across Lattice.
Affiliation:
4
0
6
1
GIFT enforces user data isolation in LLM serving with less than 11% throughput overhead, revolutionizing privacy protection in shared infrastructures.
Achieving up to 16x faster attention computation with only a 1.76% accuracy loss could revolutionize the deployment of long-context LLMs in real-time applications.
FlexServe achieves over 10X faster secure LLM inference on mobile devices without compromising privacy or performance.
Edge NPUs can outperform flagship GPUs in cost and energy efficiency for on-robot VLA model deployment, but only with hardware-aware optimizations that tackle the models' distinct compute and memory-bound phases.