Search papers, labs, and topics across Lattice.
Affiliation:
5
0
10
3
Intermediate context compression can cut GPU energy usage by over 50% while maintaining quality, challenging static compression strategies in edge RAG applications.
SparseDitto achieves up to 146.61x speedup for sparse matrix operations by dynamically customizing GPU kernels based on input patterns.
Achieving over 134% accuracy gains with just 1,000 trainable parameters reveals a game-changing approach to enhancing spatial reasoning in vision language models.
Flash-WAM achieves real-time inference for world-action models by reducing latency from 8.1 seconds to 348 milliseconds without sacrificing performance.
Moxin 7B and its variants (VLM, VLA, Chinese) offer a new suite of fully transparent, open-source multimodal models, pushing beyond simple weight sharing to enable deeper customization and collaborative research.