Search papers, labs, and topics across Lattice.
Thanks:
3
0
4
5
UnionSparse achieves up to 3.46x faster low-bit sparse LLM inference on edge GPUs by optimizing metadata handling, challenging the status quo in model efficiency.
Achieving over 2.2x inference speedup in Vision Transformers while maintaining accuracy reveals the untapped potential of hardware-software co-design in optimizing model performance.
Forget slow NTTs: Hermes' hybrid dataflow architecture delivers up to 13.6x faster throughput for homomorphic encryption than GPUs.