Search papers, labs, and topics across Lattice.
Nanjing University
4
0
5
3
Achieving nearly 5x speedup on RISC-V processors, RVANNS revolutionizes approximate nearest neighbor search by optimizing both vector representation and graph traversal.
Bole accelerates hybrid-attention LLMs by up to 4.72 times while slashing memory usage by up to 99 times, transforming the landscape of autoregressive decoding.
TIDE-MC achieves unprecedented efficiency in matrix completion, handling billion-scale datasets without running out of memory while delivering up to 11,647x speedup.
Encrypted LLM inference just got a whole lot faster: TIGER accelerates TFHE-based nonlinear layers by up to 17x using GPUs.