Search papers, labs, and topics across Lattice.
Thanks:
1
0
2
UnionSparse achieves up to 3.46x faster low-bit sparse LLM inference on edge GPUs by optimizing metadata handling, challenging the status quo in model efficiency.