Search papers, labs, and topics across Lattice.
3
0
6
3
PIVOT achieves up to 4x faster indexing for token-level sparse attention without sacrificing accuracy, transforming how we handle query processing in large models.
Exact Aumann-Shapley attributions in GNNs are now achievable with a polynomial architecture, drastically improving fidelity while slashing computational costs.
Only a subset of design interactions in heterogeneous LLM inference are binding constraints, revealing critical insights for optimizing deployment strategies.