Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
Hyper-parallel decoding on a distilled 4B LLM matches frontier model information extraction accuracy while operating at an 8% compute footprint.
LLMs can achieve better zero-shot product ranking with 57% less token usage by reasoning over structured attribute graphs instead of raw text.