Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
SRPO enables LLMs to self-reflect and transform sparse feedback into dense learning signals, achieving state-of-the-art performance with drastically reduced training costs.
Achieve LLaMA-level reasoning accuracy with 44% lower latency and 73% lower API costs by strategically offloading work from large to small models only when needed.
LLM inference gets a 2x speed boost without training, thanks to a clever technique that merges retrieval with logit-based speculation.