Search papers, labs, and topics across Lattice.
Affiliation:
4
0
7
3
This work studies adaptive routing of prompts to large language model experts to maximize response quality in an online setting with limited feedback and proposes algorithms that strategically select and observe rewards to minimize regret.
By leveraging semantic redundancy, POOL can slash confidence-estimation costs by up to 76% without sacrificing accuracy.
Retrieval-augmented reasoning can boost LRM accuracy by up to 60% during test-time scaling, transforming how models handle complex problem-solving.
Internal model signals alone can guide speculative decoding to achieve both higher accuracy and lower latency in multi-step reasoning, outperforming reward-guided approaches.