Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
0
DeepRepoQA achieves substantial performance gains in repository question answering by leveraging Monte-Carlo Tree Search for deep, multi-hop reasoning over code dependencies.
SRPO enables LLMs to self-reflect and transform sparse feedback into dense learning signals, achieving state-of-the-art performance with drastically reduced training costs.
Temporal confidence allows for adaptive computation in parallel reasoning, leading to a 32% reduction in latency without sacrificing accuracy.