Search papers, labs, and topics across Lattice.
Affiliation:
6
0
14
This paper introduces Magenta, a training-free agentic pipeline that, given only a natural-language problem, produces an answer, expresses it as a Lean 4 statement, and constructs a machine-checked proof, which achieves 100% accuracy across all evaluated olympiad benchmarks.
Shifting from solution-centric to information-centric decision-making, Iris achieves unprecedented performance in autonomous ML engineering tasks.
Pythagoras-Prover achieves state-of-the-art performance in formal proving with dramatically fewer parameters, challenging the notion that bigger models always yield better results.
LLMs can dramatically improve MIMO controller tuning by reasoning about complex interactions, achieving optimal performance with far fewer evaluations than traditional methods.
Idiom comprehension in low-resource languages suffers significantly, with literal meanings proving far more challenging than figurative interpretations, even in context-rich conversations.
Forget hand-crafted reward functions: $\text{RLR}^3$ leverages rubrics and LLMs to provide fine-grained, multi-criteria supervision, outperforming standard RLVR in vision-language tasks.