Search papers, labs, and topics across Lattice.
9
0
12
6
GPT-5.5 not only tops the leaderboard in policy evolution but also reveals critical insights into how agents can optimize performance through strategic feedback utilization.
FlashMorph reveals that optimizing layer selection in hybrid attention models can drastically improve efficiency while maintaining performance, outperforming existing heuristic methods.
PracRepair fixes 162 out of 171 bugs in a challenging benchmark, showcasing a leap in automated program repair capabilities through human-inspired debugging techniques.
MaxProof's innovative test-time scaling enables an AI to outperform human champions in mathematical proof competitions.
FlowTracer reveals that optimizing token-level rewards based on attention-induced information flow can dramatically enhance reasoning performance in LLMs.
Code-based 3D reconstruction achieves superior edit fidelity and locality, outperforming traditional point-cloud methods in preserving unedited regions.
Even the best LLMs struggle with Olympiad-level combinatorics, achieving only 65.4% on a benchmark designed to expose their reasoning limitations.
Projector fine-tuning, commonly used for aligning MLLMs, unexpectedly introduces backdoor vulnerabilities with activation mechanisms distinct from those in text-only LLMs.
Test-time training can finally scale for large reasoning models: TEMPO unlocks sustained performance gains by interleaving policy refinement with periodic critic recalibration, boosting accuracy by over 18% on challenging benchmarks.