Search papers, labs, and topics across Lattice.
6
0
7
5
Solution Hacking reveals that up to 44.1% of answers from frontier LLMs may be misleadingly credited as correct due to invalid reasoning shortcuts.
Living-Harness enables agents to learn from past failures dynamically, leading to substantial performance improvements in interactive tasks.
LLMs struggle with financial reasoning under real-world conditions, revealing critical flaws in their ability to handle complex, long-horizon tasks.
Despite achieving comparable overall scores, top-performing medical LLMs exhibit surprising differences in reasoning, evidence use, and longitudinal follow-up when evaluated on a new Chinese medical benchmark, revealing critical gaps in clinically actionable treatment planning.
Failure-driven post-training, combined with a meticulously curated 10M token STEM dataset, unlocks a 4.68% performance boost in LLM reasoning, proving that strategic data synthesis around model weaknesses is a powerful path to improvement.
An open-source ecosystem for agentic learning, complete with a trained agent and novel policy optimization, promises to accelerate research by providing a standardized, scalable platform.