Search papers, labs, and topics across Lattice.
8
0
14
3
Black-box RL can boost agent performance by nearly 15 points on complex tasks, revealing a new frontier for scalable optimization.
OPD transfers reasoning skills rather than answers, revealing a complex interplay between teacher-student origins that can either enhance or hinder model capabilities.
A two-loop configuration in LoopCoder-v2 boosts code generation performance by over 50% compared to a non-looped baseline, while more loops actually hinder results.
FORT-Searcher achieves superior performance by synthesizing training tasks that actively resist shortcut exploitation, transforming how we train deep search agents.
Building agents that can reliably automate complex, multi-step workflows over local files and tools just got a whole lot easier.
Industrial code generation gets a reasoning boost: InCoder-32B-Thinking leverages error-driven feedback and a code world model to achieve top-tier performance on complex hardware-aware tasks.
Code LLMs can achieve SOTA performance in agentic tasks by explicitly modeling the dynamic evolution of software logic across different training stages.
A new 32B code LLM trained specifically for industrial tasks crushes existing models on specialized domains like chip design and GPU kernel optimization, while remaining competitive on general coding benchmarks.