Search papers, labs, and topics across Lattice.
University of British Columbia
5
1
7
7
ModularRSI is proposed, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution that contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies.
Self-improvement in LLMs is not automatic; the pathway to effective learning depends critically on the task structure and the nature of the feedback received.
Wrapping standard CLI coding agents in a state-persisting orchestration layer beats unconstrained autonomous agents at research completeness without altering the underlying model.
Current video generation models struggle with visual reasoning, with the best achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
Mid-training with function-aware fill-in-the-middle boosts coding agent performance while preventing capability erosion in non-agentic tasks.