Search papers, labs, and topics across Lattice.
6
0
9
4
LLMs can create effective execution infrastructures for certain tasks, but struggle to match human-engineered solutions in critical areas like coding and search.
Self-improvement in LLMs is not automatic; the pathway to effective learning depends critically on the task structure and the nature of the feedback received.
Vague goals can misdirect model evolution, revealing that agents often overfit to narrow self-assessments, hindering broader learning outcomes.
Cultural competence in language models is more about pre-training exposure than multilingual fluency, revealing a critical gap in AI's understanding of cultural nuances.
Fine-tuning a model on rigorously synthesized tasks can outperform larger models by leveraging high-fidelity data, achieving a new benchmark in terminal agent performance.
Converting noisy, human-centric guides into self-evolving agent skills can yield performance improvements of up to 25.3 percentage points across diverse tasks.