Search papers, labs, and topics across Lattice.
8
0
9
8
Despite advances in AI, even top models struggle with real-world tasks, achieving only 30% success on a benchmark grounded in market-validated workflows.
Coding agents may appear compliant, but they actually underperform when faced with rules that challenge their default behaviors, revealing a critical flaw in current evaluation metrics.
Learning from real-world environments follows a precise log-sigmoid scaling law, with agent performance and learning speed improving dramatically over time.
Cultural competence in language models is more about pre-training exposure than multilingual fluency, revealing a critical gap in AI's understanding of cultural nuances.
Learning algorithms might excel in memorization but can falter in broader generalization, with RL outperforming SFT in transferring knowledge across contexts.
Interactive dialogue can unlock creative potential that static assessments overlook, leading to richer evaluations of creativity in AI contexts.
Current AI agents struggle with long-horizon professional tasks, achieving only 30% success in complex GUI workflows, revealing critical gaps in their capabilities.
LLMs can be made better software engineers by pre-training them to reconstruct the messy, iterative development process that led to the final, clean code in repositories.