Search papers, labs, and topics across Lattice.
15
0
14
6
LLMs can create effective execution infrastructures for certain tasks, but struggle to match human-engineered solutions in critical areas like coding and search.
Looping the middle layers of MoE Transformers can save up to 18% in training FLOPs while boosting performance beyond validation loss predictions.
Transforming pre-training data with reasoning annotations boosts model performance by over 2 percentage points without altering the training objective.
Vague goals can misdirect model evolution, revealing that agents often overfit to narrow self-assessments, hindering broader learning outcomes.
Self-improvement in LLMs is not automatic; the pathway to effective learning depends critically on the task structure and the nature of the feedback received.
Transforming scientific papers into multi-turn generation trajectories not only doubles the training data but also boosts academic writing benchmarks while maintaining reasoning skills.
Current omni-modal models can achieve only moderate success in interactive video assistance, revealing critical gaps in their understanding of user interactions and visual cues.
Despite advances in AI, even top models struggle with real-world tasks, achieving only 30% success on a benchmark grounded in market-validated workflows.
Separating temporal dynamics from order information could revolutionize how we model user behavior in recommendation systems.
Coding agents may appear compliant, but they actually underperform when faced with rules that challenge their default behaviors, revealing a critical flaw in current evaluation metrics.
Learning from real-world environments follows a precise log-sigmoid scaling law, with agent performance and learning speed improving dramatically over time.
Cultural competence in language models is more about pre-training exposure than multilingual fluency, revealing a critical gap in AI's understanding of cultural nuances.
Learning algorithms might excel in memorization but can falter in broader generalization, with RL outperforming SFT in transferring knowledge across contexts.
Current AI agents struggle with long-horizon professional tasks, achieving only 30% success in complex GUI workflows, revealing critical gaps in their capabilities.
LLMs can now perform entity alignment with greater interpretability and efficiency thanks to a new agent-based approach that structures the reasoning process.