Search papers, labs, and topics across Lattice.
6
0
7
3
Scalar metrics fail to capture the true diversity of AI-generated content, but diversity profiles offer a robust, multi-dimensional evaluation framework that reveals hidden biases.
No memory substrate is a one-size-fits-all solution; the right choice depends on the task, with broad retrieval boosting QA but hindering decision-making.
Leading LLM investment advisors can significantly enhance long-term investor outcomes, revealing a critical gap in traditional evaluation methods.
LLM agents can now autonomously generate complex skills with multi-file dependencies, rivaling human-authored skills, thanks to a co-evolutionary verification process that doesn't need ground truth labels.
Diffusion language models can achieve better reasoning performance by explicitly balancing generation quality and exploration, outperforming methods that prioritize only one.
Even state-of-the-art LLMs struggle to adapt to mid-task changes in long-horizon web navigation, highlighting a critical gap in their ability to handle realistic user interactions.