Search papers, labs, and topics across Lattice.
3
0
6
6
VLMs struggle with strategic decision-making in soccer, showing a preference for safer plays over optimal actions, which could reshape our understanding of their reasoning capabilities.
Even the most advanced LLMs struggle with consistent rubric verification, revealing substantial noise in scoring outputs across complex agentic scenarios.
Current reward models are surprisingly bad at judging story quality, achieving only 66% accuracy in selecting human-preferred narratives – a gap closed by a new, purpose-built reward model.