Search papers, labs, and topics across Lattice.
4
0
6
Achieving real-time, long-horizon interactive world rollouts on a single desktop GPU could revolutionize how we develop and deploy AI-driven simulations and games.
Current interactive world models fall short, with none passing the rigorous tests of WorldRoamBench designed to assess long-horizon stability across action, vision, physics, and memory.
LLM judges in multi-stakeholder settings suffer from "weighting noise" that gets *worse* as you add more stakeholders, but fixing weights upfront can stabilize the process.
Current reward models struggle to distinguish good vs. bad agent behavior in complex tool-using scenarios, especially over long horizons, revealing a critical gap in alignment research.