Search papers, labs, and topics across Lattice.
5
0
6
13
HarnessEval-W transforms world model evaluation from mere scoring to a transparent reasoning process that mirrors human judgment.
Evaluation Agent slashes evaluation time to 10% of traditional methods while providing detailed, user-tailored analyses of visual generative models.
Current video generation models struggle with law-grounded reasoning, with the best achieving only 47% on the new Apple-PI benchmark.
The fragmented field of world modeling can now be unified under a "levels x laws" taxonomy, revealing critical gaps in autonomous model revision and decision-centric evaluation.
Current video benchmarks fail to capture the nuances of animation, but AnimationBench fills the gap by rigorously evaluating character consistency, motion, and style.