Search papers, labs, and topics across Lattice.
3
0
5
5
HarnessEval-W transforms world model evaluation from mere scoring to a transparent reasoning process that mirrors human judgment.
Evaluation Agent slashes evaluation time to 10% of traditional methods while providing detailed, user-tailored analyses of visual generative models.
Current video generation models struggle with law-grounded reasoning, with the best achieving only 47% on the new Apple-PI benchmark.