Search papers, labs, and topics across Lattice.
Affiliation:
6
0
7
6
Rubric-based scoring reveals that vision-language models can significantly improve their grounding in visual evidence, enhancing reasoning and instruction adherence.
Evaluation Agent slashes evaluation time to 10% of traditional methods while providing detailed, user-tailored analyses of visual generative models.
Current video generation models struggle with law-grounded reasoning, with the best achieving only 47% on the new Apple-PI benchmark.
Forget dialogue summaries – FileGram builds user profiles directly from atomic file-system actions, unlocking a richer, more privacy-preserving approach to agent personalization.
Turns out, just showing a vision-language model the last few frames of a video stream beats fancy architectures designed for long-term memory.
Today's best AI agents can only achieve 48% accuracy when reasoning about your personal files, revealing a surprising gap between lab performance and real-world usability.