Search papers, labs, and topics across Lattice.
4
0
6
7
Everyday interactions with LLMs can foster informal learning, but only if users engage deeply and contextually with the AI.
IRT models can mislead AI evaluations, especially when benchmarks deviate from traditional testing conditions, risking inaccurate performance assessments.
Current video generation benchmarks overlook crucial aspects of physical plausibility and temporal coherence, highlighting the need for holistic evaluation metrics like PhyScore.
Turns out, the NLP metrics we use to evaluate interview systems don't actually predict whether a response is useful for qualitative research.