Search papers, labs, and topics across Lattice.
5
0
7
5
Concept recoverability in AI grading systems varies significantly by architecture, revealing hidden biases that could undermine assessment fairness.
Query-adaptive detection of active modalities boosts retrieval accuracy by 11.3% over fixed fusion methods in real-world video archives.
LLMs are surprisingly good at pinpointing what's *wrong* with student writing, even outperforming human graders in identifying relative weaknesses.
LLMs beat rule-based systems at understanding nuanced grammar in language learners, but good old-fashioned rules still win on pure syntax.
LLMs aren't equally reliable as NLG evaluators, but a Bradley-Terry extension called BT-sigma can learn judge reliability from pairwise comparisons alone, improving ranking accuracy without human supervision.