Search papers, labs, and topics across Lattice.
7
24
8
23
Inductive biases can significantly enhance performance in AI tasks with long feedback loops, challenging the dominance of purely data-driven approaches.
Risk-averse decision-making can backfire, leading to generic outputs, while Bayesian methods enhance LLM performance in high-stakes tasks like tutoring and peer review.
Isolated assessments may mask biases, but comparative evaluations can unleash hidden discrimination in LLMs, especially as model sizes grow.
Human evaluative claims can significantly enhance the quality of AI-generated peer reviews, striking a crucial balance between automation and accountability.
A staggering 73% of evaluated computer-use agents leak sensitive information across contexts, underscoring a critical gap in privacy safeguards.
Detecting AI-generated code is harder than you think: even state-of-the-art detectors fail to reliably identify machine-written code, especially when faced with distribution shifts or adversarial attacks.
LLMs that excel at math don't necessarily make good math tutors, revealing a surprising trade-off between subject matter expertise and pedagogical skill.