Search papers, labs, and topics across Lattice.
7
0
11
26
LLM-generated research ideas are systematically narrower and more focused than those of human researchers, revealing a significant gap in creative breadth.
RLMF not only boosts LLMs' ability to accurately express uncertainty but also enhances their self-assessment capabilities, fundamentally reshaping their trustworthiness.
Standard retriever evaluations hide critical weaknesses in agentic search systems, but a new benchmark and training method exposes and addresses these flaws.
LLMs are rapidly transforming peer review, but critical gaps remain in ensuring quality, fairness, and ethical considerations across the entire workflow.
Frontier models are wasted on routine GUI tasks: a step-level cascade that adaptively invokes stronger models only when lightweight monitors detect progress stalls or semantic drift slashes compute costs without sacrificing performance.
Stop generating superficial reviews: RbtAct leverages rebuttals to train LLMs to provide actionable feedback, leading to concrete revisions and improved author uptake.
Even GPT-5 struggles to reliably reproduce novel research findings, highlighting a significant gap between capability and reliability for AI agents tackling end-to-end research tasks.