Search papers, labs, and topics across Lattice.
3
0
5
3
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
A unified evaluation framework that simplifies the assessment of LLM-based agents could drastically enhance reproducibility and accelerate research breakthroughs.
ToolMaze reveals that LLMs suffer a staggering 37% drop in recovery performance when faced with implicit semantic failures, highlighting a critical vulnerability in current models.