Search papers, labs, and topics across Lattice.
University of Science and Technology of China
4
0
6
5
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
Simulating 8.3 billion diverse personas reveals nuanced user interactions that traditional evaluations miss, transforming how we assess AI systems.
Even the best vision-language models struggle with reliable evaluation of computer-using agents, but OS-Shepherd models offer a low-cost solution that matches their performance.
A unified evaluation framework that simplifies the assessment of LLM-based agents could drastically enhance reproducibility and accelerate research breakthroughs.