Search papers, labs, and topics across Lattice.
5
0
8
6
Atria Dawn Preview is introduced, a foundation agentic language model designed for scientific research and engineering workflows, with the goal of expanding the frontier of agent productivity in the real world.
SWE-Bench Pro Verified offers a more trustworthy benchmark for assessing software engineering agents, with safeguards that eliminate major leakage channels without disrupting normal agent functionality, and a task refinement that minimally corrects inconsistencies within flawed instances.
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
A unified evaluation framework that simplifies the assessment of LLM-based agents could drastically enhance reproducibility and accelerate research breakthroughs.
Automating LLM fine-tuning is now possible: a multi-agent system, TREX, matches or exceeds human performance on a diverse set of real-world tasks.