Search papers, labs, and topics across Lattice.
5
0
9
2
ModularRSI is proposed, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution that contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies.
This work introduces AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains, and uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics.
Source-style collapse can cause fine-tuned retrievers to miss relevant capabilities, but a simple TF-IDF signal can dramatically improve retrieval success rates.
Cosmos 3 sets a new benchmark for omnimodal models, outperforming existing state-of-the-art in both Text-to-Image and Image-to-Video tasks.
TACO reduces token overhead by 10% while boosting terminal agent performance by up to 4%, revolutionizing how we approach long-horizon reasoning tasks.