Search papers, labs, and topics across Lattice.
4
0
7
ModularRSI is proposed, a benchmark-disjoint, contrastive, and modular framework for generalizable harness evolution that contrasts successful and failed trajectories for the same task and aggregates evidence across tasks to identify recurring behavioral deficiencies.
Rare slang senses remain a significant challenge in semantic change detection, with Macro-F1 scores hovering around 0.5 across evaluated models.
Reasoning VLMs falter under semantic distractions, often mistaking irrelevant cues for evidence, which can lead to incorrect answers.
TACO reduces token overhead by 10% while boosting terminal agent performance by up to 4%, revolutionizing how we approach long-horizon reasoning tasks.