Search papers, labs, and topics across Lattice.
6
0
8
5
The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities.
BenchShield is presented, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation that grounds detection in a finite lifecycle model of an evaluation's reward-relevant events and achieves 96% accuracy in detecting reward hacking from infrastructure-side evidence.
Identity drift in generative agents reveals that anti-self-deception is a dominant modification behavior, challenging assumptions about agent fidelity under pressure.
Simulating 8.3 billion diverse personas reveals nuanced user interactions that traditional evaluations miss, transforming how we assess AI systems.
Transforming multimodal resources into executable agent skills boosts performance by nearly 12 percentage points, showcasing the power of diverse learning materials.
LLM agents automating productivity tasks achieve only moderate success (39-64%) while exhibiting surprisingly high rates of unsafe actions (7-33%) in realistic, multi-service workflows.