Search papers, labs, and topics across Lattice.
5
0
7
11
A unified evaluation framework that simplifies the assessment of LLM-based agents could drastically enhance reproducibility and accelerate research breakthroughs.
CUAs can achieve a 73.7% success rate on complex macOS tasks, but the secret to their performance lies in skill libraries, not just framework design.
Skills tailored to specific LLM architectures can lead to a performance boost of nearly 26 points, challenging the notion of model-agnostic skill libraries.
OpenMobile proves that high-performing mobile agents can be trained on entirely synthetic, open-source data, closing the gap with closed-source models and enabling broader research.
Decomposing GUI agent trajectories into verifiable milestones and auditing the evidence chain yields a 10% boost in RL training performance, outperforming single-judge reward systems.