Search papers, labs, and topics across Lattice.
12
0
12
4
This work proposes EvoSkill-GUI, a training-free framework in which each skill is a structured multi-file package containing retrieval metadata, executable plans, backup localization, failure-recovery rules, accessibility utilities, and failure cases.
Standard translation metrics and 235B-parameter LLM judges actively reward the worst social media translations due to cultural blindness, but progressively decaying masked cultural scaffolding enables an 8B model to rival Gemini-3.1-Pro.
PaperGym achieves a remarkable 73.48 on ResearchQA, outperforming larger models and redefining how AI can generate and evaluate research plans.
TTPO achieves label-supervised performance without any ground-truth labels, outperforming traditional methods on key benchmarks.
LLMs can reason better when they're not forced to answer in English, and a new RL method leverages this quirk to boost performance across reasoning tasks.
LLMs can learn to recognize when they lack sufficient information for reasoning and proactively ask for clarification, leading to more reliable and concise answers.
Uncertainty-driven zoom-in boosts GUI grounding accuracy by up to 13.4% without any retraining, showing that targeted attention to model uncertainty can significantly improve performance.
Offloading memory and computation to a copilot lets a 7B parameter GUI agent outperform larger models on long-horizon tasks, suggesting a path to more efficient and capable GUI automation.
Forget noisy pseudo-labels: SpatialEvo unlocks self-supervised 3D spatial reasoning by generating perfectly accurate training data directly from scene geometry.
Finally, a unified open-source framework lets you train, evaluate, and deploy GUI agents across real devices and chat platforms, closing the gap between research and real-world application.
Even frontier models like Claude Sonnet 4.6 stumble when asked to infer user preferences and proactively assist in mobile tasks, achieving less than 50% success despite excelling at explicit task execution.