Search papers, labs, and topics across Lattice.
Affiliation:
4
0
7
2
Even the strongest Android GUI agents are universally vulnerable to runtime anomalies, revealing critical flaws in their robustness.
Skill augmentation can elevate CLI agent performance beyond GUI counterparts, revealing that the true challenge lies in skill coverage rather than model capability.
Standard retriever evaluations hide critical weaknesses in agentic search systems, but a new benchmark and training method exposes and addresses these flaws.
Frontier models are wasted on routine GUI tasks: a step-level cascade that adaptively invokes stronger models only when lightweight monitors detect progress stalls or semantic drift slashes compute costs without sacrificing performance.