Search papers, labs, and topics across Lattice.
Affiliation:
5
0
8
A simple physical object can hijack an agent's internal imagination rollouts, trapping its downstream policy in an adversarial trajectory long after the trigger leaves the scene.
CoKL enables LLMs to learn new tasks without sacrificing previously acquired capabilities, striking a balance that traditional methods fail to achieve.
Closing the supervision gap in GUI agents boosts success rates from the low-30% range to over 50% through innovative skill-guided learning.
Mismatched SFT data hurting your LLM's reasoning? DART uses RL to transform it into perfectly aligned training examples, boosting generalization and efficiency.
Overcome the prohibitive cost of ground-truth labels in reinforcement learning by actively acquiring labels for only the most valuable samples, leading to stable training and improved performance even with limited annotation budgets.