Search papers, labs, and topics across Lattice.
TJUNLP Lab, School of Computer Science and Technology, Tianjin University
4
0
7
12
Selective distillation can unlock critical learning signals in reinforcement learning, leading to significant performance gains in complex tasks.
RL fine-tuning on a massive new mobile GUI dataset closes the sim2real gap, outperforming supervised methods and suggesting a path to more robust vision-language agents.
RL's superior generalization isn't about brute force, but about carefully sculpting a few key features while preserving the base model's knowledge, unlike SFT's rapid specialization.
Forget brute-force hinting: KnowRL distills knowledge into atomic units, then uses subset selection to find the *least* amount of guidance needed to supercharge LLM reasoning.