Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
Domain-specialized RL experts can be unified without negative transfer by distilling their feedback directly onto student-generated trajectories, resolving the multi-task optimization bottleneck that limits single video foundation models.
VLMs don't fail to *recognize* harmful intent when jailbroken; instead, visual inputs *shift* their internal representations into a distinct "jailbreak state," opening a new avenue for defense.