Search papers, labs, and topics across Lattice.
Affiliation:
1
0
3
3
Non-thinking inference in hybrid-thinking MLLMs suffers from widespread response-pattern failures, revealing a critical misalignment that can be mitigated with targeted reinforcement learning strategies.