Search papers, labs, and topics across Lattice.
1
0
3
0
Non-thinking inference in hybrid-thinking MLLMs suffers from widespread response-pattern failures, revealing a critical misalignment that can be mitigated with targeted reinforcement learning strategies.