Search papers, labs, and topics across Lattice.
Affiliation:, Affiliation:
2
0
4
Unified multimodal models may excel in generation and understanding, but they often falter when reasoning about their own outputs, revealing hidden weaknesses in their capabilities.
Ditch imperfect human annotations: this dual-reward RL approach trains image captioning models to be both more complete and more factually correct.