Search papers, labs, and topics across Lattice.
Central South University
4
0
6
Rejected samples, often seen as failures, can actually drive an 85.3% boost in OCR accuracy for text-rich image generation.
State-of-the-art image-to-video models are alarmingly vulnerable to visual prompt attacks, achieving up to 100% success in triggering harmful outputs.
Advanced RS MLLMs struggle with negation, but a novel learning method can dramatically enhance their understanding using minimal unlabeled data.
Forget text-centric pipelines: FlowInOne achieves SOTA multimodal generation by unifying text, layouts, and instructions into a single visual flow, outperforming both open-source and commercial systems.