Search papers, labs, and topics across Lattice.
Nanjing University
1
0
2
Fine-tuning on the new OmniVideo-100K dataset boosts model performance by over 20% in audio-visual reasoning tasks, revealing the power of structured scripts in enhancing multimodal understanding.