Search papers, labs, and topics across Lattice.
Affiliation:
4
0
7
This work proposes Vid-PRE (Video Prompt Reasoner and Enhancer), a model-agnostic prompt rewriter that offloads the cognitive burden of reasoning to a dedicated VLM, and introduces VWG-Bench, a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks.
Contrastive flow matching can upper-bound forward-KL divergence on fixed preference pairs, resolving the training instability of unbounded DPO surrogates without requiring online rollouts.
Even top-tier vision-language models consistently invert a subject's left and right, exposing a blind spot in camera- versus subject-centric spatial reasoning that rubric-guided GRPO over anatomical priors can finally resolve.
A novel framework that combines audio-grounded verification with a rich dataset leads to significant advancements in long-paragraph audio captioning.