Search papers, labs, and topics across Lattice.
5
0
7
0
Generation-guided training can significantly enhance multimodal understanding in MLLMs without any inference overhead.
CED reveals that VLMs can be trained to prioritize evidence-based reasoning over language shortcuts, leading to more reliable visual understanding.
SwanTale achieves unprecedented expressiveness in multi-speaker audio generation, outperforming existing models in both instruct and zero-shot tasks.
Salience Bias in LLMs reveals that models often ignore commonsense reasoning in favor of misleading explicit cues, with lightweight prompting showing promise in addressing this issue.
Optimal data ratios discovered through DecoupleMix yield competitive VLM performance with 80B fewer training tokens than traditional methods.