Search papers, labs, and topics across Lattice.
5
0
7
Disagreement among verifiers can be a powerful signal for identifying errors in multimodal reasoning, leading to a 5.95% performance boost without any training.
CARE-X achieves a remarkable 94.0% accuracy in visual question answering, outperforming existing models by 6 percentage points while also enhancing report quality and spatial localization.
Targeted feedback can slash calculation errors in small language models from 56.9% to 23.5%, revolutionizing their physics reasoning abilities.
Code LLMs can recognize incorrect instructions but still follow them, leading to irrecoverable semantic errors that defy traditional evaluation metrics.
Hierarchical visual concepts learned through cascaded sparse autoencoders could revolutionize how we interpret and manipulate MLLM outputs.