Search papers, labs, and topics across Lattice.
6
1
6
3
A model-free benchmark reveals that vision-language models often misread rather than misreason, exposing a performance gap that a supervised vision model can surpass.
Culturally grounded multimodal evaluation reveals significant gaps in current AI systems' understanding of Arabic contexts, with implications for both accuracy and representation.
A single design choice鈥攚hether to include parity labels鈥攃an determine if a model predicts physically impossible outcomes, with staggering accuracy implications.
Natural-language descriptions can outperform traditional image statistics in predicting pulmonary nodule malignancy, enabling a calibrated LLM to triage cases effectively.
Agreement among perturbed inputs can mislead accuracy assessments, as fine-tuning on consensus can paradoxically reduce performance.
Label-free reliability in vision-language models has a computable blind spot that can be systematically characterized and detected.