Search papers, labs, and topics across Lattice.
2
1
4
3
Agreement among perturbed inputs can mislead accuracy assessments, as fine-tuning on consensus can paradoxically reduce performance.
Label-free reliability in vision-language models has a computable blind spot that can be systematically characterized and detected.