Search papers, labs, and topics across Lattice.
3
0
7
2
Vision-language models struggle to leverage visual evidence in medical VQA, with only one model surpassing human performance on a subset of questions.
LLMs can appear competent in medical contexts, but a rigorous evaluation reveals a staggering drop in performance that questions their true clinical abilities.
LLMs are inflating IR benchmark scores, but it's hard to tell how much is real progress versus just memorization.