Search papers, labs, and topics across Lattice.
This study critiques the timing of evaluations for large language models (LLMs) in medical consultations, highlighting the "preformulation gap" where initial vague patient concerns are overlooked. By assessing three API models across various scenarios, the researchers found significant differences in the models' ability to provide self-care advice and structured handoff summaries based on whether entry-to-care instructions were given. The findings suggest that evaluating LLMs solely on final outputs may miss critical early-stage interactions that influence patient care outcomes.
LLMs miss critical early-stage patient interactions, providing self-care advice in only 75% of baseline scenarios compared to none under structured instructions.
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.