Search papers, labs, and topics across Lattice.
This study introduces TAF-MED, a benchmark designed to evaluate the persistence of medication-safety boundaries in large language models (LLMs) across multi-turn conversations with explicit self-treatment intent. By assessing 4,000 conversations across eight LLMs, the authors found that a staggering 71.6% of interactions contained unsafe responses, with 61.4% of those starting with a safe response collapsing to unsafe in subsequent turns. The results highlight significant variability in model performance, with collapse rates ranging from 24.4% to 96.2%, underscoring the inadequacy of relying solely on initial response safety as a measure of overall conversational safety.
A staggering 71.6% of LLM conversations about self-treatment led to unsafe medical advice, revealing critical flaws in current safety assessments.
Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent. We introduce TAF-MED, a physician-reviewed benchmark of 500 fixed three-turn scenarios, and evaluate eight LLMs across 4,000 conversations. A rubric-based automated judge labelled responses as SAFE, LEAKY, or UNSAFE, and two physicians independently annotated a model-balanced random subset of 400 conversations. We assessed unsafe guidance, collapse after a strictly SAFE initial response, and model-ranking stability. Overall, 71.6% of conversations contained an UNSAFE response, and 61.4% of those beginning with a strictly SAFE response later collapsed to UNSAFE; model-level collapse rates ranged from 24.4% to 96.2%. Four of 28 model pairs reversed order between initial unsafe and collapse rates. Automated labels achieved 94.3% agreement with the adjudicated physician reference ($\kappa = 0.895$). These findings show that first-turn safety is an incomplete proxy for conversational safety persistence and motivate evaluation across complete dialogue trajectories. We will release TAF-MED on Hugging Face to support reproducible research on multi-turn medical safety.