Search papers, labs, and topics across Lattice.
This study introduces WildSEEK, a novel dataset comprising 3,000 annotated information-seeking queries derived from real user interactions, aimed at evaluating the performance of language models in real-world contexts. The research reveals that over one-third of these queries are classified as high-risk, particularly highlighting the challenges LLMs face with analytical queries, where they exhibit higher failure rates in areas such as sycophantic behavior and handling of vulnerable populations. By establishing a comprehensive evaluation framework, this work lays the groundwork for assessing the reliability and safety of LLMs as they increasingly mediate information access.
Over a third of real-world information-seeking queries are high-risk, exposing critical failures in LLM responses, especially for analytical tasks.
Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of "in the wild" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our findings show that LLM responses fail more often in four criteria: sycophantic behavior, overreliance, a default US-centric perspective, and poor handling of vulnerable populations -- with failure rates being mostly higher for analytical queries. By providing methods to monitor the reliability, safety, and fairness of LLM behavior, our dataset and evaluation framework offer an empirical foundation for the broader question of how these systems should behave as they take on a growing role in information access.