Search papers, labs, and topics across Lattice.
The paper introduces SCOPE, a large-scale dataset of 241,280 counterfactual prompt pairs for evaluating group-sensitive behavior in LLMs across nine bias dimensions and 1,536 demographic groups. The dataset addresses limitations of existing fairness benchmarks by offering increased linguistic diversity, topical coverage, and support for analyzing the impact of communicative intent (Question, Recommendation, Direction, Clarification). By providing a controlled and intent-aware resource, SCOPE enables systematic investigation of fairness, robustness, and counterfactual consistency in LLMs.
LLMs exhibit surprisingly inconsistent behavior when prompted with semantically identical requests that differ only in demographic group, and SCOPE reveals the extent of this bias across diverse topics and communicative intents.
Large Language Models (LLMs) now serve as the foundation for a wide range of applications, from conversational assistants to decision support tools, making the issue of fairness in their results increasingly important. Previous studies have shown that LLM outputs can shift when prompts reference different demographic groups, even when intent and semantic content remain constant. However, existing resources for probing such disparities rely primarily on small, template-based counterfactual examples or fixed sentence pairs. These benchmarks offer limited linguistic diversity, narrow topical coverage, and little support for analyzing how communicative intent affects model behavior. To address these limitations, we introduce SCOPE (Stereotype-COnditioned Prompts for Evaluation), a large-scale dataset of counterfactual prompt pairs designed to enable systematic investigation of group-sensitive behavior in LLMs. SCOPE contains 241,280 prompts organized into 120,640 counterfactual pairs, each grounded in one of 1,438 topics and spanning nine bias dimensions and 1,536 demographic groups. All prompts are generated under four distinct communicative intents: Question, Recommendation, Direction, and Clarification, ensuring broad coverage of common interaction styles. This resource provides a controlled, semantically aligned, and intent-aware basis for evaluating fairness, robustness, and counterfactual consistency.