Search papers, labs, and topics across Lattice.
This paper introduces Automated Research Design Tracking and Assessment (ARDTrA), a novel approach for automatically detecting and evaluating the quality of causal research designs in social science literature. By leveraging a multi-turn RAG-based conversational pipeline, the authors analyze a dataset of expert-annotated papers across six counterfactual research design families, revealing that passage length significantly impacts performance. Notably, the findings indicate a disconnect between human and machine difficulty in assessing research designs, highlighting distinct challenges in automated evaluation methods.
Automated assessment of causal research designs reveals a surprising disconnect between human and machine difficulty, challenging assumptions about expert consensus.
Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Tracking and Assessment (ARDTrA), a task that involves detecting the research design used in a paper and assessing the quality of its application. We create an expert-annotated dataset of papers covering six families of counterfactual research designs and evaluate the task using a multi-turn RAG-based conversational pipeline. Across four retrieval strategies, four LLMs and six embedding models, we find that passage length is the main driver of performance, explaining 52-66% of the variance. A per-research-design analysis also shows that human and machine difficulty do not align: the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task difficulty.