Search papers, labs, and topics across Lattice.
This paper critically examines the limitations of counterfactual explanations (CEs) in the context of justification and recourse within AI systems, arguing that their naive application can obscure important design and governance choices made during the machine learning pipeline. Through four empirical experiments, the authors demonstrate that upstream decisions鈥攕uch as measurement models and validation metrics鈥攕ignificantly influence the generated counterfactuals, often more than the specifics of the CE generation method itself. The findings highlight the necessity of incorporating these upstream choices into the justification and recourse processes to avoid misleading interpretations of model behavior.
Counterfactual explanations can mislead by masking critical upstream decisions that shape model outputs, challenging their normative legitimacy in AI justification and recourse.
Counterfactual explanations (CEs) are widely used in explainable artificial intelligence (AI) to show how a model's outputs would change if the input features were manipulated. This technique is used for a range of tasks such as debugging models, explaining predictions, justifying decisions, and providing algorithmic recourse. In this paper, we explore the normative legitimacy of employing counterfactuals in real-life model deployment settings. We discuss the different stakes involved in these different purposes for which CEs are commonly employed, and find stricter requirements for justification and recourse. In particular, we find that naive application of CEs for justification and recourse can lead to ignoring contestable choices made throughout the machine learning (ML) pipeline, thus obfuscating that decisions and counterfactuals for those decisions are also artifacts of an organization's materialized design and governance choices. We demonstrate this with four empirical experiments involving interventions at stages of the ML pipeline ``upstream" of the explanation itself, and show that these affect the generated counterfactuals. We find that an organization's choices on measurement models for feature and labels, business requirements, model validation, and the metric of model success have as much or more impact on the generated counterfactuals as the specifics of the generating method. Our findings underline the need to account for such choices upon providing justification and recourse, providing a stark reminder of the relational nature of these tasks. As putative justifications or recourse recommendations, CEs do not provide adequate answers to some important "why"-questions because they preclude consideration of whether the decision-maker ought to have acted differently.