Search papers, labs, and topics across Lattice.
This paper critiques the practice of algorithmic gender prediction by distinguishing between its legitimacy and validity, arguing that while such predictions can be harmful, gender imputation may still yield valid measurements for studying gender disparities. The authors draw on transfeminist literature to highlight the different impacts of sexism on women versus transgender and nonbinary individuals, emphasizing that gender imputation can be beneficial for the former but harmful for the latter. Through three case studies, they illustrate the complexities of deploying gender imputation in a way that maximizes anti-discrimination benefits while minimizing harm.
Gender imputation can yield valid measurements for studying disparities, but its legitimacy hinges on the context and the populations affected.
Machine learning ethics researchers and critical HCI scholars have argued that algorithmically predicting gender is wrong. At the same time, other researchers rely on predicted gender labels to study gender disparities and develop algorithmic fairness techniques. How do we reconcile these two seemingly contradictory intuitions? We differentiate two ways gender prediction may be wrong: being illegitimate, thereby contributing to harm; and being invalid, thereby producing unusable measurements. Our analysis translates arguments against gender prediction into these terms of legitimacy and validity and shows how gender imputation applied for fairness purposes can be illegitimate yet still yield valid disparity measurements. We clarify this bind by drawing upon transfeminist literature to distinguish sexism that targets women and femininity from sexism that targets transgender and nonbinary people. While gender imputation can produce valid measurements for the former, it is illegitimate and harmful for the latter. We argue that practitioners should deploy gender imputation only when it would achieve anti-discrimination benefits that cannot be achieved through other reasonable means, while harms are minimized to the extent possible. We examine this tension in three case studies: auditing gender bias in generative image models, measuring gender disparities in film, and imputing gender from personal names. By disentangling legitimacy from validity, and differentiating these two forms of sexism, we show how debates over gender prediction have conflated distinct concerns, obscuring both the settings in which gender imputation can support fairness efforts and the harms towards transgender and nonbinary people that it fundamentally cannot capture. We conclude by recommending the development of more inclusive methods that address all kinds of sexism.