Search papers, labs, and topics across Lattice.
This study employs a mixed-methods human-LLM auditing framework to assess decision consistency and cognitive biases in the evaluation of neurodevelopmental disorders by both large language models (LLMs) and human experts. The findings reveal that both groups exhibit a disconnect between descriptive assessments of patients' functional levels and their support-eligibility decisions, with LLMs displaying greater intellectual humility yet mirroring the reductionist biases of human evaluators. Ultimately, the research highlights the necessity of critically examining the underlying concepts of neurodiversity that AI systems operationalize in high-stakes decision-making contexts.
LLMs express higher intellectual humility than human experts but still replicate their reductionist biases in neurodevelopmental assessments.
Large language models (LLMs) are increasingly supporting complex mental-health decisions, which depend not only on factual evidence but also value-laden interpretations. We introduce a mixed-methods human-LLM auditing framework examining decision consistency, susceptibility to cognitive heuristics, declarative intellectual humility, and the concepts operationalized in support-allocation judgments of neurodevelopmental disorders. Comparing 35 humans (18 physicians and 17 psychologists) with seven LLMs, we show that in both groups, ratings of patients' functional level were not significantly associated with support-eligibility decisions, indicating an inconsistency between descriptive assessments and final evaluative judgments. Specifically, we find that neither group showed significant susceptibility to experimental manipulations targeting anchoring and representativeness heuristics. LLMs reported higher intellectual humility than experts (U = 241, p < .001, r = .62; LLMs: M = 41.43, SD = 1.99; experts: M = 29.03, SD = 8.05), but it was unrelated to decision consistency or functional assessment. While LLMs and physicians granted support less frequently than psychologists (U = 180.50, p = .003, r = .34), they also interpreted a concept of "basic life needs" differently, primarily as biological survival and self-care, and not communicative and social needs. These findings suggest that despite expressing high levels of intellectual humility, LLMs reproduce a reductionist interpretive framework and knowledge embedded in medical decision-making. More broadly, we argue that evaluating AI in high-stakes contexts requires not only measuring accuracy, agreement, or resistance to cognitive bias, but also critical examination of the concepts of neurodiversity that AI systems operationalize.