Search papers, labs, and topics across Lattice.
This study investigates whether lesions or controlled perturbations to a multimodal language model can replicate the systematic naming errors observed in individuals with aphasia following stroke. By applying various perturbation configurations to the LLaVA 1.6 model, the authors found that six out of seven error categories emerged at clinically comparable proportions, successfully matching the error profiles of 97.8% of participants across multiple categories. These findings establish a quantitative framework for using language models as digital twins to simulate individual aphasic error patterns in picture naming tasks.
Language models can now mimic the complex error patterns of individuals with aphasia, potentially transforming our understanding of language disorders.
Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) whether lesions or controlled perturbations to a multimodal language model can reproduce different types of errors in picture naming, and (2) whether the framework can reproduce the complete error profile of individual persons with aphasia (PWAs). Using LLaVA 1.6, we evaluated perturbation configurations that varied the layer, proportion, and amount of noise applied to model units. We examined 278 PWAs on the Philadelphia Naming Test, classifying responses into seven categories using a validated neural classifier. Six of seven response categories (correct, semantic, mixed, unrelated, neologism, no response errors) emerged at clinically-comparable proportions across distinct parameter space regions, with formal paraphasia being the exception. Searching the perturbation space revealed configurations that reproduced the individual error profile in at least six of seven categories for 97.8% of PWAs and in all seven categories for 79.5% of PWAs. Monte Carlo baselines confirmed that this matching reflects joint inter-category structure rather than marginal overlap. These results establish a quantitative framework for reproducing individual aphasic error patterns in picture naming. They suggest the potential for language models to serve as digital twins of individuals with post-stroke aphasia.