Search papers, labs, and topics across Lattice.
This study employs Exploratory Factor Analysis (EFA) to compare the latent structures underlying human and LLM performance on quantitative reasoning and chemistry assessments. The findings reveal that while subject-matter experts can interpret most human-derived factors, they struggle to ascribe meaningful interpretations to LLM-derived factors, indicating a significant divergence in the cognitive constructs employed by LLMs versus humans. This insight underscores the limitations of current evaluation methods that assume alignment between human and AI cognitive processes.
LLMs operate on fundamentally different cognitive constructs than humans, leaving experts unable to interpret their performance metrics meaningfully.
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.