Search papers, labs, and topics across Lattice.
This study introduces CogArena, a benchmark designed to evaluate cognitive ability structures in large language models (LLMs) through a multimethod framework across 13 paradigms. The findings reveal that while there are positive correlations among cognitive-task scores, the expected dimensional profiles are not consistently stable across different model families, with only a small advantage observed for targeted scaffolds. Ultimately, the results suggest that the cognitive ability dimensions proposed may not be robustly established, challenging the validity of current cognitive profiling approaches in LLMs.
Cognitive ability dimensions in LLMs may not be as stable or meaningful as previously thought, with evidence suggesting a boundary conclusion on their validity.
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.