Search papers, labs, and topics across Lattice.
This study introduces a pipeline for profiling AI agents and workplace tasks based on shared cognitive capabilities, addressing the challenge organizations face in determining which tasks can be automated or should remain human-led. By inferring an agent's capabilities from benchmark performance and eliciting task requirements from domain experts, the framework allows for independent updates as models and roles evolve. Validation on synthetic agents and profiling of six AI systems revealed significant differences in cognitive dimensions across AI models, providing a comparative tool for assessing AI suitability in various occupational contexts.
AI systems vary more in cognitive capabilities than in model families, revealing a nuanced landscape for task automation in the workplace.
Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where systems will succeed or fail in practice, while human judgements of model capabilities quickly become outdated. We introduce a pipeline that profiles agents and tasks using a shared set of core cognitive capabilities. Cognitive capability profiling infers an agent's capabilities from performance on a benchmark battery annotated for the cognitive demands of each item. Task requirements weighting elicits from domain experts the relative importance of these same capabilities for their work. As both use a common set of cognitive dimensions, they can be updated independently as models and roles change, and combined to estimate AI suitability at the level of a domain, organisation, role, or individual duty. We validate capability recovery on synthetic agents, profile six AI systems, and elicit task requirements from 410 employees across six occupational domains. AI systems differed more across cognitive dimensions than across model families, while workplace activities converged on a shared cognitive core. The resulting scores provide a comparative scoping tool for identifying promising candidates for piloting and areas where current systems are unlikely to be well suited. We discuss extending the framework to profile human workers alongside AI systems, moving from AI suitability towards human-machine task allocation.