Microsoft ResearchJun 9, 2026arXiv:2606.10956

Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?

Tengchao Lv, Dongdong Zhang, Jiayu Ding, Yilin Jia, Yuzhong Zhao, Yupan Huang, Wenshan Wu, Xiangyang Zhou, Shaohan Huang, Nan Yang, Li Dong, Lei Cui, Furu Wei

AI Summary

This study evaluates the performance of seven frontier Large Language Models (LLMs) on a standardized Office proficiency exam, specifically designed to assess their document-automation capabilities in a complex, multi-application environment. The evaluation, based on China's National Computer Rank Examination, reveals that even the most advanced models score only 36.6% in single-turn tasks, while a more sophisticated agentic system with feedback and iterative repair achieves 68.8%, still falling short of the community reference score of 95.5%. These findings highlight the significant limitations of current LLMs in executing fine-grained Office document automation, despite advancements in related technologies.

Key Contribution

LLMs struggle with Office automation, scoring only 36.6% on a standardized proficiency exam, revealing a critical gap in their capabilities.

Abstract

The deployment of Large Language Model (LLM) agents for computer automation is accelerating, yet their ability to navigate complex, professional-grade productivity software is largely untested. We argue that Office automation is an ideal environment for benchmarking document-automation capability, as it requires long-horizon planning and reasoning, precise parameter configuration, and multi-application integration. To quantify this capability, we introduce an evaluation based on China's National Computer Rank Examination (NCRE), featuring 200 comprehensive practical-operation tasks across Word, Excel, and PowerPoint. Each task is scored on a 100-point rubric scale using 7,118 machine-gradable criteria, and Score Rate (SR) denotes the mean percentage of rubric points earned across these tasks. We benchmark 7 frontier LLMs and observe stark limitations: single-turn models score a maximum of 36.6%. A stronger agentic system with execution feedback, iterative repair, and broader Office automation access reaches 68.8%, but remains below the 95.5% community-reference score used as a scoring sanity check. Ultimately, our experiments demonstrate that despite recent advancements in code generation, achieving reliable fine-grained Office document automation remains a significant challenge for current code-generating LLM and agent systems.

Eval Frameworks & Benchmarks Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?

Related Papers