Search papers, labs, and topics across Lattice.
This paper introduces the Scaffolded Task Design (STaD) framework, which systematically identifies compositional skill gaps in large language models (LLMs) by generating controlled variations of benchmark tasks. By employing a scaffolding approach, STaD provides structured support that allows researchers to probe model behavior more effectively, revealing specific reasoning skills that LLMs lack. Experiments across six models demonstrate that each model exhibits unique failure points, underscoring the necessity for targeted improvements in LLM capabilities.
LLMs reveal distinct compositional skill gaps that traditional benchmarks fail to capture, highlighting the need for tailored interventions.
Benchmarks are often used as a standard to understand LLM capabilities in different domains. However, aggregate benchmark scores provide limited insight into compositional skill gaps of LLMs and how to improve them. To make these weaknesses visible, we propose Scaffolded Task Design (STaD) framework. STaD generates controlled variations of benchmark tasks based on the concept of scaffolding, which introduces structured, incremental support in a step-by-step manner. Rather than inspecting failures individually, this approach enables systematic and scalable probing of model behavior by identifying the specific reasoning skill compositions they lack. Treating the LLM as a black box, our experiments on six models of varying sizes reveal multiple failure points in three reasoning benchmarks and highlight each model's unique and distinct skill gaps.