Search papers, labs, and topics across Lattice.
This paper introduces DigitalCoach, a multimodal dataset comprising 72 expert-novice coaching sessions that captures 22,752 dialogue turns and 28.1 hours of screen interactions across various software applications. The study evaluates state-of-the-art models in their ability to coach humans on software use, revealing that while models provide more direct instructions, they lack the depth of human coaching, such as explanations and error diagnoses. Ultimately, the findings highlight that model-generated coaching leads to passive learning, indicating significant gaps in visual grounding and engagement compared to human coaches.
Models may give clearer instructions, but they fail to engage learners deeply, resulting in passive instruction-following rather than active understanding.
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce DigitalCoach, a multimodal dataset of 72 human expert-novice computer use coaching sessions consisting of 22,752 dialogue turns grounded in 28.1 hours of screen and input event recordings across five software applications. We use DigitalCoach to evaluate whether state-of-the-art models can teach humans how to use computers. Automated evaluation shows that models differ from humans in how they coach: models provide more direct instructions, but fewer explanations, error diagnoses, and knowledge-check questions. When we fix the coaching method, models produce utterances similar to human references yet poorly grounded in visual context. Interactive evaluation confirms that model coaches cause learners to passively follow instructions without deeper engagement and fall short in visual grounding. DigitalCoach lays a foundation for collaborative and proactive computer use coaching agents.