Search papers, labs, and topics across Lattice.
The Teaching Monster Challenge benchmarks AI agents' Pedagogical Content Knowledge (PCK) by requiring them to generate instructional videos tailored to specific learner personas. While the systems excel at content generation, they struggle significantly with presentation and adaptation to the learner's needs, revealing a gap in their pedagogical capabilities. Additionally, the study highlights limitations in automatic judging, as the LLM-judge fails to accurately rank the top-performing systems compared to human evaluators, indicating a need for improved assessment methods alongside teaching systems.
AI agents can generate content but often miss the mark on effective teaching, struggling to adapt lessons to individual learners.
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been benchmarked. To measure it, we introduce the Teaching Monster Challenge, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion. Each system is given a topic and a learner persona and must generate a complete instructional video. Every video is screened by an LLM-judge, ranked by crowd pairwise voting, and finalized by an expert panel. The first edition shows that today's systems handle the content well but are far weaker at presenting it and adapting it to the learner. The same process exposes a limit of automatic judging. The LLM-judge separates a clear low-performing tail but ranks the strongest systems poorly. The strongest systems receive nearly identical scores from the judge, so its ranking of them does not match human preference. Progress therefore requires not only better teaching systems but also better automatic judges, and we release the benchmark, rubric, and human judgments as a testbed for both.