Search papers, labs, and topics across Lattice.
This dissertation introduces an AI-guided learning framework that enhances knowledge and skill acquisition through three innovative systems: AIxSpeed, FastPerson, and Profy. AIxSpeed optimizes audio playback speed based on listening difficulty, achieving an average playback factor of 1.30x while improving user satisfaction. FastPerson and Profy further streamline learning by providing multimodal video summaries and effective pronunciation practice, respectively, demonstrating significant reductions in content consumption time and improvements in skill acquisition without sacrificing comprehension or intelligibility.
AIxSpeed lets learners consume audio content 30% faster while enhancing comprehension and satisfaction, reshaping how we approach learning efficiency.
Audio and video have become major learning media, but learners face two persistent challenges: the time cost of consuming long-form content sequentially and the lack of scalable feedback for imitation-based skill acquisition. This dissertation proposes an AI-guided learning framework that supports three interconnected stages: Consume, Understand, and Imitate. It develops and evaluates three systems. AIxSpeed dynamically adjusts audio playback speed at the phoneme level using speech-recognition-model confidence as a proxy for listening difficulty. FastPerson generates multimodal video summaries that preserve visual and auditory information and lets learners switch between summarized and full versions by chapter. Profy learns proficiency from largely unannotated speech data and visualizes classifier-relevant regions and model-derived acoustic distances to support pronunciation practice. Technical and user evaluations show that AIxSpeed achieved average playback factors of 1.30x on LibriSpeech and 1.29x on UME-ERJ and received higher mean opinion scores than matched constant-speed playback; FastPerson reduced viewing time by 53% with no statistically significant difference in quiz scores compared with normal playback; and Profy showed an observed improvement in pronunciation intelligibility, with non-overlapping pre- and post-practice confidence intervals. Together, these systems demonstrate how deep learning can support efficient content consumption, multimodal understanding, and repeated skill practice while retaining learner access to the original material.