Search papers, labs, and topics across Lattice.
This systematic literature review evaluates the evolution of Optical Character Recognition (OCR) technologies over the past decade, focusing on advancements in AI models that enhance text recognition across diverse languages and formats. By analyzing 97 studies published between January 2015 and January 2025, the review identifies key models, their strengths, limitations, and the ongoing challenges in the field, such as handling underrepresented languages and variability in handwritten text. The findings underscore the need for innovative approaches like self-supervised learning and multimodal AI to improve OCR accuracy and support real-time applications.
OCR technologies are evolving rapidly, yet significant challenges remain in recognizing diverse scripts and handwritten text, demanding innovative solutions for real-time applications.
Optical Character Recognition (OCR) for text recognition using machine vision has significantly improved, particularly when handling heterogeneous textual data. Traditional OCR models struggle with script variations, writing styles, and degraded documents. Advancements in technology are leading to new AI models with improved architecture for handling multiple languages and complex data formats. Despite this progress, a comprehensive evaluation of OCR advancements remains limited. Based on the established preferred reporting items for systematic reviews and meta-analysis (PRISMA) guidelines, this literature review presents an extensive assessment of OCR research to trace the evolution of AI models over the past decade. It explores the transition in AI models, application domains, data types, linguistic coverage, and challenges. Through a detailed analysis of 97 selected studies published during January 2015 - January 2025, key OCR models are identified, and their performance, strengths, and limitations are analyzed. The findings highlight how OCR technologies have evolved to address structured and unstructured text, scene text recognition, and multilingual processing. Unresolved challenges include limited resources for underrepresented languages, high variability in handwritten text, visual similarity among characters, and constraints in real-time OCR applications. To address these issues, several promising approaches are proposed. Key suggestions include self-supervised learning, multimodal AI, automated machine learning (AutoML), AI-assisted postprocessing, tiny machine learning (TinyML), and the creation of joint corpora for script matching. The future recommendations aim to enhance OCR accuracy and tackle the challenges identified for real-time industrial applications. This study will guide future research and establish a foundation for OCR field.