Search papers, labs, and topics across Lattice.
This systematic mapping study synthesizes 33 primary studies on the quality of AI-based software, revealing six key challenge categories that hinder effective quality assessment. The most significant challenge identified is the inadequacy of existing quality assessment models, alongside issues related to non-functional requirement management and quality assurance practices. These findings underscore the urgent need for collaboration between researchers and industry practitioners to establish standardized quality assessment frameworks for AI software.
Existing quality assessment models for AI software are failing, highlighting a critical gap that could jeopardize software reliability.
Artificial Intelligence (AI) is increasingly embedded in modern software systems, raising important questions about how its quality should be defined, assessed, and assured. This paper presents a Systematic Mapping Study (SMS) on the quality of AI-based software. The study synthesizes primary studies published between January 2020 and January 2026 and selected from five electronic data sources. A total of 33 primary studies were included after automated search, screening, and snowballing. The results identify six recurring challenge categories, with the most prominent being limitations in existing quality assessment models, followed by issues in non-functional requirement management, quality-aware development, and quality assurance. The findings suggest a call for collaboration of researchers and industrial practitioners with standardization organizations, that could possibly devise comprehensive quality assessments and their measurement methods.