Search papers, labs, and topics across Lattice.
This paper conducts a rapid review to establish guidelines for using Generative AI (GenAI) tools in systematic literature reviews (SLRs), addressing the challenges of ensuring rigor, reliability, and transparency. The authors present the GUEST framework, which outlines process recommendations for researchers to effectively integrate GenAI in SLRs while emphasizing the necessity of human oversight. Key findings reveal that while GenAI can assist with repetitive tasks, it is not yet capable of conducting unsupervised systematic studies, highlighting the importance of careful evaluation and reporting practices.
GenAI can streamline systematic literature reviews, but without human oversight, it risks compromising rigor and reliability.
Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require. Objectives: To support researchers intending to conduct SLRs using GenAI or those conducting empirical studies evaluating how well GenAI supports SLR tasks. Methods: First, we conducted a rapid review to identify studies that propose guidelines for evaluating and using GenAI and LLMs to support SLRs. Second, we drew on thought experiments, relevant guidance from the literature, and our own experience conducting SLRs and evaluating tools to develop recommendations for how to use and assess GenAI in the context of SLRs. Results: We discuss the problems researchers face when evaluating GenAI for SLRs. We identify and explain process issues to consider when planning, conducting, and reporting both SLRs using GenAI and evaluations of GenAI tools. Finally, we summarize our results as a set of process recommendations, which we name GUEST (GenAI Use and Evaluation in SLR Tasks). Conclusion: We argue that GenAI requires human oversight and is not currently capable of unsupervised systematic studies. However, it offers the prospect of cost-effective assistance for some repetitive tasks and for additional validation of some complex tasks. Our GUEST recommendations should help software engineering researchers both to conduct and report trustworthy SLRs using GenAI and to provide rigorous independent evaluation studies.