Search papers, labs, and topics across Lattice.
This study details the development and validation of the GenAI Literacy Test (GenAIT), an 18-item multiple-choice assessment designed to measure high school students' understanding of generative AI across technical, practical, and human-impact domains. Utilizing a large-scale survey of 7,432 Estonian students, the authors employed confirmatory factor analysis and item response theory to evaluate the test's psychometric properties, revealing adequate reliability and good model fit. Notably, the findings indicate that GenAIT scores do not correlate with students' perceptions of AI usefulness or ease of use, challenging assumptions that frequent AI use equates to conceptual understanding.
Frequent use of generative AI tools does not guarantee a solid understanding of their underlying concepts, as shown by the GenAIT test results.
There is growing international interest in generative AI (GenAI) literacy and its assessment among high school students, but objective assessment in this population remains underdeveloped. This article reports the iterative development and validation of the GenAI Literacy Test (GenAIT), an 18-item multiple-choice test measuring high school students'conceptual knowledge about GenAI, with content spanning technical, practical, and human-impact domains. Expert review of relevance, clarity, and comprehensiveness provided evidence of content validity. In a large-scale survey of 7432 Estonian high school students, we evaluated the psychometric functioning of the Estonian-language GenAIT using confirmatory factor analysis, classical test theory, and item response theory. Results supported approximate unidimensionality, broadly adequate reliability for group-level research (marginal reliability = .72, KR-20 = .69), and good fit of a three-parameter logistic model (RMSEA = .013, TLI = .987, CFI = .990, SRMSR = .021). Measurement precision was sufficient for the majority of students but varied substantially across the latent trait, with lower precision for lower scoring students. GenAIT is therefore more suitable for group-level research than high-stakes individual classification. GenAIT scores were unrelated to perceived usefulness and perceived ease of use, and negatively associated with LLM use frequency, suggesting that frequent use and favorable perceptions of AI should not be treated as proxies for conceptual understanding.