Search papers, labs, and topics across Lattice.
This study conducts a blind Turing Test to evaluate the performance of leading out-of-the-box LLMs on three Italian legal professional exams: the Bar, Judges, and Notary exams. The results indicate that while some LLMs can match or surpass human performance in adversarial legal argumentation and doctrinal analysis, they consistently underperform in the notary exam, which demands rigorous goal-directed legal planning. The research highlights both the strengths and limitations of LLMs in legal contexts, revealing specific failure patterns across different tasks.
Some LLMs can outperform humans in legal argumentation, but all struggle with the complex planning required for notary exams.
The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and recurring legal failure patterns. Although limited to out-of-the-box systems, the findings provide qualitative evidence on the current scope and boundaries of the legal competence of LLMs across distinct professional tasks.