Search papers, labs, and topics across Lattice.
The authors evaluate retrieval-augmented in-context learning (RetICL) and structured statutory definitions across multiple LLMs for detecting criminally defamatory online posts under German law. Dynamic exemplar retrieval offers negligible gains over random few-shot baselines and underperforms optimized static prompts, with underlying base model selection driving nearly all performance variance. Despite aggressively over-predicting criminality, all tested models still miss 26–57% of true defamatory violations, highlighting a fundamental failure mode in nuanced legal boundary detection.
Retrieval-augmented prompting fails to beat hand-tuned static exemplars for legal moderation, and even top LLMs miss up to 57% of criminal hate speech despite aggressive over-policing.
With hate speech being ubiquitous online, automatic detection is crucial, in particular when it comes to criminally relevant social media posts. We study a variety of retrieval-based in-context learning (RetICL) strategies for detecting defamatory offences under {\S}{\S} 185-187 StGB (the subject of GermEval 2026 Subtask 4). Few-shot prompting beats zero-shot, but retrieval-based approaches offer only marginal gains over random demonstrations, and even fall behind an optimised static set of demonstrations. Providing concrete legal knowledge helps, yet model choice outweighs every other system choice. Models over-predict criminal relevance while still missing 26-57% of criminally relevant posts, suiting them for triage rather than autonomous moderation.