Search papers, labs, and topics across Lattice.
This study outlines a multi-stage workflow designed to prioritize substantive review needs for the extensive web portfolios of German statutory health insurance (SHI) funds, addressing the limitations of generic AI-text detection in identifying specific review requirements. By analyzing 56,198 pages from 84 SHI websites, the workflow employed a combination of deterministic screening, model-assisted triage, and in-depth review to generate 35,998 review records, with a significant focus on transparency, legal framing, and medical content. The findings highlight that while the workflow effectively routed 21,452 pages for case review, it does not validate AI authorship or autonomous detection, emphasizing the necessity for human adjudication in public claims.
A novel workflow prioritizes review needs in health insurance content, revealing that 33.3% of pages flagged for review contained significant AI-related failure signals.
Background: German statutory health insurance (SHI) funds publish web portfolios that exceed continuous specialist review capacity. Their content can shape health and benefit expectations. Generic AI-text detection does not identify medical, benefit, legal, or editorial review needs. Objective: To characterize a multi-stage workflow that prioritizes substantive review needs while separating AI-provenance signals from quality claims. Methods: We analyzed 56,198 pages from 84 SHI websites or sub-sites. The workflow combined deterministic screening, model-assisted triage and in-depth review, minimum evidence checks, temporal-validity safeguards, and paired-model comparison. It is reproducibility-bounded, not a validated detector. Production code is proprietary; reproducibility rests on frozen derived tables and paired-comparison artifacts. The 300-page lower-priority check was a single-model, risk-enriched routing stress test, not a human-reference evaluation. Results: All pages received a review state. The workflow generated 35,998 review records and routed 21,452 to case review. The workload concentrated in transparency, legal framing, medical content, contradictions, and AI-related failure-mode signals. A quoted passage was locatable in captured page text for 31,347 records, confirming literal occurrence rather than factual correctness. The routing stress test surfaced a signal on 100/300 pages (33.3% within the sample). Across 182 matched cases, two models agreed in 75.8% (kappa = 0.532; 95% CI 0.415-0.649). Conclusions: The workflow produces a prioritized workload, not error prevalence or final legal, medical, or insurer-level findings. It neither proves AI authorship nor validates autonomous detection. Paired-model agreement quantifies consistency, not correctness or sufficient triage performance; public claims require human adjudication.