Search papers, labs, and topics across Lattice.
This study addresses the challenge of improving the realism of synthetic clinical benchmarks used in AI applications within healthcare, which often pass utility checks but lack structural realism. By formulating the problem as utility-constrained realism improvement, the authors revise a care-gap benchmark derived from Synthea-generated patients while ensuring it meets operational utility standards. The results demonstrate that two deterministic revisions significantly enhance the realism of the benchmarks without compromising their utility, highlighting the need for explicit optimization of synthetic benchmark quality beyond mere utility compliance.
Synthetic clinical benchmarks can be made significantly more realistic without sacrificing operational utility, challenging the assumption that utility alone guarantees quality. WHY_IT MATTERS: This work could fundamentally change how synthetic benchmarks are evaluated and optimized in healthcare AI, ensuring they are both useful and realistic.
Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especially in privacy-sensitive healthcare settings where operational data are hard to access. We study how to improve such benchmarks without breaking the downstream utility checks already used in practice. We formulate benchmark revision as utility-constrained realism improvement: dataset changes should increase realism while staying above an operational utility floor. We instantiate this idea on a care-gap benchmark derived from Synthea-generated patients exercised through demonstration electronic health record workflows and then processed by the same downstream pipeline as operational data. Realism is measured through missingness structure, simplicity, structural plausibility, and population alignment. The baseline benchmark is extremely thin: sampled-pair missingness is 79.44%, only 12.75% of rows are actionable, 38.94% of patients have zero actionable measures, and top-three token concentration reaches 100.0%. Two deterministic revisions improve these panels while remaining above the current utility floor, whereas a naive densification control preserves unrealistic templating. We further show that internal benchmark realism and source fidelity to an aggregate operational reference are related but distinct objectives. These results suggest that synthetic benchmark quality should be optimized explicitly, with utility treated as one constraint rather than as sufficient evidence of realism.