Search papers, labs, and topics across Lattice.
This paper introduces ProxyGuard, a novel framework for ensuring reliability in randomized data release mechanisms by controlling errors through bounded risks and a sealed target set. The method effectively corrects for multiplicity and evaluates independent mechanism draws, significantly improving power from 5.6% to 64.2% at a reliability level of 0.95 in a registered study. By providing finite-sample reliability guarantees without requiring independent target batches, ProxyGuard enhances the robustness of data releases in statistical analyses.
Directly improving reliability in randomized data releases, ProxyGuard boosts statistical power dramatically while ensuring valid inference.
Researchers often choose a proxy dataset from many releases, transformations, or seeds. Search can make an invalid release appear adequate, while one adequate release does not establish that its generator is reliable. ProxyGuard controls both errors using prespecified bounded risks and a sealed target set. Named-release mode corrects for multiplicity and certifies specific releases. Direct shared-target mode evaluates independent mechanism draws on a common target, lower-bounds their favorable-score rate, and subtracts a bound on favorable scores contributed by invalid releases. Conditional on the target, release scores are independent, yielding a finite-sample mechanism-reliability guarantee without independent target batches or assumptions on release-level $p$-value dependence. We show that the mean-only penalty is sharp and derive a smooth-score certificate with additive target concentration. In a registered three-requirement study, direct mode raises power from 5.6\% to 64.2\% at reliability 0.95, while named mode remains stronger under high-signal evidence. Prospective audits span full-pipeline Rice--TVAE, which retrains on every draw, and a non-tabular text mechanism.