Search papers, labs, and topics across Lattice.
This paper critiques the limitations of Fr茅chet Inception Distance (FID) and Kernel Inception Distance (KID) in evaluating generative models, particularly their inability to capture nuanced distributional differences. The authors introduce ZID (Z-resolved Integrated Diagnostic), a novel metric that provides a comprehensive analysis by reporting a ranking index, a permutation p-value for distributional equality, and a signed dispersion readout. In controlled experiments, ZID effectively identifies a wider range of deviations than FID, including cases where FID fails to reflect significant changes in model performance, such as mode collapse.
ZID reveals critical distributional insights that FID and KID overlook, enabling researchers to diagnose generative model failures with unprecedented precision.
Generative models are commonly ranked by Fr茅chet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance obtain FID $24.7$ versus $58.6$ for held-out real images (lower is better). Moreover, FID and KID are scalar discrepancies that are unchanged when the two samples are exchanged and therefore do not encode the direction of a dispersion change: under-dispersion, as can occur in mode collapse, versus over-dispersion. We introduce \textbf{ZID} (\emph{Z-resolved Integrated Diagnostic}), which combines six standardized location- and dispersion-sensitive arms from a rank graph (RISE) and Gaussian kernels (GPK at two bandwidths). Rather than asking one scalar to serve incompatible roles, ZID reports three linked outputs: an index for ranking departure magnitude, a permutation $p$-value for testing distributional equality, and a signed dispersion readout for diagnosis. In controlled experiments, ZID detects a broad range of departures, and its score tracks increasing severity along the corresponding sweeps, including cases in which FID is flat or reversed. On DiT-XL/2 and SiT-XL/2 guidance sweeps, ZID detects departure from real data, and its signed readout labels the high-guidance diversity collapse as under-dispersion.