Search papers, labs, and topics across Lattice.
This study evaluates the effectiveness of protein backbone generation models in producing genuinely novel folds by introducing the Domain Retrieval Rate (DRR) metric, which assesses the extent to which generated structures align with known protein domains. The analysis of eight models reveals that while many outputs contain known structural elements, the true emergence of new folds is significantly lower than previously reported, highlighting a discrepancy in the perceived novelty of generated structures. Additionally, the authors introduce RetFold, a zero-training baseline that efficiently generates protein backbones by retrieving known domains and optimizing their connections, demonstrating a substantial reduction in computational cost.
Most protein backbone generation models may be recycling known structures rather than creating truly novel folds, as revealed by the new Domain Retrieval Rate metric.
Protein backbone generation models are often credited with exploring novel fold space based solely on low full-chain similarity to known proteins, yet this cannot distinguish a genuinely new fold from a novel assembly of known structural units. We first ask whether this granularity mismatch alone explains the reported rates, and introduce the Domain Retrieval Rate (DRR), the fraction of generated backbones for which any constituent domain matches a known domain in CATH S40. Applied to eight backbone generation models spanning diffusion and flow-matching paradigms, DRR finds locally alignable known structure in most outputs, while the fraction containing a substantially covered complete domain is considerably smaller and depends on the scoring convention. To calibrate what retrieval alone can achieve, we propose RetFold, a zero-training baseline that constructs backbones by retrieving CATH domains and refining inter-domain connections through geometry-based helix-linker optimization, at two orders of magnitude lower cost on CPU alone.