Search papers, labs, and topics across Lattice.
This paper introduces Atomic Policy Optimization (APO), an unsupervised framework for predicting the 3D structures of atomic systems, addressing the limitations of supervised preference learning that relies on costly ground-truth labels. By employing a dual-reward mechanism that reinforces dominant latent structural modes and enforces thermodynamic stability, APO enables self-correction of physically plausible configurations. Extensive benchmarks reveal that APO outperforms fully supervised models in match rates and structural fidelity, establishing a new state-of-the-art in the field.
Intrinsic physical consistency outperforms noisy supervised learning in predicting atomic structures, achieving unprecedented accuracy without the need for costly ground-truth labels.
Predicting the 3D structures of atomic systems is fundamental to advancing material science and drug discovery. While flow-matching models (, FlowDPO) have recently shown promise in this domain, their performance relies heavily on alignment with ground-truth coordinates via supervised preference learning. However, obtaining experimental labels for novel crystal phases or de novo proteins is prohibitively expensive, creating a bottleneck for structural modeling in data-scarce regimes. In this work, we propose (Atomic Policy Optimization), a fully unsupervised alignment framework that eliminates the need for ground-truth reference structures. APO adapts group-relative policy optimization to 3D atomic environments, utilizing a novel dual-reward mechanism: (i) a that reinforces the policy's dominant latent structural modes through eigen-decomposition of sample similarities, and (ii) a that enforces thermodynamic stability. Our framework enables the model to ``self-correct''by identifying physically plausible configurations within sampled groups. Extensive benchmarks on crystal and antibody structure prediction demonstrate that APO consistently outperforms fully supervised baselines, achieving a new state-of-the-art in match rates and structural fidelity. Furthermore, we show that APO effectively straightens probability paths, significantly improving inference efficiency. Our results suggest that intrinsic physical consistency can serve as a superior guide for alignment compared to noisy, supervised coordinate matching.