Search papers, labs, and topics across Lattice.
This paper introduces DegradeQuery, a context-aware prediction framework that leverages label-missing records from PROTAC databases to enhance degradation prediction by employing a counterfactual tuple pretraining objective. By contrasting existing tuples with alternatives formed by modifying the target protein or E3 ligase, the model learns contextual associations without needing pseudo-labels for activity. The approach achieves a notable improvement in performance on the PROTAC-8K benchmark, with an area under the ROC curve of 0.9065, highlighting the potential of utilizing incomplete data for effective model training.
Incompletely labeled PROTAC databases can still yield powerful predictors for protein degradation, achieving state-of-the-art performance without extensive experimental labels.
Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joint outcome of the degrader molecule and its biological context. Although public databases contain thousands of structured molecule-target-E3 records, degradation measurements are available for only a small fraction of them. Existing supervised approaches therefore leave most recorded chemical-biological relationships unused. We introduce DegradeQuery, a context-aware prediction framework that converts these label-missing records into a pretraining signal. Its counterfactual tuple pretraining objective contrasts recorded tuples with alternatives formed by replacing the target, the E3 ligase, or both, enabling the model to learn contextual associations without assigning activity pseudo-labels. The resulting representation is then fine-tuned to predict degradation from the complete molecule-target-E3 context. On the official PROTAC-8K benchmark, DegradeQuery achieves an area under the receiver operating characteristic curve of 0.9065 and an accuracy of 0.8500, outperforming the compared methods. Controlled analyses further show that the improvement is primarily attributable to tuple-level pretraining, can be recovered using only label-missing records, and remains complementary to protein language model representations. These findings demonstrate that incompletely labeled PROTAC databases contain useful relational supervision and provide a practical route for learning context-aware degradation predictors from scarce experimental labels.