Improving AI-based Synthesizability Scores for Next-Generation Protein Degradation Drugs
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
PROteolysis TArgeting Chimeras (PROTACs) are an emerging class of therapeutic modality that enables targeted protein degradation. A PROTAC consists of three components: a target-binding warhead, a chemical linker, and an E3 ligase ligand. By simultaneously binding a Protein Of Interest (POI) and an E3 ubiquitin ligase, the PROTAC induces ubiquitination of the target protein, leading to its recognition and degradation by the proteasome. Despite their therapeutic potential, the modular structure of PROTACs often necessitates complex multi-step synthesis. Evaluating their synthesizability typically requires extensive laboratory synthesis and validation, making the process time-consuming, costly, and expertise-intensive. In this thesis, we present a systematic framework for PROTAC synthesizability assess ment based on retrosynthetic planning and machine learning. A dataset with 27,099 PROTACs was collected, curated, and preprocessed to generate retrosynthesis-based synthesizability labels and scores. PROTAC-Splitter was used to decompose each PROTAC into its constituent warhead, linker, and E3 ligase ligand components, while AiZynthFinder was employed to evaluate retrosynthetic solvability and generate syn thesizability scores for both complete molecules and individual components. Relevant features, including molecular fingerprints, molecular descriptors, and component-level synthesizability information, were subsequently used to train machine learning models based on Random Forest, XGBoost, and Multi-Layer Perceptron architectures. Retrosynthetic analysis revealed a 57% agreement between component-level and whole-PROTAC synthesizability, demonstrating that component-level assessment captures meaningful information about overall synthetic feasibility. Linkers were identified as the primary source of synthetic difficulty, highlighting their importance in determining PROTAC synthesizability. Overall, our best ML classification model achieved a ROC-AUC of 0.958, the best regression model achieved an (R2) score of 0.497, increasing to 0.661 after filtering noisy labels. These results demonstrate that component-level retrosynthetic information can be leveraged to construct efficient surrogate models for PROTAC synthesizability assessment, providing a scalable alternative to computationally expensive retrosynthetic planning.