Engineering selective amyloid precursor protein inhibitors by machine learning and deep mutational scanning

Deep mutational scanning (DMS) has proven effective for mapping protein–protein interactions (PPIs), but it cannot provide complete coverage of the mutation landscape, particularly for multi-mutant variants. To address this limitation, we trained machine-learning (ML) models on previously generated DMS data for a stabilized amyloid precursor protein inhibitor (APPI) binding to either of two serine proteases, mesotrypsin and kallikrein-6 (KLK6), which are implicated in various human disorders. We combined the models to accurately predict the binding selectivity of APPI variants, including double-mutant variants, for the two serine proteases. We achieved a Pearson correlation of 0.937 between predicted log2 selectivity enrichment ratios and DMS-derived values. We further validated the predictions of our combined model by yeast-surface-display measurements and inhibition assays of purified APPI variants and revealed epistatic interactions that shape protease selectivity. Guided by binding selectivity predictions, we identified highly selective APPI variants, including the most selective mesotrypsin inhibitor reported to date. Together, these findings support the use of DMS and ML as a framework for predicting PPI selectivity and prioritizing selective therapeutic protein variants.


This is a companion discussion topic for the original entry at https://onlinelibrary.wiley.com/doi/10.1002/pro.70712?af=R