A Stochastic Landscape Approach for Protein Folding State Classification.
Michael FaranDhiman RayShubhadeep NagUmberto RaucciMichele ParrinelloGili BiskerPublished in: Journal of chemical theory and computation (2024)
Protein folding is a critical process that determines the functional state of proteins. Proper folding is essential for proteins to acquire their functional three-dimensional structures and execute their biological role, whereas misfolded proteins can lead to various diseases, including neurodegenerative disorders like Alzheimer's and Parkinson's. Therefore, a deeper understanding of protein folding is vital for understanding disease mechanisms and developing therapeutic strategies. This study introduces the Stochastic Landscape Classification (SLC), an innovative, automated, nonlearning algorithm that quantitatively analyzes protein folding dynamics. Focusing on collective variables (CVs) - low-dimensional representations of complex dynamical systems like molecular dynamics (MD) of macromolecules - the SLC approach segments the CVs into distinct macrostates, revealing the protein folding pathway explored by MD simulations. The segmentation is achieved by analyzing changes in CV trends and clustering these segments using a standard density-based spatial clustering of applications with noise (DBSCAN) scheme. Applied to the MD-based CV trajectories of Chignolin and Trp-Cage proteins, the SLC demonstrates apposite accuracy, validated by comparing standard classification metrics against ground-truth data. These metrics affirm the efficacy of the SLC in capturing intricate protein dynamics and offer a method to evaluate and select the most informative CVs. The practical application of this technique lies in its ability to provide a detailed, quantitative description of protein folding processes, with significant implications for understanding and manipulating protein behavior in industrial and pharmaceutical contexts.