Publication: Sequence and Structure-Based Machine Learning Model to Predict Cancer Neoantigen Immunogenicity
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Accurate prediction of the immunogenic potential of tumor-specific mutant peptides remains a central bottleneck in personalized cancer vaccine development. Existing computational approaches rely primarily on peptide–MHC binding affinity and sequence-derived features, often using non-cancer-specific models with limited biological interpretability. This thesis directly addresses these limitations by developing a novel pipeline that incorporates a comprehensive, literature-justified set of sequence and structural features to predict neopeptide immunogenicity with high accuracy, interpretability, and cancer specificity. The model was trained on a curated dataset of 2,263 experimentally validated peptide–MHC complexes, with three-dimensional structures generated using a custom AlphaFold-based pipeline. A total of 60 features spanning six mechanistic categories relevant to T cell activation were extracted and used to train, fine-tune, and evaluate a machine learning model. The resulting model achieved performance comparable to leading published methods, corresponding to a 45% relative improvement over the random baseline. Ablation analysis demonstrated that while incorporating explicit structural features offers strong predictive value, sequence does well to implicitly encode for structural geometries and interaction energetics. Post-hoc interpretability analysis additionally revealed key determinants of immunogenicity: peptide binding affinity to the MHC, foreignness relative to the self-proteome, and T cell receptor recognition constraints imposed by biochemical properties. Notably, AlphaFold confidence metrics emerged as strong predictors, suggesting that structure prediction biases act as proxies for peptide–MHC stability and presentation-related properties. Together, the model developed in this thesis contributes to the growing field of computational immunogenicity prediction by demonstrating that accurate prediction can be achieved within a fully interpretable framework, allowing for a better mechanistic understanding of how specific sequence and structural peptide properties affect the immune response in the context of innate cancer immunity. By linking predictive performance to underlying biological constraints, this work improves neopeptide selection and engineering and supports the development of next-generation personalized cancer vaccine immunotherapies.