Maximizing Interpretability in Credit Scoring through Decision Tree and Ensemble Learning Algorithms

datacite.rightsrestricted
dc.contributor.advisorKlusowski, Jason
dc.contributor.authorAnteneh, Aemu
dc.date.accessioned2022-07-29T14:52:08Z
dc.date.accessioned2026-09-28T16:12:17Z
dc.date.available2022-07-29T14:52:08Z
dc.date.available2026-09-28T16:12:17Z
dc.date.created2022-04-05
dc.date.issued2022-07-29
dc.description.abstractThe tension between interpretability and accuracy is heavily discussed in machine learning due to the huge implications this tradeoff has in a variety of applications. One such application is within credit scoring, specifically the FICO Score due to the black box nature of its algorithm. The lack of transparency of this model makes it susceptible to bias and inconsistencies, but the proprietary nature of FICO makes the existing literature on its implementation sparse. We address this issue by using a Decision Tree model to replicate the FICO Score algorithm in an interpretable way while maintaining good accuracy in credit score classification. This choice of algorithm is motivated by previous success of tree-based methods in determining creditworthiness, as well as the sequential nature of the online myFICO Credit Score Estimator resembling the decision making process of trees. We then use the ensemble learning methods Random Forest and Gradient Boosting to examine what bias may exist in FICO’s algorithm by comparing permutation feature importance scores between two datasets, one that mimics the factors that FICO uses in scoring (our “traditional” dataset) and another that includes that information as well as additional features that are supposedly non diagnostic (our “enhanced” dataset). Across all three models, there is a meaningful difference between the top significant features when the additional variables are included in the dataset, suggesting that the FICO algorithm may be implicitly placing heavy weight on unethical variables in determining credit scores. Furthermore, accuracy of each of our methods using both sub datasets was consistently good. This study gives evidence that a transparent credit scoring model that maintains performance is not only achievable, but is an extremely important step to monitoring bias in assigning credit score classifications.en_US
dc.format.mimetypeapplication/pdf
dc.identifier.urihttp://arks.princeton.edu/ark:/88435/dsp01g158bm478
dc.identifier.urihttps://theses-dissertations.princeton.edu/handle/88435/dsp01g158bm478
dc.language.isoenen_US
dc.titleMaximizing Interpretability in Credit Scoring through Decision Tree and Ensemble Learning Algorithmsen_US
dc.typePrinceton University Senior Theses
pu.certificateProgram in Cognitive Scienceen_US
pu.contributor.authorid920208954
pu.date.classyear2022en_US
pu.departmentOperations Research and Financial Engineeringen_US
pu.mudd.walkinNoen_US
pu.pdf.coverpageSeniorThesisCoverPage

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
ANTENEH-AEMU-THESIS.pdf
Size:
1.13 MB
Format:
Adobe Portable Document Format
Download

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
LICENSE.txt
Size:
249 B
Format:
Plain Text
Description:
Download