Maximizing Interpretability in Credit Scoring through Decision Tree and Ensemble Learning Algorithms
| datacite.rights | restricted | |
| dc.contributor.advisor | Klusowski, Jason | |
| dc.contributor.author | Anteneh, Aemu | |
| dc.date.accessioned | 2022-07-29T14:52:08Z | |
| dc.date.accessioned | 2026-09-28T16:12:17Z | |
| dc.date.available | 2022-07-29T14:52:08Z | |
| dc.date.available | 2026-09-28T16:12:17Z | |
| dc.date.created | 2022-04-05 | |
| dc.date.issued | 2022-07-29 | |
| dc.description.abstract | The tension between interpretability and accuracy is heavily discussed in machine learning due to the huge implications this tradeoff has in a variety of applications. One such application is within credit scoring, specifically the FICO Score due to the black box nature of its algorithm. The lack of transparency of this model makes it susceptible to bias and inconsistencies, but the proprietary nature of FICO makes the existing literature on its implementation sparse. We address this issue by using a Decision Tree model to replicate the FICO Score algorithm in an interpretable way while maintaining good accuracy in credit score classification. This choice of algorithm is motivated by previous success of tree-based methods in determining creditworthiness, as well as the sequential nature of the online myFICO Credit Score Estimator resembling the decision making process of trees. We then use the ensemble learning methods Random Forest and Gradient Boosting to examine what bias may exist in FICO’s algorithm by comparing permutation feature importance scores between two datasets, one that mimics the factors that FICO uses in scoring (our “traditional” dataset) and another that includes that information as well as additional features that are supposedly non diagnostic (our “enhanced” dataset). Across all three models, there is a meaningful difference between the top significant features when the additional variables are included in the dataset, suggesting that the FICO algorithm may be implicitly placing heavy weight on unethical variables in determining credit scores. Furthermore, accuracy of each of our methods using both sub datasets was consistently good. This study gives evidence that a transparent credit scoring model that maintains performance is not only achievable, but is an extremely important step to monitoring bias in assigning credit score classifications. | en_US |
| dc.format.mimetype | application/pdf | |
| dc.identifier.uri | http://arks.princeton.edu/ark:/88435/dsp01g158bm478 | |
| dc.identifier.uri | https://theses-dissertations.princeton.edu/handle/88435/dsp01g158bm478 | |
| dc.language.iso | en | en_US |
| dc.title | Maximizing Interpretability in Credit Scoring through Decision Tree and Ensemble Learning Algorithms | en_US |
| dc.type | Princeton University Senior Theses | |
| pu.certificate | Program in Cognitive Science | en_US |
| pu.contributor.authorid | 920208954 | |
| pu.date.classyear | 2022 | en_US |
| pu.department | Operations Research and Financial Engineering | en_US |
| pu.mudd.walkin | No | en_US |
| pu.pdf.coverpage | SeniorThesisCoverPage |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- ANTENEH-AEMU-THESIS.pdf
- Size:
- 1.13 MB
- Format:
- Adobe Portable Document Format
Download
License bundle
1 - 1 of 1
Loading...
- Name:
- LICENSE.txt
- Size:
- 249 B
- Format:
- Plain Text
- Description:
Download