Extracting Empirical Statistics on Gender Bias From Word Embeddings: A Multilingual Analysis

datacite.rightsrestricted
dc.contributor.advisorNarayanan, Arvind
dc.contributor.authorAbdulhusein, Neamah
dc.date.accessioned2017-07-20T14:19:24Z
dc.date.accessioned2026-09-29T21:55:45Z
dc.date.available2017-07-20T14:19:24Z
dc.date.available2026-09-29T21:55:45Z
dc.date.created2017-05-06
dc.date.issued2017-5-6
dc.description.abstractRecent analyses on gender and other categories of bias in word embeddings have focused their efforts on developing algorithmic methods to remove these unwanted biases from the vector space models. However, little work has been done on drawing connections between the biases present in word embeddings and how they can inform us on the state of the environment in which the language is spoken. In this work, we explore the question of whether we can extract empirical information on the world around us from the vector space models of the languages we speak. We explore this question specifically with regards to gender bias. We conduct an intra-lingual experiment to determine whether gender associations of sports words in a language model L can predict the percentage of female participants from countries that speak L in those sports in the Summer Olympics. We conduct two inter-lingual experiments to determine whether the gender score of a language can be used to predict country-specific gender statistics, such as the UN Gender Inequality Index (GII). Our intra-lingual experiments in English, Spanish and Portuguese show highly significant results. Our inter-lingual experiment on predicting UN GII also shows significant potential for using the proposed metrics to predict country and language specific empirical statistics.en_US
dc.identifier.urihttp://arks.princeton.edu/ark:/88435/dsp01pz50gz72c
dc.identifier.urihttps://theses-dissertations.princeton.edu/handle/88435/dsp01pz50gz72c
dc.language.isoen_USen_US
dc.titleExtracting Empirical Statistics on Gender Bias From Word Embeddings: A Multilingual Analysisen_US
dc.typePrinceton University Senior Theses
pu.certificateCenter for Statistics and Machine Learningen_US
pu.contributor.advisorid960831988
pu.contributor.authorid960889762
pu.date.classyear2017en_US
pu.departmentComputer Scienceen_US
pu.pdf.coverpageSeniorThesisCoverPage

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
neamah_written_final_report_bound.pdf
Size:
995.32 KB
Format:
Adobe Portable Document Format
Download