Repository logo

Thesis Central

Communities & Collections
Browse
Log In
  1. Home
  2. Browse by Author

Browsing by Author "Abdulhusein, Neamah"

Filter results by typing the first few letters
Now showing 1 - 1 of 1
  • Results Per Page
  • Sort Options
  • Loading...
    Thumbnail Image
    Item

    Extracting Empirical Statistics on Gender Bias From Word Embeddings: A Multilingual Analysis

    (2017-5-6) Abdulhusein, Neamah; Narayanan, Arvind

    Recent analyses on gender and other categories of bias in word embeddings have focused their efforts on developing algorithmic methods to remove these unwanted biases from the vector space models. However, little work has been done on drawing connections between the biases present in word embeddings and how they can inform us on the state of the environment in which the language is spoken. In this work, we explore the question of whether we can extract empirical information on the world around us from the vector space models of the languages we speak. We explore this question specifically with regards to gender bias. We conduct an intra-lingual experiment to determine whether gender associations of sports words in a language model L can predict the percentage of female participants from countries that speak L in those sports in the Summer Olympics. We conduct two inter-lingual experiments to determine whether the gender score of a language can be used to predict country-specific gender statistics, such as the UN Gender Inequality Index (GII). Our intra-lingual experiments in English, Spanish and Portuguese show highly significant results. Our inter-lingual experiment on predicting UN GII also shows significant potential for using the proposed metrics to predict country and language specific empirical statistics.

© 2024 The Trustees of Princeton University. All rights reserved.

  • Privacy policy
  • Accessibility
  • Send Feedback