SLAI Natural Language Processing
Lab 4: Measuring Differences in Word Usage with Embeddings
Word embeddings are commonly used in many NLP applications, from text generation to classification. By calculating distances from other words in the embedding space, we are able to gain some understanding of the contexts that words are used in. In this lab, we will compare neighbors of words across embeddings trained on data from 1900 and 1990 to see which words have changed in the intervening period and how usage differs.
Learning goals
Once you complete this lab, you should:
- Understand one method of determining language change using word embeddings
- Know how to use the
KeyedVectorsclass of the gensim word2vec library
Load embeddings
The load_emb function in can be used to load word embeddings as KeyedVectors. You can see that the function has one parameter, emb_set, which should be the year that you are loading embeddings for.
Because you are loading embeddings for 1900 and 1990, those should be the arguments that you use to call the function!
Filtering vocabulary
The load_emb function automatically ensures that you are working with a shared vocabulary space, which means that you only consider words that are in the vocabulary of the 1900 and 1990 embeddings. There are also a number of other vocabulary filtering rules are implemented, most of which are drawn from Gonen et al.
You should figure out which words are in the vocabulary, using key_to_index as demonstrated in the code that was given to you.
Computing scores
To compute the extent to which word usage has changed between 1900 and 1990, we will use the method from Gonen et al.1 The method works as follows:
- For each word in the vocabulary, find the 500 nearest neighbors in each embedding space.
- We will find the “nearest neighbors” using cosine similarity.
- Formally, let’s call these neighbors of word \(w\) \(NN_{1900}(w)\) and \(NN_{1990}(w)\) for 1900 and 1990, respectively.
- The score for a word can be computed using the following equation
- \[score(w) = - | NN_{1900}(w) \cap NN_{1990}(w) |\]
- \(\cap\) (intersection) is the words that appear in both \(NN_{1900}(w)\) and \(NN_{1990}(w)\)
- \(\vert x \vert\) computes the cardinality of set \(x\), which is the number of items in the set
- You may find it useful to use the python set library when computing scores
Compute scores for all words in the vocabulary using a loop (strong suggestion: use tqdm). Store the values in a dictionary as you go.
Which words have changed?
Finally, you’ll want to determine which words have changed the most between 1900 and 1990. Think about whether words with the highest or lowest scores have changed the most!
To visualize the results, I recommend printing the 20 words with the most change between 1900 and 1990, along with their ten nearest neighbors in each of the 1900 and 1990 embeddings.
To get the words with the most extreme scores, you can use the
get_max_or_min_keysfunction.
See what you think of the results - can you see how and why some words have changed between 1900 and 1990?
Extensions
If you’ve finished the main part of the assignment, consider building on your analysis in one of the following ways.
Differences in word usage during COVID-19
What if we want to determine how word usage has changed more recently by comparing the periods before and after start of the COVID-19 pandemic?
I haven’t provided pre-existing embeddings for pre and post-COVID, so you will need to train your own. I have provided data from city-related subreddits (e.g., r/Minneapolis, r/nyc) from December 2018 and December 2020, which are located in data/other. Each line in the file is one sentence.
See this page for sample code that shows how you can train your own word2vec embeddings using gensim. I recommend setting min_count=50 so that you don’t have too many words in your embeddings.
Training the model is slow, so don’t start this unless you have a decent amount of time left. Parameters like
epochswill affect the time that is taken. It took approximately 3 minutes per corpus when I tested it on colab.
Vocab Filtering
In the prior section, the vocabulary was pre-filtered for you. Here, you should use the vectors_for_all function (see usage in load_emb) to ensure that the vocab overlaps for your 2018 and 2020 embeddings.
New results
After training a model and filtering your vocab, you should be able to apply your code from the main part of the lab to see how word usage has changed following the start of the COVID pandemic.
Bias in embeddings
There is a separate repl.it project to complete this extension!
If you would like to explore how word embeddings reflect human biases as was described in the readings, you can do so using similar methodology to what you did for this assignment. You’ll need to load embeddings, then compute various word similarity scores. Specifically, I have provided data for replicating the occupation-gender association test.
The paper includes a lot of math for statistical tests, but my recommendation is to write the described \(s(w, A, B)\) function from page 3 of the paper, and try to recreate the plot and pearson’s correlation shown in Fig. 1. You can use matplotlib’s scatter function to make a plot.
Data on percentage of women by profession from Caliskan et al. are given in data/professions.json.2
The female and male attribute words used are:
female_attributes = ["female", "woman", "girl", "sister", "she", "her", "hers", "daughter"]
male_attributes = ["male", "man", "boy", "brother", "he", "him", "his", "son"]
GloVe embeddings (an alternative to word2vec that was used in Caliskan’s paper3) can be loaded by calling the load_emb function:
# load glove embeddings
emb = load_emb("glove", prefilter=False)
Credits
Thanks to ziyin-dl on github, who uploaded the yearly pretrained embeddings from google books n-grams. Thanks as well to the authors of the following papers, from which algorithms and data from this assignment were drawn.
- Measuring changes in word embeddings using nearest neighbors: Gonen, H., Jawahar, G., Seddah, D., & Goldberg, Y. (2020). Simple, Interpretable and Stable Method for Detecting Words with Usage Change across Corpora. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 538–555.
- Bias in word embeddings: Caliskan, A., Bryson, J. J., & Narayanan, A. (2017). Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334), 183–186.
- Google books n-grams: Michel, J.-B., Shen, Y. K., Aiden, A. P., Veres, A., Gray, M. K., Google Books Team, … & Aiden, E. L. (2011). Quantitative Analysis of Culture Using Millions of Digitized Books. Science, 331(6014), 176–182.
Thanks to the maintainers of Primer Spec from EECS 485 at the University of Michigan, which is being used to style this webpage.
-
The description of the method has been simplified from Gonen et al., as they define it for an arbitrary number of nearest neighbors. ↩
-
This data is drawn from Caliskan’s supplemental material, but has been pre-processed following the method described in the paper to make it easier to work with. ↩
-
The GloVe embeddings in the data folder are cut off after the first 81K words due to replit’s space constraints. Embeddings for all of the words that you will need are in this set of 81K. ↩