SLAI Natural Language Processing
Lab 2: Generating Movie Plots
In this lab, we are going to write a program that generates text from a prompt, like this fantastic (and completely fake) story about unicorns in the Andes from GPT-3.
What you’ll see is a story that is completely grammatically correct and coherent, but doesn’t actually make sense. The stories that you generate will be slightly grammatically correct and coherent, but make no sense at all, which is part of the fun!
The data that we will use is movie plot summaries, using the CMU movie summary corpus.
Learning goals
Once you complete this lab, you should:
- Be able to use the
np.randomlibrary to make random choices in your program - Understand how markov models can be used to generate text from bigram counts
- Know how to use part-of-speech tagging in the nltk toolkit
Counting n-grams
N-Grams is a term we use for consecutive pairs (or triples, or quads, etc.) of words from a text. So for example, if we have the sentence
I do not like green eggs and ham
Then the 2-Grams (also known as bigrams) of this sentence are I do, do not, not like, like green, green eggs, eggs and, and and ham. Using information about counts of bigrams in a text, we can create a program that will generate text “in the style of” any type of text we already have.
I have provided a function called ngram_transitions that takes a list of strings (in your case, movie summaries) as input and returns a dictionary of dictionaries that store the probability of transitioning from that n-gram to any other n-gram.
It uses the special tokens <s> and </s> to represent the beginning and end of strings.
If we run this code,
texts = ["I am Sam", "Sam I am", "I do not like green eggs and ham"]
bigram_transitions = ngram_transitions(texts, 2)
print("bigrams:")
print(bigram_transitions)
print()
trigram_transitions = ngram_transitions(texts, 3)
print("trigrams:")
print(trigram_transitions)
we get the following output:
bigrams:
{('<s>',): {'I': 0.6666666666666666, 'Sam': 0.3333333333333333}, ('I',): {'am': 0.6666666666666666, 'do': 0.3333333333333333}, ('am',): {'Sam': 0.5, '</s>': 0.5}, ('Sam',): {'</s>': 0.5, 'I': 0.5}, ('do',): {'not': 1.0}, ('not',): {'like': 1.0}, ('like',): {'green': 1.0}, ('green',): {'eggs': 1.0}, ('eggs',): {'and': 1.0}, ('and',): {'ham': 1.0}, ('ham',): {'</s>': 1.0}}
trigrams:
{('<s>', '<s>'): {'I': 0.6666666666666666, 'Sam': 0.3333333333333333}, ('<s>', 'I'): {'am': 0.5, 'do': 0.5}, ('I', 'am'): {'Sam': 0.5, '</s>': 0.5}, ('am', 'Sam'): {'</s>': 1.0}, ('Sam', '</s>'): {'</s>': 1.0}, ('<s>', 'Sam'): {'I': 1.0}, ('Sam', 'I'): {'am': 1.0}, ('am', '</s>'): {'</s>': 1.0}, ('I', 'do'): {'not': 1.0}, ('do', 'not'): {'like': 1.0}, ('not', 'like'): {'green': 1.0}, ('like', 'green'): {'eggs': 1.0}, ('green', 'eggs'): {'and': 1.0}, ('eggs', 'and'): {'ham': 1.0}, ('and', 'ham'): {'</s>': 1.0}, ('ham', '</s>'): {'</s>': 1.0}}
Note: the keys in the dictionary are “tuples”, which you can think of as a fixed-size list that is displayed with () instead of []. To change a list into a tuple, you can use this code:
words = ["hello", "world"]
t_words = tuple(words)
print(t_words) # --> prints ("hello", "world")
print(type(t_words)) # --> prints <class 'tuple'>
print(type(words)) # --> <class 'list'>
The n-gram transitions function has already been called for you, but take a look at the output!
Random text generation
Next, we’ll use the numpy library to help us randomly select some words. The plan will be as follows:
- Your text should start with a
<s>token, which represents the beginning of the text- This is prepared for you in the correct format in the
mainfunction, ascurrent_token_tuple
- This is prepared for you in the correct format in the
- Look at all the potential next tokens, and choose one of them randomly, using the probabilities as weights.
- Add the new token to a list, then use it as a starting word for your next bigram.
- Repeat until you generate a
</s>token, which marks the end of the plot summary.- Use a while loop to generate until you reach a specific token
- Print all of the words in the list - that’s your summary!
You’ll want to generate the next word using the dictionary keys as weights, which will determine the probability of choosing each word. To do so, I recommend using np.random.choice, where p is the probabilies associated with each word. To import numpy, add this to the top of your file:
import numpy as np
Here’s the numpy random demo we reviewed in class, for your reference.
You’ll need an ordered list of the keys and values in the n-gram transitions sub-dictionary associated with your previous word to do this. Here’s one way to generate those lists:
tokens = list(transitions[current_token_tuple].keys())
weights = list(transitions[current_token_tuple].values())
Python will automatically ensure that the indices match up between the lists.
Moving to plot summaries
Once your code is working with green eggs and ham, move on to using movie plot summaries. Load 100 summaries like this:
summaries = load_summaries(100)
If that works, feel free to increase the number of summaries to 42306 to use the full dataset. Caution: this will be slow!
Beyond bigrams
Extend this idea beyond bigrams. Change the value of n when you call the ngram_transitions function to integrate the value of n in generating n-grams. For example, if n is 5, then you should use 5-grams for generating the text. You would randomly select a sequence of 4 tokens from the text, look at all 5-grams that start with those 4 tokens, then choose one at random. Use the 5th token from that 5-grams to be the next token, then repeat.
To generate your first token with 5-grams, your first
current_token_tuplemust be("<s>", "<s>", "<s>", "<s>").
It is possible to write your random selection code so that the value of n does not matter; try to do so, so that you can easily compare the output for different values!
Worksheet question: what are some advantages that you see with 5-grams vs. bigrams? What are some disadvantages?
Extensions
If you have extra time, try improving your program to generate more coherent, grammatical, or interesting plot summaries. These extensions can be implemented in any order!
Using parts of speech
One of the challenges in making the above work well is that your program has no knowledge of grammar. We can sometimes improve the approach by using part-of-speech tagging, often abbreviated as POS tagging. The idea is that you run a tool which tries to identify a grammatical part-of-speech for each token in your file. Once you’ve done this, you use your entire N-gram approach not on the original tokens in the data, but on the parts-of-speech.
In python, you can use nltk.pos_tag to get part-of-speech tags for your text. Look at the documentation to see how it works. You’ll see lots of parts-of-speech that you may not recognize, since it uses many more parts than you’ll learn about in a typical natural language class. If you want to see what any of the tags in particular means, you can do it at the interactive Python prompt. Here’s an example:
>>> import nltk
>>> nltk.help.upenn_tagset('NNP')
NNP: noun, proper, singular
Motown Venneboerger Czestochwa Ranzer Conchita Trumplane Christos
Oceanside Escobar Kreisler Sawyer Cougar Yvette Ervin ODI Darryl CTCA
Shannon A.K.C. Meltex Liverpool ...
So now your goal is to use the exact same N-gram approach that you used above, except that instead of using actual tokens, you should use part-of-speech. This may seem strange, because instead of generating text, you’re just going to generate a sequence of parts of speech, like
PRP VBP NNP NNP PRP VBP DT JJ DT ...
POS to tokens
You will then need to convert this to tokens. To do so, you should construct a dictionary where the keys are part-of-speech. For each key the value is a dictionary consisting of words that appear in the original text that were tagged with that part-of-speech. The dictionary should store probabilities like the dictionary of word transitions.
The purpose of this dictionary is to convert the parts-of-speech that you generated into words. For each part-of-speech, look it up in that dictionary, and then choose one of the words at random, using the values as probabilities as we did above.
You might find it easier to start with only bigrams, then add on 3+ grams!
You’ll then have to decide if this approach works better or worse than the purely word-based approach.
Combining N-Grams and POS tags
See what you can do to combine the two approaches of using n-grams and part-of-speech tags!
Children’s/Family or Horror?
If you look at the movie summaries, you might realize that they come from all different kinds of movies! Something that might make your summaries confusing is that they will use bigrams from both family-friendly films and horror movies, perhaps even in the same sentence.
However, we don’t just have movie summaries, we also have a movie metadata file, which includes information about a movie’s genre, release date, and more. Can you use this file to generate summaries that start with the same word for different genres, and do they differ?
You can read the metadata into a pandas dataframe with this code:
metadata = pd.read_csv("data/movie.metadata.tsv", sep="\t",
header=None,
names=["id", "freebase_id", "name", "release_date", "box_office",
"runtime", "languages", "countries", "genres"])
Then, use the pandas merge function to merge the dataframe with the summary dataframe to map summaries to genres. You may also need to use json.loads to convert the genre string into a dictionary.
Advanced Language Modeling
See page 16 of this book chapter on n-gram language models to learn about backoff and interpolation, and experiment with using these methods in your program.
Credits
Many thanks to Dave Musicant, Andy Exley, and James Ryan, from whom I’ve used some of their ideas on presenting this material.
Thanks to the maintainers of Primer Spec from EECS 485 at the University of Michigan, which is being used to style this webpage.