SLAI Natural Language Processing

Lab 3: Language Identification

In this lab, you will use text classification algorithms to identify the language of sentence. For language identification, we will use the Naive Bayes machine learning algorithm with character n-gram features.

You will start with classifying English and Spanish, and if you have time, you will create a model that distinguishes between 8 languages!

Learning goals

Once you complete this lab, you should:

Provided functions

Functions are provided for loading data and for creating n-grams, normalizing, and performing argmax.

Examples of how to use the functions get_char_ngrams, normalize and argmax are provided in the examples function on replit.

get_char_ngrams

You used ngrams in the last lab, so you should be familiar with them. For this lab, we will work with character ngrams, which are ngrams made up of individual characters instead of words.

The inputs to get_char_ngrams are a string string and an integer n, which represents the number of characters that should be in each ngram. The function returns a list of strings.

Here are a few examples:

string = "SLAI CS"
print(get_char_ngrams(string, 2)) # --> prints ["SL", "LA", "AI", "I ", " C", "CS"]
print(get_char_ngrams(string, 3)) # --> prints ["SLA", "LAI", "AI ", "I C", " CS"]

normalize

This function takes a dictionary of counts and normalizes them by dividing each value by the total counts. It returns a dictionary of the normalized counts, which are probabilities. It does not modify the original dictionary.

If the input is a dictionary of dictionaries, it normalizes each sub-dictionary.

Here’s an example:

counts = {"spa": 50, "eng": 150}
probs = normalize(counts)
print(probs) # --> prints {'spa': 0.25, 'eng': 0.75}

argmax

The argmax function takes a dictionary with numbers as values, and returns the key with the highest value.

Here’s an example:

probs = {"spa": 0.7, "eng": 0.3}
highest_key = argmax(probs)
print(highest_key) # --> prints "spa"

Predict labels

We will start by replicating the example of predicting the language of able from the worksheet. The priors and likelihoods have been computed for you, and are available to use in main.

To predict the label of a new sentence \(s\), you’ll need to collect all of the character bigrams from \(s\). We will call these the features, \(f\), and compute the probability of the language, \(l\). Then, you’ll need to compute the following product:

\[P(l|s) = P(l) \prod_{f\in s} P(f | l)\]

Remember, the \(\prod\) symbol means that you will multiply all of the probabilities together.
For example, \(\prod_{i=1}^3 i = 1 * 2 * 3 = 6\).

You will want to compute this product for each language, and store the probabilities that a sentence is written in each of the languages.

Predicting label will require a nested for-loop structure. In the outer loop, you’ll loop through the possible languages. Then, your inner loop will loop through all of the character n-grams.

You might see KeyError if a character n-gram isn’t present in your likelihood dictionary. To avoid this, check if the key is in your likelihoods. If it’s not, multiply by 0 if it’s in the vocab, and 1 otherwise. The following example will need to be modified, but shows how to handle a KeyError:

if ngram in likelihoods[lang]:
    prob *= likelihoods[lang][ngram]
elif ngram in vocab:
    prob *= 0
else:
    prob *= 1

Once you have all of the probabilities, you should use either an if statement or your argmax function to select the most probable language.

Make sure that the final probability matches the probability from your worksheet before continuing

Test on more data

Testing on a single sentence probably doesn’t tell us much. Use the load_test function to load 2000 samples from the test data:

from util import load_test
test_sentences, test_langs = load_test(avg_samples_per_language=100)

Note that the sentence test_sentences[i] has the language test_langs[i]. Then, try to predict the language of all of these sentences using your code.

Hint: you’ll need to add a third loop that loops through sentences, surrounding your loop that goes through the two languages.

How good are your results?

The most straightforward method for seeing how good you model is is to compute the accuracy. Given a list of predictions and a list of true labels for the same sentences, you will to determine how many indices have the same language. Then, divide that value by the number of sentences in the test set to get the accuracy score.

Print the accuracy score that you get when testing on the test sentences.

Train a better language identification model

Your results probably aren’t very good because your model only knows about three words, “blanco”, “ablaze”, and “hablo”, which clearly don’t cover the entirety of the English or Spanish languages.

Load a large sample of training data and see if it improves results.

from util import load_train
train_sentences, train_langs = load_train(avg_samples_per_language=1000)

The calculation of priors and likelihoods should remain the same!

Extensions

If you have extra time, think about ways in which you might be able to improve your language identification algorithm. You should start by implementing the multilingual version, then you may proceed in any order.

Multilingual version

Change your data loading as follows:

train_sentences, train_langs = load_train_data(avg_samples_per_language=1000, binary=False)
test_sentences, test_langs = load_test_data(avg_samples_per_language=100, binary=False)

This will mean that you are now loading data in eight different languages instead of two. Now, update your code to work with 8 languges. You will probably want to:

  1. Store your likelihoods in a dictioanry of dictionaries, if you aren’t already
  2. Use argmax to compute the most probable language, if you aren’t already

Confusion Matrix

A confusion matrix is a n x n matrix, where n is the number of labels in your dataset. The columns represent the actual labels, and the rows represent the predicted labels. Let’s look at an example:

y_true = ["fra", "fra", "ger", "eng", "ita", "eng"]
y_pred = ["fra", "eng", "ger", "eng", "eng", "eng"]

For this data, we would have the following confusion matrix:

  fra ger eng ita
fra 1 0 0 0
ger 0 1 0 0
eng 1 0 2 1
ita 0 0 0 0

As you can see, the number at row n and column m represents the number of times that a sentence written in the language from column m was predicted to be written in langue n.

Now, create your own confusion matrix in python. One good way to do so is using a 2D numpy array storing integers. Here’s an example of using the array - you will need to add a loop, and a way to connect each language to a row and a column.

>>> import numpy as np
>>> counts = np.zeros((3, 3), dtype=np.int32)
>>> print(counts)
[[0 0 0]
 [0 0 0]
 [0 0 0]]
>>> counts[1, 2] += 1
>>> print(counts)
[[0 0 0]
 [0 0 1]
 [0 0 0]]

This should help you to gain a more thourough understanding of the strengths and weaknesses of your model, compared to the accuracy score.

Add-one smoothing

In your classifier, you set the probability to zero if an ngram is not found in your likelihood dictionary. However, this means that even if there is lots of evidence that a text is in English, such as “I like to eat sjkdls”, you will say that the probability that it is in English is 0% after processing “sjkdls”, if one character n-gram appears in Spanish!

A better thing to do is to use Laplace (add-one) smoothing, which you will do after you have collected word counts but before you normalize them. First, create a set containing all of the character bigrams in the vocabulary. Then, loop through your count dictionaries to add one to each value. Finally, normalize!

How about character bigrams that aren’t in the vocabulary at all, for any language? Then, there presence shouldn’t affect your prediction, so you can safely ignore them entirely.

Preprocessing your text

If you followed the instructions word-for-word, you won’t have done much text preprocessing. Can you add preprocessing steps to improve your predictions, such as converting the text to lowercase?

N-gram size

Is character bigrams the best choice of n-gram size, or would trigrams work better? Try out different values from 1 to 5 to see what works best.

As you do this, keep an eye on your confusion matrix in addition to your accuracy scores. If you see any interesting patterns, it is worth considering a) If and why they are stronger for some languages than others b) If they could be fixed by the other extensions, if you haven’t yet implemented them.

Credits

Thanks to Winston Wu for sharing the material from his Multilingual NLP class, from which I got the idea for this lab.

Thanks to the maintainers of Primer Spec from EECS 485 at the University of Michigan, which is being used to style this webpage.