SLAI Natural Language Processing

Research Project Ideas

For your research project, you will be completing a NLP project with one or two other students. The list below gives some project ideas, but you are welcome to propose another project. These ideas are somewhat vague, so you need to write a short project proposal to make your project idea concrete. Your project should use NLP techniques, ideally building upon what we have learned in class.

Predicting review scores

Websites such as Amazon and Yelp have large collections of reviews; having text linked to star ratings makes these data sources useful for supervised machine learning. Find a dataset from a review site, and build a classifier to predict whether a review is positive or negative. Try to figure out what makes people like different restaurants and products! A recent paper studying how positivity appears in negative reviews is a good place to get some ideas about work in this area!

Sentiment and emotion over time

There are numerous models and lexicons to determine sentiment and emotions from text. Using these algorithms, explore how sentiment has changed over time or in relation to an event. Tweet Emotion Dynamics: Emotion Word Usage in Tweets from US and Canada is a recent paper that you might want to look at to get some ideas on how to proceed.

Fake news detection

There are a number of datasets for identification of fake news and misinformation. Build a classifier to detect fake news, and explore which features are the most effective. After doing so, you may want to try applying your classifier to a new domain, such as identifying COVID-19 misinformation. My lab has released a fake news dataset, which is described in this paper along with algorithms for detecting fake news.

Demographic differences is word associations

Prior work has shown that people with different backgrounds give different answers when asked to perform a word association task (Garimella et al., 2017). My lab has collected a corpus of word similarity scores, tagged with demographic informaition for each annotator. A typical evaluation of word embedding methods has been to calculate correlations with word similarity scores. In this project, you will compare word similarity scores to various word embeddings, and see if correlations differ based on demographics.

This data is not publicly available, and you may not share it outside of SLAI without permission to do so.

Machine translation bias

Machine translation has become very powerful in recent years, but researchers have also begun to explore biases that are created and amplified by these systems. For example, a recent research paper showed that Google Translate often defaults to the male form of adjectives when translating sentences from English to Icelandic such as “I am strong” that mentioned positive personality traits. Come up with a language pair and a methodology to study this phenomenon; it would be helpful to have a background in the language that you are studying, so that you know how gender markers may differ from those that are common in English.

Ongoing and past shared tasks

The NLP community frequently hosts shared tasks, in which researchers from various universities are given access to the same training data and compete to build the best system. These tasks often maintain publicly accessible data after the completion of the competition. Oftentimes, senior researchers spend a long time working on these tasks, so you may not be able to win the competition with a few weeks of work; however, you can still come up with and test out an interesting approach!

Here are a few examples:

If none of these interest you, you can search for more shared tasks on the Corpora email list, or by looking at previous SemEval tasks. Be careful; not all tasks have openly accessible data, and you may not be able to join an ongoing task.

Various data sources