SLAI Natural Language Processing
Research Project Ideas
For your research project, you will be completing a NLP project with one or two other students. The list below gives some project ideas, but you are welcome to propose another project. These ideas are somewhat vague, so you need to write a short project proposal to make your project idea concrete. Your project should use NLP techniques, ideally building upon what we have learned in class.
Predicting review scores
Websites such as Amazon and Yelp have large collections of reviews; having text linked to star ratings makes these data sources useful for supervised machine learning. Find a dataset from a review site, and build a classifier to predict whether a review is positive or negative. Try to figure out what makes people like different restaurants and products! A recent paper studying how positivity appears in negative reviews is a good place to get some ideas about work in this area!
Sentiment and emotion over time
There are numerous models and lexicons to determine sentiment and emotions from text. Using these algorithms, explore how sentiment has changed over time or in relation to an event. Tweet Emotion Dynamics: Emotion Word Usage in Tweets from US and Canada is a recent paper that you might want to look at to get some ideas on how to proceed.
Fake news detection
There are a number of datasets for identification of fake news and misinformation. Build a classifier to detect fake news, and explore which features are the most effective. After doing so, you may want to try applying your classifier to a new domain, such as identifying COVID-19 misinformation. My lab has released a fake news dataset, which is described in this paper along with algorithms for detecting fake news.
Demographic differences is word associations
Prior work has shown that people with different backgrounds give different answers when asked to perform a word association task (Garimella et al., 2017). My lab has collected a corpus of word similarity scores, tagged with demographic informaition for each annotator. A typical evaluation of word embedding methods has been to calculate correlations with word similarity scores. In this project, you will compare word similarity scores to various word embeddings, and see if correlations differ based on demographics.
This data is not publicly available, and you may not share it outside of SLAI without permission to do so.
Machine translation bias
Machine translation has become very powerful in recent years, but researchers have also begun to explore biases that are created and amplified by these systems. For example, a recent research paper showed that Google Translate often defaults to the male form of adjectives when translating sentences from English to Icelandic such as “I am strong” that mentioned positive personality traits. Come up with a language pair and a methodology to study this phenomenon; it would be helpful to have a background in the language that you are studying, so that you know how gender markers may differ from those that are common in English.
Ongoing and past shared tasks
The NLP community frequently hosts shared tasks, in which researchers from various universities are given access to the same training data and compete to build the best system. These tasks often maintain publicly accessible data after the completion of the competition. Oftentimes, senior researchers spend a long time working on these tasks, so you may not be able to win the competition with a few weeks of work; however, you can still come up with and test out an interesting approach!
Here are a few examples:
- HaHackathon: Detecting and Rating Humor and Offense
- SemEval 2021 Task 5: Toxic Spans Detection
- TREC Clinical Trials Track
If none of these interest you, you can search for more shared tasks on the Corpora email list, or by looking at previous SemEval tasks. Be careful; not all tasks have openly accessible data, and you may not be able to join an ongoing task.
Various data sources
- Twitter: there is a lot of work in NLP using data from Twitter, and the Tweepy API makes it fairly easy to retrieve tweets. However, the API has some pretty severe restrictions on how many tweets you are allowed to search.
- Twitter also does not allow researchers to share Tweets directly, they may only share tweet IDs. There are many interesting Twitter datasets, such as this one on COVID stance detection, but you will need to use a service like the Twitter Hydrator to turn IDs into text.
- Reddit: Reddit has become a popular alternative to Twitter in recent years; the third-party pushshift API makes it easy to collect historical data from Reddit.
- Yelp: Yelp has released an easy-to-use dataset for academics, which you can download here.
- Project Gutenberg: free .txt files for out-of-print books
- Wikipedia: did you know that you can download wikipedia?
- A bunch of NLP datasets: here’s a long list of NLP datasets, but note that not all of the links work.