Assignment 1 · 9% of the subject
Preprocessing and text classification
- 01Tokenise 943 tweets with NLTK's TweetTokenizer, lowercase them, drop tokens without letters and stop words, and build bags of words.
- 02Split hashtags into words with MaxMatch over NLTK's word list, using WordNet lemmas to match inflected forms.
- 03Run MaxMatch again from right to left and score both outputs with an add-one smoothed Brown unigram model.
- 04Make a stratified 70/15/15 split, then tune Naive Bayes and Logistic Regression on the development set only.
- 05Compare both models on the test set (accuracy and macro-F1) and inspect the top features for each country.