Skip to content
NLP/Playground

Assignment 1 · Questions 1, 5–8

Tweet Geolocator

The task: given a 943-tweet sample labelled with one of ten countries, turn each tweet into a bag of words and predict where it was posted. Below, your text goes through the same preprocessing and the same two trained models, re-implemented in the browser.

Interactive

Where was this tweeted?

Predictions update as you type. Tokens the models never saw in training are ignored, just as sklearn's DictVectorizer ignores them.
68/280

Preprocessing (Question 1)

keptunseen in trainingremoved

  • heading
  • to
  • the
  • footy
  • this
  • arvo
  • ,
  • melbourne
  • weather
  • is
  • great
  • for
  • once
  • !

TweetTokenizer → lowercase → keep tokens containing a letter → drop NLTK stop words → bag of words with 6 distinct tokens.

Loading model weights (≈230 kB)…

Question 8

What each model learned

Naive Bayes favours frequent words in each country's tweets, so “i'm”, “great” and “love” rank highly everywhere. Logistic Regression weights pick out distinctive handles, links and place names instead.

Top 20 features for Australia, ranked by log P(word | country), as printed in the submitted notebook. Handles and links are abbreviated to protect the people behind them.

  1. 1great-4.767
  2. 2i'm-5.168
  3. 3melbourne-5.168
  4. 4one-5.168
  5. 5even-5.452
  6. 6something-5.452
  7. 7keep-5.452
  8. 8good-5.452
  9. 9get-5.452
  10. 10shows-5.849
  11. 11friends-5.849
  12. 12little-5.849
  13. 13like-5.849
  14. 14two-5.849
  15. 15awesome-5.849
  16. 16australian-5.849
  17. 17forward-5.849
  18. 18australia-5.849
  19. 19come-5.849
  20. 20away-5.849

Questions 5–7

How well it worked

A stratified 70/15/15 split gives 660 training, 141 development and 142 test tweets. Hyper-parameters were tuned on the development set only. Ten roughly balanced classes put chance at about 10%, and the largest class is 11% of the data.

Naive Bayes · test accuracy

30.3%

alpha = 0.05 · 95% CI 23.2–38.0%

Naive Bayes · macro-F1

0.303

95% CI 0.220–0.376

Logistic Regression · test accuracy

31.0%

C = 5 · 95% CI 23.2–38.7%

Logistic Regression · macro-F1

0.282

95% CI 0.208–0.346

Naive Bayes on the dev setbest alpha = 0.05 · 28.4%
15%20%25%30%0.050.92.51050100alpha
Logistic Regression on the dev setbest C = 5 · 27.0%
15%20%25%30%0.010.1152075C

Rows: true country · Columns: predicted · 142 test tweets

Confusion matrix for Naive Bayes
True countryaucadegbidmyphsgusza
au5102111211
ca0601101033
de4010001101
gb3104100141
id3000511122
my3001060113
ph4200304110
sg2000021711
us0203021133
za0301024032

F1 by country

  • Australia0.26
  • Canada0.40
  • Germany0.22
  • United Kingdom0.30
  • Indonesia0.38
  • Malaysia0.41
  • Philippines0.28
  • Singapore0.48
  • United States0.18
  • South Africa0.13

Numbers come from re-running the original notebook. The browser port reproduces every test-set prediction and all four scores (parity tests in web/src/lib/geolocator/model.test.ts).

Added in 2026 · analysis around the original results

How sure can we be?

The submitted notebook reported single numbers. With 142 test tweets, each number carries a lot of sampling noise, so here are the same results with 95% intervals, a paired test of the two models and a check of how well their probabilities are calibrated. The point estimates are unchanged. Intervals come from 10,000 bootstrap resamples of the test tweets (seed 90042). The method and its limits are on the methods page.

Test metrics with 95% intervals · 142 tweets · bootstrap, 10,000 resamples, seed 90042

  • Accuracy

    Naive Bayes
    0.303, 95% confidence interval 0.232 to 0.380
    Logistic Regression
    0.310, 95% confidence interval 0.232 to 0.387
    LR − NB (paired)
    +0.007, 95% confidence interval −0.056 to +0.070
  • Macro-F1

    Naive Bayes
    0.303, 95% confidence interval 0.220 to 0.376
    Logistic Regression
    0.282, 95% confidence interval 0.208 to 0.346
    LR − NB (paired)
    −0.022, 95% confidence interval −0.089 to +0.048
  • Calibration error (ECE)

    Naive Bayes
    0.308, 95% confidence interval 0.244 to 0.383
    Logistic Regression
    0.058, 95% confidence interval 0.046 to 0.151
  • Brier score

    Naive Bayes
    0.940, 95% confidence interval 0.838 to 1.042
    Logistic Regression
    0.835, 95% confidence interval 0.785 to 0.884

Wilson intervals for accuracy agree with the bootstrap: Naive Bayes 0.233 to 0.383, Logistic Regression 0.240 to 0.390. Chance is 0.100 and always guessing the largest class scores 0.106. Lower is better for ECE and Brier score. ECE is biased upwards on small samples, which is why Logistic Regression's interval sits mostly above its estimate.

Same tweets, two models

McNemar's test looks only at tweets where the models disagree.

Paired outcomes of the two models on the test tweets
LR rightLR wrong
NB right3211
NB wrong1287
Exact p-value
1.00
χ² (continuity-corrected)
0.00, p 1.00
Odds ratio (NB-only ÷ LR-only)
0.92, 95% confidence interval 0.40 to 2.08
Cohen's g
0.02

23 of 142 tweets separate the models, 11 in favour of Naive Bayes and 12 in favour of Logistic Regression. The test set gives no evidence that either is more accurate, and the accuracy difference could plausibly be anywhere from −5.6 to +7.0 percentage points.

F1 by country, with 95% intervals

  • Naive Bayes
  • Logistic Regression
aun = 15

NB0.26, 95% confidence interval 0.06 to 0.43LR0.29, 95% confidence interval 0.08 to 0.49

can = 15

NB0.40, 95% confidence interval 0.16 to 0.61LR0.16, 95% confidence interval 0.00 to 0.36

den = 8

NB0.22, 95% confidence interval 0.00 to 0.60LR0.00, 95% confidence interval 0.00 to 0.00

gbn = 15

NB0.30, 95% confidence interval 0.07 to 0.52LR0.25, 95% confidence interval 0.00 to 0.48

idn = 15

NB0.38, 95% confidence interval 0.11 to 0.61LR0.45, 95% confidence interval 0.25 to 0.63

myn = 15

NB0.41, 95% confidence interval 0.16 to 0.63LR0.31, 95% confidence interval 0.09 to 0.51

phn = 15

NB0.28, 95% confidence interval 0.07 to 0.48LR0.36, 95% confidence interval 0.10 to 0.61

sgn = 14

NB0.48, 95% confidence interval 0.22 to 0.69LR0.55, 95% confidence interval 0.30 to 0.75

usn = 15

NB0.18, 95% confidence interval 0.00 to 0.36LR0.31, 95% confidence interval 0.09 to 0.52

zan = 15

NB0.13, 95% confidence interval 0.00 to 0.29LR0.12, 95% confidence interval 0.00 to 0.28

Each country has 8 to 15 test tweets, so per-country intervals are wide. Logistic Regression never predicts Germany, so its Germany F1 is 0 in every resample. The full table is in the model card.

Calibration: do the probabilities mean what they say?

Test tweets are grouped by the model's top probability. A calibrated model sits on the diagonal. Points below it are overconfident, and the whiskers are Wilson 95% intervals for each group's accuracy. Dot area shows how many tweets are in the group.

Naive BayesECE 0.308, 95% confidence interval 0.244 to 0.383
000.250.250.50.50.750.7511mean top-class probabilityaccuracyProbability 0.1 to 0.2: 15 tweets, mean confidence 0.106, accuracy 0.000 (Wilson 95% CI 0.00 to 0.20)Probability 0.2 to 0.3: 7 tweets, mean confidence 0.258, accuracy 0.000 (Wilson 95% CI 0.00 to 0.35)Probability 0.3 to 0.4: 14 tweets, mean confidence 0.345, accuracy 0.214 (Wilson 95% CI 0.08 to 0.48)Probability 0.4 to 0.5: 15 tweets, mean confidence 0.445, accuracy 0.333 (Wilson 95% CI 0.15 to 0.58)Probability 0.5 to 0.6: 13 tweets, mean confidence 0.547, accuracy 0.154 (Wilson 95% CI 0.04 to 0.42)Probability 0.6 to 0.7: 21 tweets, mean confidence 0.657, accuracy 0.238 (Wilson 95% CI 0.11 to 0.45)Probability 0.7 to 0.8: 13 tweets, mean confidence 0.758, accuracy 0.385 (Wilson 95% CI 0.18 to 0.64)Probability 0.8 to 0.9: 11 tweets, mean confidence 0.862, accuracy 0.364 (Wilson 95% CI 0.15 to 0.65)Probability 0.9 to 1.0: 33 tweets, mean confidence 0.959, accuracy 0.576 (Wilson 95% CI 0.41 to 0.73)15714151321131133
Logistic RegressionECE 0.058, 95% confidence interval 0.046 to 0.151
000.250.250.50.50.750.7511mean top-class probabilityaccuracyProbability 0.1 to 0.2: 39 tweets, mean confidence 0.155, accuracy 0.231 (Wilson 95% CI 0.13 to 0.38)Probability 0.2 to 0.3: 49 tweets, mean confidence 0.246, accuracy 0.224 (Wilson 95% CI 0.13 to 0.36)Probability 0.3 to 0.4: 24 tweets, mean confidence 0.341, accuracy 0.375 (Wilson 95% CI 0.21 to 0.57)Probability 0.4 to 0.5: 12 tweets, mean confidence 0.445, accuracy 0.417 (Wilson 95% CI 0.19 to 0.68)Probability 0.5 to 0.6: 5 tweets, mean confidence 0.550, accuracy 0.600 (Wilson 95% CI 0.23 to 0.88)Probability 0.6 to 0.7: 8 tweets, mean confidence 0.651, accuracy 0.500 (Wilson 95% CI 0.22 to 0.78)Probability 0.7 to 0.8: 3 tweets, mean confidence 0.763, accuracy 0.333 (Wilson 95% CI 0.06 to 0.79)Probability 0.8 to 0.9: 2 tweets, mean confidence 0.840, accuracy 1.000 (Wilson 95% CI 0.34 to 1.00)394924125832

Naive Bayes is overconfident. It gives 33 test tweets a top probability above 0.9 and gets 19 of them right. Logistic Regression's mean top probability (0.307) is close to its accuracy (0.310).