Assignment 1 · Questions 1, 5–8
Tweet Geolocator
The task: given a 943-tweet sample labelled with one of ten countries, turn each tweet into a bag of words and predict where it was posted. Below, your text goes through the same preprocessing and the same two trained models, re-implemented in the browser.
Interactive
Where was this tweeted?
DictVectorizer ignores them.Preprocessing (Question 1)
keptunseen in trainingremoved
- heading
- to
- the
- footy
- this
- arvo
- ,
- melbourne
- weather
- is
- great
- for
- once
- !
TweetTokenizer → lowercase → keep tokens containing a letter → drop NLTK stop words → bag of words with 6 distinct tokens.
Question 8
What each model learned
Top 20 features for Australia, ranked by log P(word | country), as printed in the submitted notebook. Handles and links are abbreviated to protect the people behind them.
- 1great-4.767
- 2i'm-5.168
- 3melbourne-5.168
- 4one-5.168
- 5even-5.452
- 6something-5.452
- 7keep-5.452
- 8good-5.452
- 9get-5.452
- 10shows-5.849
- 11friends-5.849
- 12little-5.849
- 13like-5.849
- 14two-5.849
- 15awesome-5.849
- 16australian-5.849
- 17forward-5.849
- 18australia-5.849
- 19come-5.849
- 20away-5.849
Questions 5–7
How well it worked
Naive Bayes · test accuracy
30.3%
alpha = 0.05 · 95% CI 23.2–38.0%
Naive Bayes · macro-F1
0.303
95% CI 0.220–0.376
Logistic Regression · test accuracy
31.0%
C = 5 · 95% CI 23.2–38.7%
Logistic Regression · macro-F1
0.282
95% CI 0.208–0.346
Rows: true country · Columns: predicted · 142 test tweets
| True country | au | ca | de | gb | id | my | ph | sg | us | za |
|---|---|---|---|---|---|---|---|---|---|---|
| au | 5 | 1 | 0 | 2 | 1 | 1 | 1 | 2 | 1 | 1 |
| ca | 0 | 6 | 0 | 1 | 1 | 0 | 1 | 0 | 3 | 3 |
| de | 4 | 0 | 1 | 0 | 0 | 0 | 1 | 1 | 0 | 1 |
| gb | 3 | 1 | 0 | 4 | 1 | 0 | 0 | 1 | 4 | 1 |
| id | 3 | 0 | 0 | 0 | 5 | 1 | 1 | 1 | 2 | 2 |
| my | 3 | 0 | 0 | 1 | 0 | 6 | 0 | 1 | 1 | 3 |
| ph | 4 | 2 | 0 | 0 | 3 | 0 | 4 | 1 | 1 | 0 |
| sg | 2 | 0 | 0 | 0 | 0 | 2 | 1 | 7 | 1 | 1 |
| us | 0 | 2 | 0 | 3 | 0 | 2 | 1 | 1 | 3 | 3 |
| za | 0 | 3 | 0 | 1 | 0 | 2 | 4 | 0 | 3 | 2 |
F1 by country
- Australia0.26
- Canada0.40
- Germany0.22
- United Kingdom0.30
- Indonesia0.38
- Malaysia0.41
- Philippines0.28
- Singapore0.48
- United States0.18
- South Africa0.13
Numbers come from re-running the original notebook. The browser port reproduces every test-set prediction and all four scores (parity tests in web/src/lib/geolocator/model.test.ts).
Added in 2026 · analysis around the original results
How sure can we be?
Test metrics with 95% intervals · 142 tweets · bootstrap, 10,000 resamples, seed 90042
Accuracy
- Naive Bayes
- 0.303, 95% confidence interval 0.232 to 0.380
- Logistic Regression
- 0.310, 95% confidence interval 0.232 to 0.387
- LR − NB (paired)
- +0.007, 95% confidence interval −0.056 to +0.070
Macro-F1
- Naive Bayes
- 0.303, 95% confidence interval 0.220 to 0.376
- Logistic Regression
- 0.282, 95% confidence interval 0.208 to 0.346
- LR − NB (paired)
- −0.022, 95% confidence interval −0.089 to +0.048
Calibration error (ECE)
- Naive Bayes
- 0.308, 95% confidence interval 0.244 to 0.383
- Logistic Regression
- 0.058, 95% confidence interval 0.046 to 0.151
Brier score
- Naive Bayes
- 0.940, 95% confidence interval 0.838 to 1.042
- Logistic Regression
- 0.835, 95% confidence interval 0.785 to 0.884
Wilson intervals for accuracy agree with the bootstrap: Naive Bayes 0.233 to 0.383, Logistic Regression 0.240 to 0.390. Chance is 0.100 and always guessing the largest class scores 0.106. Lower is better for ECE and Brier score. ECE is biased upwards on small samples, which is why Logistic Regression's interval sits mostly above its estimate.
Same tweets, two models
McNemar's test looks only at tweets where the models disagree.
| LR right | LR wrong | |
|---|---|---|
| NB right | 32 | 11 |
| NB wrong | 12 | 87 |
- Exact p-value
- 1.00
- χ² (continuity-corrected)
- 0.00, p 1.00
- Odds ratio (NB-only ÷ LR-only)
- 0.92, 95% confidence interval 0.40 to 2.08
- Cohen's g
- 0.02
23 of 142 tweets separate the models, 11 in favour of Naive Bayes and 12 in favour of Logistic Regression. The test set gives no evidence that either is more accurate, and the accuracy difference could plausibly be anywhere from −5.6 to +7.0 percentage points.
F1 by country, with 95% intervals
- Naive Bayes
- Logistic Regression
NB0.26, 95% confidence interval 0.06 to 0.43LR0.29, 95% confidence interval 0.08 to 0.49
NB0.26, 95% confidence interval 0.06 to 0.43
LR0.29, 95% confidence interval 0.08 to 0.49
NB0.40, 95% confidence interval 0.16 to 0.61LR0.16, 95% confidence interval 0.00 to 0.36
NB0.40, 95% confidence interval 0.16 to 0.61
LR0.16, 95% confidence interval 0.00 to 0.36
NB0.22, 95% confidence interval 0.00 to 0.60LR0.00, 95% confidence interval 0.00 to 0.00
NB0.22, 95% confidence interval 0.00 to 0.60
LR0.00, 95% confidence interval 0.00 to 0.00
NB0.30, 95% confidence interval 0.07 to 0.52LR0.25, 95% confidence interval 0.00 to 0.48
NB0.30, 95% confidence interval 0.07 to 0.52
LR0.25, 95% confidence interval 0.00 to 0.48
NB0.38, 95% confidence interval 0.11 to 0.61LR0.45, 95% confidence interval 0.25 to 0.63
NB0.38, 95% confidence interval 0.11 to 0.61
LR0.45, 95% confidence interval 0.25 to 0.63
NB0.41, 95% confidence interval 0.16 to 0.63LR0.31, 95% confidence interval 0.09 to 0.51
NB0.41, 95% confidence interval 0.16 to 0.63
LR0.31, 95% confidence interval 0.09 to 0.51
NB0.28, 95% confidence interval 0.07 to 0.48LR0.36, 95% confidence interval 0.10 to 0.61
NB0.28, 95% confidence interval 0.07 to 0.48
LR0.36, 95% confidence interval 0.10 to 0.61
NB0.48, 95% confidence interval 0.22 to 0.69LR0.55, 95% confidence interval 0.30 to 0.75
NB0.48, 95% confidence interval 0.22 to 0.69
LR0.55, 95% confidence interval 0.30 to 0.75
NB0.18, 95% confidence interval 0.00 to 0.36LR0.31, 95% confidence interval 0.09 to 0.52
NB0.18, 95% confidence interval 0.00 to 0.36
LR0.31, 95% confidence interval 0.09 to 0.52
NB0.13, 95% confidence interval 0.00 to 0.29LR0.12, 95% confidence interval 0.00 to 0.28
NB0.13, 95% confidence interval 0.00 to 0.29
LR0.12, 95% confidence interval 0.00 to 0.28
Each country has 8 to 15 test tweets, so per-country intervals are wide. Logistic Regression never predicts Germany, so its Germany F1 is 0 in every resample. The full table is in the model card.
Calibration: do the probabilities mean what they say?
Test tweets are grouped by the model's top probability. A calibrated model sits on the diagonal. Points below it are overconfident, and the whiskers are Wilson 95% intervals for each group's accuracy. Dot area shows how many tweets are in the group.
Naive Bayes is overconfident. It gives 33 test tweets a top probability above 0.9 and gets 19 of them right. Logistic Regression's mean top probability (0.307) is close to its accuracy (0.310).