Skip to content
NLP/Playground

How the numbers were produced

Methods and decisions

Where the data came from, how each method works, how the results were evaluated, what the evaluation assumes, and what I would do differently. The original coursework results are reported as submitted. Everything added in 2026 is analysis around them.

Data provenance

Where the data came from

Two kinds of data feed the site. The course provided a labelled tweet sample, and NLTK provided public corpora. Raw tweets are never published, as the data statement below explains.

Course tweets

943 tweets labelled with one of ten countries, provided by the COMP90042 teaching team in 2024. The site serves only derived parameters, aggregate results and test-set predictions.

NLTK corpora

NLTK's words list, WordNet 3.0 and the Brown corpus (425 hashtags are segmented against them). Hangman uses Brown word types, split into training words and 1,000 test words.

Re-run, not re-implemented

scripts/build_a1_artefacts.py and build_a2_artefacts.py execute the submitted notebooks' cells verbatim with pinned NLTK 3.8.1 and scikit-learn 1.3.2, compare each cell's output with the submission, and export small artefacts. The TypeScript ports are tested against that Python output.

Methods

What each part does

The coursework methods are unchanged. The uncertainty analysis is new, and its helpers live in web/src/lib/stats with unit tests against SciPy, statsmodels, scikit-learn and base R.

Hashtag Segmenter (A1 Q2–Q4)

Greedy MaxMatch from the left and from the right over NLTK's word list, with WordNet lemmatisation. An add-one smoothed Brown unigram model keeps the split with the higher log-probability (DR-001).

Tweet Geolocator (A1 Q1, Q5–Q8)

Tweets are tokenised, lowercased, filtered and turned into bags of words. Multinomial Naive Bayes (alpha 0.05) and Logistic Regression (C 5) were tuned on the development set only. The model card below has the full details.

Hangman guessers (A2 Q2–Q7)

Random, unigram, length-conditioned unigram, left-context bigram, and my context n-gram guesser (DR-002). An optional LLM guesser runs on the visitor's own key (DR-005, DR-006).

Statistics added in 2026

  • Percentile bootstrap intervals for means, macro-F1, per-class F1, ECE and Brier score.
  • Wilson score intervals for proportions of independent items, such as accuracy and words solved.
  • Word-level cluster bootstrap for per-call rates in the LLM harness, because calls within a word are correlated. A rate with no events gets an exact one-sided bound instead, because the bootstrap has no width there.
  • Student-t intervals next to the bootstrap below 30 items, where percentile intervals are too narrow.
  • Paired bootstrap intervals for differences between players on the same items.
  • McNemar's exact and continuity-corrected tests, with the odds ratio and Cohen's g.
  • Cohen's d_z and win, tie and loss counts as effect sizes for paired differences.
  • Reliability diagrams and expected calibration error with ten equal-width bins.

Evaluation design

How results are compared

Every comparison uses held-out items, every interval states its sample size, and every random choice has a fixed, displayed seed.

Held out and paired

The geolocator is scored on 142 test tweets that played no part in training or tuning. Both models see the same tweets, so their difference is tested pair by pair. The Hangman guessers are compared on one common split, with every model trained on the same 39,234 words and playing the same test words (DR-004). The LLM is scored the same way, with invalid replies and repeats counted as mistakes.

Fixed seeds
Assignment 1 splitrandom_state = 1118472The submitted stratified 70/15/15 split.
Assignment 2 splitsPYTHONHASHSEED = 3Pins Python's set order, so the seed-1 and seed-40 splits can be reproduced.
Paired word sampleseed 2024Orders the 1,000 seed-40 test words. The LLM harness plays the first words of that order.
Bootstrap resamplesseed 90042, B = 10,000Every interval on the site unless the page says otherwise.

Assumptions and limitations

What the numbers rely on

Assumptions

  • Test tweets and test words are exchangeable draws from the population the results should describe, which the bootstrap needs.
  • The country labels in the course data are correct. The assignment did not say how they were assigned.
  • Brown word types stand in for English words a Hangman player might face.
  • Words, not calls, are the independent unit of an LLM run. Claude Haiku 4.5 runs at temperature 0, which makes it close to repeatable, though providers do not guarantee this. Claude Sonnet 5.5 and OpenAI reasoning models run with the provider's default sampling and may think first, so repeat runs can differ.

Limitations

  • 142 test tweets give accuracy intervals about ±7 percentage points wide, and 8 to 15 tweets per country make per-country F1 very uncertain.
  • The submitted Assignment 2 split cannot be recovered, because it depended on Python's hash randomisation. The re-run uses a pinned seed.
  • I kept split seed 40 after seeing test scores in 2024, which flatters my guesser, and the common split inherits that.
  • Percentile bootstrap intervals undercover on small samples, such as a 15-word LLM run, and ECE is biased upwards on small test sets.
  • There is no gold segmentation for the hashtags, so the segmenter has win counts, not accuracy.
  • LLM results depend on the model version and the date of the run, and list prices change.

Next time

What I'd change

  • Fix every split seed before looking at a test score, and report results across several seeds.
  • Hand-label about 100 hashtags so the segmenter gets an accuracy with an interval.
  • Retrain the geolocator without handle and link features and measure what they add.
  • Fix the Q6 left-context lookup and use proper smoothing, then test it on unseen words.
  • Use BCa or t-based intervals as a check whenever a sample has fewer than about 30 items.
  • Run each LLM prompt several times per word and add a second prompt, to measure variation.

docs/model-card.md

Model card: Tweet Geolocator

The card for the two classical models behind the geolocator page.

Two bag-of-words classifiers from COMP90042 Assignment 1 (University of Melbourne, 2024 Semester 1) that guess which of ten countries a tweet was posted from. They run in the browser at /geolocator. Every evaluation number below, including the counts under known failure modes, is computed from the exported test-set predictions (most of them at build time by web/src/lib/geolocator/evaluation.ts), and a unit test checks that this card quotes the same numbers.

Model details

PropertyNaive BayesLogistic Regression
TypeMultinomial Naive BayesMultinomial logistic regression, lbfgs, max_iter=10000
Regularisationalpha = 0.05C = 5
Chosen from13 values of alpha on the development set12 values of C on the development set
Libraryscikit-learn (submitted with 1.3.0, re-run with 1.3.2)same
  • Input. One tweet. It is lowercased, split with NLTK's TweetTokenizer, tokens without a letter are dropped, and the 179 words of NLTK's 2024 English stop-word list are removed. The remaining tokens become a bag of word counts.
  • Features. The 3,221 distinct tokens of the 660 training tweets. Tokens never seen in training are ignored. Handles and links stay as features but are published only as hashes.
  • Output. A probability for each of ten countries. The codes are au (Australia), ca (Canada), de (Germany), gb (United Kingdom), id (Indonesia), my (Malaysia), ph (Philippines), sg (Singapore), us (United States) and za (South Africa).
  • Author. Sunchuangyu (Rin) Huang, as individual coursework. The browser port reproduces every test-set prediction of the notebook.

Intended use

Teaching and demonstration. The demo shows how a bag-of-words classifier turns text into a country guess, which tokens push the guess, and how uncertain the evaluation of a small test set is.

Out-of-scope uses

The models are not suitable for real-world location inference. Do not use them to work out where a person is or lives, to profile people, or to inform moderation, policing, advertising or any other decision about a person. They only know ten countries, so a post from anywhere else still receives one of the ten labels. They are right on about three tweets in ten.

Training data

  • Source. 943 tweets provided by the COMP90042 teaching team for the assignment, each labelled with one country. The assignment did not document how the labels were collected or when the tweets were posted.
  • Class balance. 100 tweets each for au, ca, gb, id, my, ph, us and za, 93 for sg and 50 for de.
  • Split. A stratified 70/15/15 split with random_state=1118472 gives 660 training, 141 development and 142 test tweets. Hyperparameters were tuned on the development set only, as the assignment required.
  • What is published. Only derived parameters, never tweet text. See the data statement and DR-003.

Evaluation

The test set has 142 tweets, 15 for most countries, 14 for sg and 8 for de. Chance is 10% and always guessing the largest class scores 10.6%. Intervals are 95% percentile bootstrap intervals over the 142 tweets (10,000 resamples, seed 90042). Accuracy also has a Wilson interval.

MetricNaive BayesLogistic Regression
Accuracy0.303 (0.232 to 0.380)0.310 (0.232 to 0.387)
Accuracy, Wilson interval0.233 to 0.3830.240 to 0.390
Macro-F10.303 (0.220 to 0.376)0.282 (0.208 to 0.346)
Expected calibration error, 10 bins0.308 (0.244 to 0.383)0.058 (0.046 to 0.151)
Brier score (lower is better)0.940 (0.838 to 1.042)0.835 (0.785 to 0.884)
Mean top-class probability0.6110.307

Paired comparison. Both models were scored on the same tweets, so the difference is resampled in pairs. Logistic Regression minus Naive Bayes is +0.007 in accuracy (−0.056 to 0.070) and −0.022 in macro-F1 (−0.089 to 0.048). Only 23 tweets separate the two models: 11 are classified correctly by Naive Bayes alone and 12 by Logistic Regression alone. McNemar's exact test gives p = 1.00, the odds ratio of the discordant pairs is 0.92 (0.40 to 2.08) and Cohen's g is 0.02. The test set gives no evidence that either model is more accurate.

F1 by country. The intervals are bootstrap 95% intervals. With 8 to 15 test tweets per country they are wide.

CountryTest tweetsNaive BayesLogistic Regression
au150.256 (0.065 to 0.432)0.294 (0.077 to 0.488)
ca150.400 (0.160 to 0.611)0.160 (0.000 to 0.364)
de80.222 (0.000 to 0.600)0.000 (0.000 to 0.000)
gb150.296 (0.074 to 0.516)0.250 (0.000 to 0.476)
id150.385 (0.111 to 0.609)0.455 (0.250 to 0.625)
my150.414 (0.160 to 0.625)0.313 (0.087 to 0.512)
ph150.276 (0.069 to 0.483)0.364 (0.095 to 0.606)
sg140.483 (0.222 to 0.690)0.552 (0.296 to 0.750)
us150.176 (0.000 to 0.357)0.313 (0.087 to 0.519)
za150.125 (0.000 to 0.294)0.118 (0.000 to 0.278)

Known failure modes

  • Overconfident Naive Bayes. Naive Bayes gives 33 test tweets a top probability above 0.9, and only 19 of those 33 are right. Its probabilities should be read as scores, not as chances. Logistic Regression's mean top-class probability (0.307) is close to its accuracy (0.310).
  • A favourite class for each model. Naive Bayes predicts Australia for 24 test tweets and only 5 of them are Australian. Logistic Regression predicts Indonesia for 29 test tweets and only 10 of them are Indonesian. The other 19 include 4 tweets each from Germany, the United Kingdom, Malaysia and the Philippines.
  • Germany. Germany has half as much training data as most countries. Logistic Regression never predicts Germany, so its F1 for Germany is 0.
  • Unseen and short tweets. Tokens that never appeared in training are ignored. A tweet made only of unseen tokens is classified from the class prior (Naive Bayes) or the intercepts (Logistic Regression) alone.
  • Memorised accounts. Logistic Regression's strongest features include handles and links of specific accounts (Question 8). They identify people who posted in the training data, not places, and they do not help with anyone else.
  • Neighbouring varieties. English-language tweets from au, ca, gb, us and za share most of their vocabulary, and Malay and Indonesian share many words, so these groups are hard to separate with word counts.

Ethical considerations and privacy

  • The tweets were written by real people. The site publishes no tweet text, and handles and links are hashed in the published vocabulary. Hashing is pseudonymisation, not anonymisation, as DR-003 explains.
  • Guessing location from text is a privacy risk even when a model is weak, and country predictions can act as a proxy for a person's language background. Those uses are out of scope.
  • Text a visitor types stays in the browser, because both models run client-side. It is sent to an AI provider only if the visitor presses the optional LLM comparison button, with their own key, and that call is recorded in the browser's audit log at /ai-log.

Caveats and recommendations

  • With 142 test tweets, accuracy is only known to within about ±7 percentage points.
  • Hyperparameters were chosen on 141 development tweets, where neighbouring settings differ by one or two tweets. The chosen values are not reliably better than their neighbours.
  • The ECE estimate is biased upwards on small samples, and the bootstrap interval inherits that bias. This is why Logistic Regression's interval sits mostly above its point estimate.
  • Report the intervals together with the point estimates whenever these results are quoted.

docs/data-statement.md

Data statement

This statement covers the data behind the NLP Playground: where it came from, what the site publishes, and why the raw tweets are not published.

Datasets used

DatasetUsed bySourcePublished on the site?
943 labelled tweets (as1-data.json)Geolocator, SegmenterCOMP90042 teaching team, 2024 Semester 1No. Only derived parameters and aggregate results.
NLTK words listSegmenterNLTK dataYes, inside the front-coded lexicon
WordNet 3.0 lemmas and exception listsSegmenterNLTK data (Princeton WordNet)Yes, the noun and verb parts needed by the lemmatiser
Brown corpusSegmenter, HangmanNLTK dataOnly counts. Word counts for the Segmenter, letter counts and two 1,000-word test lists for Hangman.
NLTK English stop words (179-word 2024 list)GeolocatorNLTK dataYes, as code

The tweet dataset

  • Who wrote it. Members of the public on Twitter. The tweets contain @handles, t.co links, names and personal news.
  • Who provided it. The COMP90042 teaching team, for Assignment 1. The assignment did not document how the tweets were sampled, when they were posted or how the country labels were assigned.
  • Size and balance. 943 tweets. 100 each for Australia, Canada, the United Kingdom, Indonesia, Malaysia, the Philippines, the United States and South Africa, 93 for Singapore and 50 for Germany.
  • Languages. Mostly English, with some Malay, Indonesian and Tagalog, plus slang and emoji.

Why the raw tweets are not published

  1. They are other people's posts. The authors did not agree to have their tweets republished on a portfolio site, and many tweets name or link to identifiable accounts.
  2. The dataset was provided for the subject. It belongs to the teaching team and was shared for the assignment, not for redistribution.
  3. The demo does not need them. The browser models need only parameters derived from the tweets, and the evaluation needs only predictions and labels.

What the site publishes instead

  • The model vocabulary, with the 620 handle and link features replaced by 64-bit FNV-1a hashes, plus Naive Bayes counts per country and Logistic Regression weights.
  • Question 8's top features per country, with handles abbreviated to the first character (for example @m••••) and links cut to the domain.
  • For the 142 test tweets, the row position in the data file, the true country, both models' predictions and their class probabilities. These feed the confidence intervals, McNemar's test and the calibration plots. They contain no text.
  • Aggregate counts for the Segmenter, such as how often each MaxMatch direction wins.

The reasoning, the options I rejected and the limits of hashing are in DR-003.

Known limits of these protections

  • Hashing is pseudonymisation. Someone who suspects a handle can hash it and check whether it appears in the vocabulary.
  • Rare words in the plain-text vocabulary may have come from a single tweet. They cannot be turned back into tweets, but someone who already has the dataset could match them.
  • The raw data file is still in this private repository, in the original submission folder, so the notebooks can be re-run. The parity-test fixtures in web/src/lib/__fixtures__ include the tokenised tweets and their hashtags. None of these is served by the site.
  • The repository history also keeps the first version of two published artefacts, from commit 305d34f, before handles and links were hashed: web/public/data/geolocator-model.json has a plain-text vocabulary with 338 @handles and 282 links, and web/src/lib/data/geolocator-report.json names handles 42 times. Any Vercel deployment built from a commit before d8cf950 served that vocabulary.
  • Before the repository is made public, all of these must be removed from the working tree and from the history: coursework/as1-data.json, the tokenised fixtures (a1-tokens.json, a1-hashtags.json), and the 305d34f versions of geolocator-model.json and geolocator-report.json. Any old Vercel deployment that served the plain-text vocabulary must be deleted too.

Text typed by visitors

The geolocator runs in the browser, so text typed into it is not sent anywhere by default. If a visitor presses the optional LLM comparison button, their text and the list of ten countries are sent from their browser to the AI provider they chose, with their own key. The call is recorded in their browser's audit log, which they can view and export at /ai-log.

docs/ai-use-statement.md

AI use statement

This statement describes how the NLP Playground uses generative AI. It is informed by the Australian Government's policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is a description of practice, not a claim of compliance or certification.

What AI does on this site

  • LLM Hangman evaluation (/hangman/llm-eval). A language model plays Hangman on the same seeded words as the five n-gram guessers, one letter per call, and is scored with the same game loop and the same statistics.
  • LLM geolocation comparison (/geolocator). A language model guesses which of the ten countries a post typed by the visitor came from, shown next to the two classical models.

Both features are optional. The rest of the site, including every original algorithm, runs without AI and without a key.

What AI never does here

  • It never replaces or changes the original coursework results. The submitted numbers are shown as they were reported, and the analysis around them is computed without AI.
  • It never runs without a visitor's action and a visitor's own API key.
  • It never receives the course tweets or any other data from the dataset.
  • Its outputs are never used to make a decision about a person.

Data sent to the provider

  • Hangman. The masked word pattern (for example _ a _ _ e _), the letters already guessed and the mistake count. The secret word is never sent.
  • Geolocation. The text the visitor typed and the list of ten countries.
  • Never sent. The API key travels only in the request header to the provider the visitor chose. It is never sent to this site, never logged and never committed to the repository.

Requests go directly from the visitor's browser to Anthropic (api.anthropic.com) or OpenAI (api.openai.com). The provider's own terms and retention policy apply to those requests.

Transparency and records

  • Every AI output on the site is labelled "AI-generated".
  • Every call is written to an audit log in the visitor's browser (IndexedDB). An entry records the time, feature, provider, model, the full prompt, the raw and parsed output, latency, token usage and the human decision. For the Hangman evaluation it also records how the reply was scored (hit, miss, repeat, invalid, truncated or refused). The log is at /ai-log, where it can be filtered by feature, decision or run, exported as JSON or CSV, and cleared. If a call cannot be written to the log, the page says so.
  • Replies are requested as JSON matching a schema and validated again before use. A reply that fails validation is recorded and counted as invalid, not silently repaired. A reply cut off at the token limit and a refusal by the provider are recorded with the provider's stop reason and counted separately from invalid replies.
  • A refusal is not retried on another model. The site does not turn on Anthropic's server-side refusal fallbacks, so every result comes from the model the visitor chose. Refusals are counted, and any tokens the provider reports for them are recorded and included in the running cost (DR-006).

Human in the loop

  • The visitor starts every call or evaluation run, after seeing what will be sent, the expected number of calls and tokens and, for Anthropic models, the estimated cost at list price. OpenAI's prices change often, so for OpenAI models the site shows the token estimate and asks the visitor to check the price. The visitor sees the result before anything else happens with it. A run stops when the visitor presses Stop or leaves the page.
  • For each geolocation answer the visitor can accept it, reject it, or record their own country as an edit. For an evaluation run they can accept or reject the whole run, and the decision is stored on every call the run made, including calls for a word that was not finished.
  • Evaluation results are reported with uncertainty and with the number of words played, so a small run cannot be mistaken for a strong result.

How AI was used to build the site

The Assignment 1 notebook (2024) declares that GitHub Copilot and GPT-4 were used mainly to draft function docstrings. The Assignment 2 notebook has no AI declaration. The 2026 revival and this upgrade were built with Claude Code as a coding assistant. Changes go through pull requests, and every port is checked against the notebooks' own output by automated parity tests.

docs/decisions

Decision records

Each record states the decision first, then the options, the reasons, what happened (including the weak numbers) and what I would change. Past records are never edited. A new record supersedes an old one.
  1. DR-001AcceptedPick forward or reverse MaxMatch with a unigram language modelWhen forward and reverse MaxMatch split a hashtag differently, I keep the split with the higher add-one smoothed Brown unigram log-probability.2024-03 (submitted), analysis added 2026-10
  2. DR-002AcceptedA context n-gram Hangman guesser that looks at both neighboursFor each blank I add up a letter distribution chosen by how much context is visible: "left _ right" counts when both neighbours are known, a one-sided table when one is known, and unigram counts when neither is. I guess the highest-scoring letter that has not been tried.2024 Semester 1 (submitted), analysis added 2026-10
  3. DR-003AcceptedPublish derived parameters, never the tweetsThe site serves only numbers derived from the course tweets, with Twitter handles and links replaced by hashes. It never serves tweet text.2026-10 (revival)
  4. DR-004AcceptedCompare the Hangman guessers pair by pair on one common splitFor the uncertainty analysis, every guesser is trained on the same seed-40 training words and plays the same seeded sample of seed-40 test words, and I report paired differences with bootstrap confidence intervals. The submitted averages stay on the site unchanged as the faithful reference.2026-10
  5. DR-005AcceptedOptional LLM features run only on the visitor's own key, from the browser, with a local audit logThe LLM features are optional. They call Anthropic or OpenAI directly from the visitor's browser with a key the visitor pastes in, every call is written to an audit log in the visitor's own browser, and the site works fully without a key.2026-10
  6. DR-006AcceptedCall the providers with plain fetch, and keep refusal fallbacks off so one model is evaluatedThe browser calls Anthropic and OpenAI with plain fetch through two small adapters instead of the providers' SDKs, and the site does not turn on Anthropic's server-side refusal fallbacks, so every result comes from the model the visitor chose and a refusal is counted as a refusal.2026-10