Skip to content
NLP/Playground
All decision records

Decision record · DR-003

Publish derived parameters, never the tweets

Status
Accepted
Date
2026-10 (revival)
Applies to
/geolocator, web/public/data, web/src/lib/data

Decision in one line

The site serves only numbers derived from the course tweets, with Twitter handles and links replaced by hashes. It never serves tweet text.

Context

The geolocator demo needs the two trained models in the browser. The models were trained on 943 tweets labelled with one of ten countries. The dataset belongs to the COMP90042 teaching team and was provided for the assignment, not for redistribution. The tweets were written by real people and contain @handles and t.co links that point back to them.

Decision

The site publishes the following derived artefacts and nothing else from the dataset.

ArtefactWhat it holds
geolocator-model.jsonThe 3,221-feature vocabulary, Naive Bayes counts per country and Logistic Regression weights rounded to 1e-4. The 620 handle and link features are stored as 64-bit FNV-1a hashes.
geolocator-report.jsonTuning curves, metrics, confusion matrices and the Question 8 top features, with handles shown as @m••••.
geolocator-eval.jsonFor the 142 test tweets, the row position in the data file, the true country, both models' predictions and class probabilities.
segmenter-summary.jsonSeven counts about which MaxMatch direction wins. No hashtags.

Tokens a visitor types are hashed the same way before lookup, so the browser models give the same predictions as the notebook.

Options considered

  1. Ship the raw tweets and train in the browser. Rejected. It republishes other people's posts and the teaching team's dataset.
  2. Ship the trained parameters with a plain-text vocabulary. This was the first version of the revival. It published the handles and links of the people in the dataset.
  3. Ship the parameters with hashed handles and links (chosen). Predictions are unchanged and the handle list cannot be read off the file.
  4. Drop handle and link features. The browser models would no longer match the submitted notebook, which breaks the parity tests that show the port is faithful.
  5. Serve predictions from an API. It needs a server, and the parameters can still be estimated by querying it.

Why

Option 3 was the only one that kept the demo faithful to the submitted models while keeping other people's posts and identities off the site.

What happened

  • All 142 test-set predictions and all four Question 7 scores are reproduced in the browser, and the parity tests pass with the hashed vocabulary.
  • Hashing is pseudonymisation, not anonymisation. Anyone who already suspects a handle can hash it and check whether it is in the vocabulary, which is exactly how typed tokens are matched. The hashes stop someone reading or scraping the list of handles. They do not stop a targeted check of one handle.
  • The plain-text vocabulary still contains rare words that occurred in a single tweet. The per-country counts lose word order and tweet boundaries, so the tweets cannot be rebuilt from them, but a person who already has the dataset could match rare words to tweets.
  • The site does not serve the raw data, but the repository still holds it. The original submission folder includes coursework/as1-data.json so the notebooks can be re-run, and the parity fixtures in web/src/lib/__fixtures__ include the tokenised tweets and their hashtags. The repository is private.
  • Option 2 shipped first, so the history still has it. Commit 305d34f added web/public/data/geolocator-model.json with a plain-text vocabulary of 338 @handles and 282 links, and web/src/lib/data/geolocator-report.json with 42 handle mentions. Commit d8cf950 replaced both, but the old blobs remain in git history, and any Vercel deployment built from an earlier commit served the plain-text vocabulary.

What I'd change

  • Retrain a variant without handle and link features and report how much accuracy they buy, with a paired confidence interval, so the privacy cost has a number next to it.
  • Before the repository is made public, take these out of the working tree and the history, or keep the repository private: the raw data file, the tokenised fixtures (a1-tokens.json and a1-hashtags.json), and the 305d34f versions of geolocator-model.json and geolocator-report.json. Delete any Vercel deployment built before d8cf950 as well.
  • Drop vocabulary entries seen in only one training tweet if a future version is retrained.