Decision record · DR-004
Compare the Hangman guessers pair by pair on one common split
- Status
- Accepted
- Date
- 2026-10
- Applies to
- /hangman/benchmark, /hangman/llm-eval, scripts/build_a2_artefacts.py
Decision in one line
For the uncertainty analysis, every guesser is trained on the same seed-40 training words and plays the same seeded sample of seed-40 test words, and I report paired differences with bootstrap confidence intervals. The submitted averages stay on the site unchanged as the faithful reference.
Context
The notebook scored the random, unigram, length-unigram and bigram guessers on the test words of a seed-1 split, and my context guesser on the test words of a seed-40 split, each trained on its own split's training words. Comparing averages from two different sets of words mixes two effects, the guesser and the difficulty of the words it happened to get. The averages also came with no measure of uncertainty.
Decision
scripts/build_a2_artefacts.py now refits the Question 3 to 5 count tables on the seed-40
training words with the notebook's own code (the same Counter loops and
cal_bigram_counts), and passes those tables to the notebook's guesser functions as keyword
arguments. The context guesser's unigram fallback uses the same refit table, which also
removes a train/test leak in the submitted fallback (DR-002). All five
guessers then play the 1,000 seed-40 test words with the 26-mistake budget. The site
reports three things from those games.
- The mean mistakes per guesser, with a 95% percentile bootstrap CI from 10,000 resamples drawn with seed 90042.
- Each guesser's paired difference from the context guesser. Words are resampled in pairs, and Cohen's d_z and the number of words each guesser did better, the same or worse on are shown next to the interval.
- A control to redraw a smaller seeded sample of words (default seed 2024) and re-run everything in a Web Worker.
Options considered
- Leave the unpaired averages. Faithful, but it cannot say whether a gap is larger than the noise from which words were drawn.
- Play every guesser on the union of both test sets. Most seed-1 test words are in the seed-40 training words and the other way round, so every model would be scored partly on words it was trained on.
- One common split, paired resampling (chosen).
- Many split seeds with paired resampling within each. Better, because it also shows how much the split itself moves the results. The split depends on Python's string-hash seed, so each new split needs a full re-run of the notebook. Left for later.
Why
Pairing removes the word-difficulty variation that every guesser shares. The standard deviation of the per-word difference between the bigram and context guessers is about 3.3 mistakes, so a paired interval on 1,000 words is about ±0.2, which is enough to rank the guessers with confidence. Bootstrap percentile intervals need no assumption about the shape of the mistake distribution, which is skewed and bounded at 0 and 26.
What happened
- The TypeScript port reproduces the Python per-word mistakes of all five guessers on all 1,000 common-split words, so the browser re-run and the build-time numbers agree exactly.
- Common-split means (95% CI): random 16.748 (16.612 to 16.883), unigram 10.209 (9.951 to 10.459), length unigram 10.159 (9.907 to 10.407), bigram 8.820 (8.577 to 9.063), context n-gram 7.381 (7.140 to 7.628).
- Every guesser makes more mistakes than the context guesser, and no paired interval includes zero. The smallest gap is the bigram guesser at 1.44 (1.24 to 1.63).
- Refitting moved the averages by up to 0.15 mistakes against the per-split re-runs (bigram 8.674 on its own split, 8.820 here). That is the size of effect that a change of test words produces, which is why the unpaired comparison was not reliable at the second decimal place.
- The common split still uses seed 40, the seed I kept after looking at test scores in 2024 (see DR-002), so it does not remove that selection effect.
What I'd change
- Repeat the analysis over several split seeds and report the spread between splits next to the bootstrap intervals.
- Use BCa intervals, or a t-interval on the paired differences as a check, when the sample is small (the LLM harness defaults to 15 words, where percentile intervals are too narrow).