Assignment 2 · Questions 2–6
Hangman benchmark
Each guesser plays every word in a held-out set of 1,000 Brown word types and keeps going until it solves the word. The score is the average number of wrong guesses. The assignment awarded full marks for the open-ended Q6 player if it averaged under 7.6.
Results
Average mistakes per word (lower is better)
re-run 95% CI 16.62–16.92
re-run 95% CI 9.97–10.49
re-run 95% CI 9.97–10.48
re-run 95% CI 8.42–8.93
re-run 95% CI 7.13–7.62
The submitted notebook kept only the averages, so the submitted figures have no interval. The re-run intervals are percentile bootstrap intervals over its 1,000 per-word results (10,000 resamples, seed 90042). They are wide enough that the 7.6 full-marks threshold sits inside the context model's interval.
Added in 2026 · paired comparison
All five guessers on the same words
1,000 seed-40 test words (sample seed 2024) · 26-mistake budget · bootstrap 10,000 resamples, seed 90042
Q6Context n-gram(reference)
- Mean mistakes [95% CI]
- 7.38, 95% confidence interval 7.14 to 7.63
Q5Bigram
- Mean mistakes [95% CI]
- 8.82, 95% confidence interval 8.58 to 9.06
- Paired difference vs Q6Context n-gram
- +1.44, 95% confidence interval +1.24 to +1.63
- Cohen's dz
- 0.44
- Better / same / worse
- 257 / 132 / 611
Q4Length unigram
- Mean mistakes [95% CI]
- 10.16, 95% confidence interval 9.91 to 10.41
- Paired difference vs Q6Context n-gram
- +2.78, 95% confidence interval +2.55 to +3.00
- Cohen's dz
- 0.77
- Better / same / worse
- 156 / 89 / 755
Q3Unigram
- Mean mistakes [95% CI]
- 10.21, 95% confidence interval 9.95 to 10.46
- Paired difference vs Q6Context n-gram
- +2.83, 95% confidence interval +2.59 to +3.05
- Cohen's dz
- 0.77
- Better / same / worse
- 168 / 74 / 758
Q2Random
- Mean mistakes [95% CI]
- 16.75, 95% confidence interval 16.61 to 16.88
- Paired difference vs Q6Context n-gram
- +9.37, 95% confidence interval +9.12 to +9.61
- Cohen's dz
- 2.33
- Better / same / worse
- 12 / 7 / 981
| Player | Mean mistakes per word [95% CI] | Plot | Paired difference vs Q6Context n-gram | dz | Better / same / worse |
|---|---|---|---|---|---|
| Q6Context n-gram(reference) | 7.38, 95% confidence interval 7.14 to 7.63 | · | · | · | |
| Q5Bigram | 8.82, 95% confidence interval 8.58 to 9.06 | +1.44, 95% confidence interval +1.24 to +1.63 | 0.44 | 257 / 132 / 611 | |
| Q4Length unigram | 10.16, 95% confidence interval 9.91 to 10.41 | +2.78, 95% confidence interval +2.55 to +3.00 | 0.77 | 156 / 89 / 755 | |
| Q3Unigram | 10.21, 95% confidence interval 9.95 to 10.46 | +2.83, 95% confidence interval +2.59 to +3.05 | 0.77 | 168 / 74 / 758 | |
| Q2Random | 16.75, 95% confidence interval 16.61 to 16.88 | +9.37, 95% confidence interval +9.12 to +9.61 | 2.33 | 12 / 7 / 981 |
Positive differences mean more mistakes than the context model. dz is the mean difference divided by the standard deviation of the per-word differences. Better / same / worse counts words where a guesser made fewer, equal or more mistakes than the context model.
Smaller samples give wider intervals. Try 100 words a few times with different seeds to see how much a ranking can move.
Can a large language model beat the n-grams? The LLM evaluation plays the same seeded words with your own API key and reports the same paired statistics.
LLM evaluationDistribution
Mistakes per word, 0 to 26
Verify
Same numbers, computed client-side
set of letters.Re-run it in your browser
5,000 games of the TypeScript port in a Web Worker, on the same two test splits.
| Guesser | Python re-run | Browser (TS port) | Difference | Time |
|---|---|---|---|---|
| Q2Random | 16.771 | — | — | |
| Q3Unigram | 10.226 | — | — | |
| Q4Length unigram | 10.224 | — | — | |
| Q5Bigram | 8.674 | — | — | |
| Q6Context n-gram | 7.371 | — | — |
Every guesser matches Python on every word, not just on average. Several guessers iterate set(ascii_lowercase) - set(guessed), whose order depends on string hashes. With the hash seed pinned, the port replays CPython's set layout from the 26 letter hashes, so ties and the random guesser's picks come out the same.