Skip to content
NLP/Playground

Assignment 2 · Questions 2–6

Hangman benchmark

Each guesser plays every word in a held-out set of 1,000 Brown word types and keeps going until it solves the word. The score is the average number of wrong guesses. The assignment awarded full marks for the open-ended Q6 player if it averaged under 7.6.

Results

Average mistakes per word (lower is better)

Submitted figures come straight from the notebook. The Python re-run executes the same cells with a pinned hash seed (PYTHONHASHSEED=3), so its train/test split, and every number below, can be reproduced.
Q2Random
submitted16.204
re-run16.771

re-run 95% CI 16.62–16.92

Q3Unigram
submitted10.136
re-run10.226

re-run 95% CI 9.97–10.49

Q4Length unigram
submitted10.092
re-run10.224

re-run 95% CI 9.97–10.48

Q5Bigram
submitted8.669
re-run8.674

re-run 95% CI 8.42–8.93

Q6Context n-gram
submitted7.493
re-run7.371

re-run 95% CI 7.13–7.62

The submitted notebook kept only the averages, so the submitted figures have no interval. The re-run intervals are percentile bootstrap intervals over its 1,000 per-word results (10,000 resamples, seed 90042). They are wide enough that the 7.6 full-marks threshold sits inside the context model's interval.

Added in 2026 · paired comparison

All five guessers on the same words

The notebook scored Q2–Q5 on one test split and Q6 on another, so its averages are not a like-for-like comparison. Here every guesser is fitted on the same seed-40 training words, using the notebook's own counting code, and plays the same 1,000 test words. Differences are resampled word by word. On this split the context model makes 1.44 fewer mistakes per word than the bigram model (95% CI 1.24 to 1.63). Why the split was changed is recorded in DR-004.

1,000 seed-40 test words (sample seed 2024) · 26-mistake budget · bootstrap 10,000 resamples, seed 90042

  • Q6Context n-gram(reference)

    Mean mistakes [95% CI]
    7.38, 95% confidence interval 7.14 to 7.63
  • Q5Bigram

    Mean mistakes [95% CI]
    8.82, 95% confidence interval 8.58 to 9.06
    Paired difference vs Q6Context n-gram
    +1.44, 95% confidence interval +1.24 to +1.63
    Cohen's dz
    0.44
    Better / same / worse
    257 / 132 / 611
  • Q4Length unigram

    Mean mistakes [95% CI]
    10.16, 95% confidence interval 9.91 to 10.41
    Paired difference vs Q6Context n-gram
    +2.78, 95% confidence interval +2.55 to +3.00
    Cohen's dz
    0.77
    Better / same / worse
    156 / 89 / 755
  • Q3Unigram

    Mean mistakes [95% CI]
    10.21, 95% confidence interval 9.95 to 10.46
    Paired difference vs Q6Context n-gram
    +2.83, 95% confidence interval +2.59 to +3.05
    Cohen's dz
    0.77
    Better / same / worse
    168 / 74 / 758
  • Q2Random

    Mean mistakes [95% CI]
    16.75, 95% confidence interval 16.61 to 16.88
    Paired difference vs Q6Context n-gram
    +9.37, 95% confidence interval +9.12 to +9.61
    Cohen's dz
    2.33
    Better / same / worse
    12 / 7 / 981

Positive differences mean more mistakes than the context model. dz is the mean difference divided by the standard deviation of the per-word differences. Better / same / worse counts words where a guesser made fewer, equal or more mistakes than the context model.

Draw your own sample

Same guessers and words, a different seeded sample. Runs in a Web Worker.

Smaller samples give wider intervals. Try 100 words a few times with different seeds to see how much a ranking can move.

Can a large language model beat the n-grams? The LLM evaluation plays the same seeded words with your own API key and reports the same paired statistics.

LLM evaluation

Distribution

Mistakes per word, 0 to 26

From the Python re-run. Better models move their mass to the left. The random guesser re-seeds before every pick, so it tries the same fixed letter sequence on every word and needs most of the alphabet.
Random μ = 16.77 [16.62, 16.92]
Unigram μ = 10.23 [9.97, 10.49]
Length unigram μ = 10.22 [9.97, 10.48]
Bigram μ = 8.67 [8.42, 8.93]
Context n-gram μ = 7.37 [7.13, 7.62]

Verify

Same numbers, computed client-side

The TypeScript port replays each guesser on the same test words. It matches the Python re-run word for word, including the order in which Python iterates a set of letters.

Re-run it in your browser

5,000 games of the TypeScript port in a Web Worker, on the same two test splits.

GuesserPythonBrowserDifference
Q2Random16.771——
Q3Unigram10.226——
Q4Length unigram10.224——
Q5Bigram8.674——
Q6Context n-gram7.371——

Every guesser matches Python on every word, not just on average. Several guessers iterate set(ascii_lowercase) - set(guessed), whose order depends on string hashes. With the hash seed pinned, the port replays CPython's set layout from the 26 letter hashes, so ties and the random guesser's picks come out the same.