Skip to content
NLP/Playground

Assignment 2 · Questions 1–7

Hangman vs n-grams

The assignment asked for automatic Hangman players built from character-level language models trained on word types from the Brown corpus. Each player is scored on 1,000 test words held out from its own training split, so it has to generalise to words it has never seen.

Loading n-gram counts and test words…

The players

Five guessers, one idea each

All five ported functions take the masked word and the letters guessed so far and return one letter. Each one adds a single new source of evidence.
  1. Q2

    Random

    Meant to pick a random unguessed letter, but it re-seeds NumPy before every pick, so it tries the same sequence (n, f, z, j, …) on every word, exactly as the original did.

  2. Q3

    Unigram

    Always guesses the most frequent letter it has not tried yet.

  3. Q4

    Length unigram

    Letter frequencies measured only on training words of the same length.

  4. Q5

    Bigram

    Scores each blank by the letter that tends to follow its left neighbour.

  5. Q6

    Context n-gram

    Uses both neighbours when it can (CBOW-style), one side otherwise, unigram as a last resort.

The context model (Q6) was my open-ended answer. When both neighbours of a blank are known it uses counts of “left _ right” triples, which is the CBOW idea applied to characters. With one known neighbour it falls back to a reversed bigram table, and with none it uses unigram frequencies. Probabilities are summed over all blanks.