Assignment 2 · Questions 1–7
Hangman vs n-grams
The assignment asked for automatic Hangman players built from character-level language models trained on word types from the Brown corpus. Each player is scored on 1,000 test words held out from its own training split, so it has to generalise to words it has never seen.
The players
Five guessers, one idea each
Q2
Random
Meant to pick a random unguessed letter, but it re-seeds NumPy before every pick, so it tries the same sequence (n, f, z, j, …) on every word, exactly as the original did.
Q3
Unigram
Always guesses the most frequent letter it has not tried yet.
Q4
Length unigram
Letter frequencies measured only on training words of the same length.
Q5
Bigram
Scores each blank by the letter that tends to follow its left neighbour.
Q6
Context n-gram
Uses both neighbours when it can (CBOW-style), one side otherwise, unigram as a last resort.
The context model (Q6) was my open-ended answer. When both neighbours of a blank are known it uses counts of “left _ right” triples, which is the CBOW idea applied to characters. With one known neighbour it falls back to a reversed bigram table, and with none it uses unigram frequencies. Probabilities are summed over all blanks.