Skip to content
NLP/Playground

Assignment 2 · evaluation harness · bring your own key

LLM vs n-gram Hangman

Can a large language model, asked for one letter at a time, play Hangman better than character n-gram counts from 2024? This page runs the model on the same seeded words as the five notebook guessers and scores it with the same rules and the same paired statistics.

It is optional and uses your own API key, from your browser. The rest of the site works without one.

Run it

Race a model against the n-grams

Choose how many words to play. Calls run one at a time, and every call is written to the AI audit log in your browser.
Loading n-gram models and words…

Protocol

How the comparison is kept fair

The harness changes only the player. Words, rules, budget and statistics are the ones the n-gram guessers are measured with.
  1. 01

    Same words

    A seeded sample of the seed-40 test words (default seed 2024, 15 words). The 15 words are the first 15 of the 1,000-word paired benchmark, so the n-gram numbers line up exactly.

  2. 02

    Same rules

    The notebook's game loop with a 26-mistake budget. A wrong letter, a repeated letter and a reply that is not one letter a–z each cost one mistake. A reply cut off at the token limit or refused by the provider also costs one, so every turn moves the game on, but both are counted separately from invalid guesses.

  3. 03

    One letter per call

    Each turn sends the pattern, the letters tried and the mistake count, and asks for JSON {"letter": …}. The reply is checked against a schema with zod. The secret word is never sent.

  4. 04

    Paired statistics

    Mean mistakes with 95% bootstrap intervals and the paired difference from the context n-gram, plus Student-t intervals below 30 words, where the bootstrap is too narrow. Invalid and repeat rates use a cluster bootstrap over words, because calls within a word are not independent, and an exact one-sided bound when there are no events. Words solved has a Wilson interval. Latency, tokens and cost are reported too.

System prompt

You are playing Hangman. A secret English word (a word type from the Brown corpus) is hidden. Each turn you see the word pattern, with _ for unknown letters, and the letters already guessed. Guess the single most useful next letter. Wrong letters, repeated letters and anything that is not exactly one letter from a to z each cost one mistake. Reply with JSON only, in the form {"letter": "e"}.

Example turn (what is sent)
Pattern: _ a _ _ e _ (6 letters)
Already guessed: a, e, s, t
Mistakes so far: 2 of 26
Next letter?

Expected reply: {"letter": "r"}. Claude Haiku 4.5, the default, runs at temperature 0 with a 1,024-token limit. Claude Sonnet 5.5 and OpenAI reasoning models run with the provider's default sampling and a 4,096-token limit, because their thinking counts against it. Other OpenAI models run at temperature 0.

Limits worth knowing. Language models may have seen Brown-corpus word lists in training, but they never see the secret word here, only the pattern. Fifteen words make a smoke test, not a verdict, and the page says so when a run is that small. The seed-40 split was kept in 2024 after seeing the context guesser's test score (DR-002), so the context guesser, the reference here, is slightly flattered on these words, and a gap in the LLM's favour is if anything understated. The design and its trade-offs are in DR-005 and DR-006.