Skip to content
NLP/Playground

Assignment 1 · Questions 2–4

Hashtag Segmenter

Hashtags glue words together with no spaces. This demo re-runs the MaxMatch tokeniser I wrote for the assignment in both directions and lets a Brown-corpus unigram language model pick the more plausible split.

Segment a hashtag

Lowercased, then matched against NLTK's word list with WordNet lemmas and scored with Brown counts (263k lexicon entries). Up to 40 characters.

Loading the word list, WordNet morphology and Brown counts (≈600 kB)…

Method

How it works

The TypeScript port is checked against the notebook. All 425 hashtags in the dataset get the same forward and reverse segmentations, and log-probabilities match to 9 decimal places.
  1. 01

    Greedy longest match

    From the current position, try every fragment up to the end of the hashtag and keep the longest one that is in NLTK's word list. If nothing matches, emit a single character and move on.

  2. 02

    Inflections via WordNet

    The word list only has lemmas, so a fragment also counts if its WordNet lemma is listed (verb first, then noun). That is how “images” matches through “image”.

  3. 03

    Two directions

    Reversed MaxMatch does the same thing from the right-hand end. The two greedy passes often disagree, e.g. all·is·well versus al·li·swell.

  4. 04

    Let a language model decide

    Each segmentation is scored with a unigram model over the lowercased Brown corpus, using add-one smoothing in log space. The less negative score wins.

log P(w₁…wₙ) = Σ log ( (countBrown(wᵢ) + 1) / (N + V) )

N = 1,161,192 tokens · V = 49,815 types

Added in 2026 · Question 4 in numbers

Which direction wins, and why

The notebook printed both scores for every hashtag where the directions disagree. Counting them shows what the unigram model actually rewards. There is no gold segmentation, so these are win counts, not accuracy.

Hashtags where the directions disagree

124 of 425

1 exact tie

Reverse MaxMatch wins

67 of 123

54.5% · Wilson 95% CI 45.7%–63.0%

The split with fewer tokens wins

58 of 59

98.3% · Wilson 95% CI 91.0%–99.7%

The reverse interval includes 50%, so these 123 hashtags cannot tell which direction is better in general. The model nearly always prefers fewer tokens, because every token adds a negative log-probability, from about −2.85 for “the” to about −14.0 for a token Brown has never seen. The decision and its weak spots are written up in DR-001.