Assignment 1 · Questions 2–4
Hashtag Segmenter
Hashtags glue words together with no spaces. This demo re-runs the MaxMatch tokeniser I wrote for the assignment in both directions and lets a Brown-corpus unigram language model pick the more plausible split.
Segment a hashtag
Lowercased, then matched against NLTK's word list with WordNet lemmas and scored with Brown counts (263k lexicon entries). Up to 40 characters.
Method
How it works
- 01
Greedy longest match
From the current position, try every fragment up to the end of the hashtag and keep the longest one that is in NLTK's word list. If nothing matches, emit a single character and move on.
- 02
Inflections via WordNet
The word list only has lemmas, so a fragment also counts if its WordNet lemma is listed (verb first, then noun). That is how “images” matches through “image”.
- 03
Two directions
Reversed MaxMatch does the same thing from the right-hand end. The two greedy passes often disagree, e.g. all·is·well versus al·li·swell.
- 04
Let a language model decide
Each segmentation is scored with a unigram model over the lowercased Brown corpus, using add-one smoothing in log space. The less negative score wins.
log P(w₁…wₙ) = Σ log ( (countBrown(wᵢ) + 1) / (N + V) )
N = 1,161,192 tokens · V = 49,815 types
Added in 2026 · Question 4 in numbers
Which direction wins, and why
Hashtags where the directions disagree
124 of 425
1 exact tie
Reverse MaxMatch wins
67 of 123
54.5% · Wilson 95% CI 45.7%–63.0%
The split with fewer tokens wins
58 of 59
98.3% · Wilson 95% CI 91.0%–99.7%
The reverse interval includes 50%, so these 123 hashtags cannot tell which direction is better in general. The model nearly always prefers fewer tokens, because every token adds a negative log-probability, from about −2.85 for “the” to about −14.0 for a token Brown has never seen. The decision and its weak spots are written up in DR-001.