Skip to content
NLP/Playground

Showcase · recorded with Playwright

Guided tour

Three short walkthroughs of the main workflows, and a screenshot of every key feature. A Playwright script records them from a production build of this site with fixed inputs, so they can be re-made whenever the site changes. Every video has captions and a written transcript.

Workflows

Watch the main journeys

Each video is muted, with the step shown in a caption bar at the bottom of the frame. The same steps are listed beside it, and the transcript gives the time of each one.

0:42 · 1.3 MB · MP4 Download the video

Walkthrough 1

Segment a hashtag

Type a hashtag, step through forward and reverse MaxMatch one match at a time, and see which split the Brown unigram language model prefers.

  1. 1Open the Hashtag Segmenter. Both MaxMatch directions run in the browser.
  2. 2Type a hashtag: #1stsundayofoctober
  3. 3Step through forward and reverse MaxMatch, one dictionary match at a time
  4. 4Play to the end. Forward MaxMatch ends with ofo, c, tobe and r
  5. 5The unigram LM scores both splits. The higher log-probability wins
  6. 6Per-token Brown counts and log-probabilities explain the verdict
  7. 7Across the notebook's hashtags, win counts come with Wilson 95% intervals
Transcript

Silent screen recording of segment a hashtag on the NLP Playground, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Open the Hashtag Segmenter. Both MaxMatch directions run in the browser.
  2. 0:02Type a hashtag: #1stsundayofoctober
  3. 0:07Step through forward and reverse MaxMatch, one dictionary match at a time
  4. 0:17Play to the end. Forward MaxMatch ends with ofo, c, tobe and r
  5. 0:21The unigram LM scores both splits. The higher log-probability wins
  6. 0:27Per-token Brown counts and log-probabilities explain the verdict
  7. 0:34Across the notebook's hashtags, win counts come with Wilson 95% intervals
Try it yourself

0:54 · 3.2 MB · MP4 Download the video

Walkthrough 2

Geolocate a tweet

Type a tweet, compare the Naive Bayes and Logistic Regression probabilities and the tokens behind them, then check the intervals, the calibration plot and the model card.

  1. 1Open the Tweet Geolocator
  2. 2Type a tweet. It is tokenised and filtered exactly as in Assignment 1
  3. 3Naive Bayes and Logistic Regression each give a probability for ten countries
  4. 4The tokens pushing each model towards its prediction
  5. 5Test accuracy and macro-F1 with bootstrap 95% intervals and a paired McNemar test
  6. 6Reliability diagrams: Naive Bayes is overconfident, Logistic Regression much less so
  7. 7The model card on the methods page reports the same intervals, failure modes and limits
Transcript

Silent screen recording of geolocate a tweet on the NLP Playground, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Open the Tweet Geolocator
  2. 0:02Type a tweet. It is tokenised and filtered exactly as in Assignment 1
  3. 0:12Naive Bayes and Logistic Regression each give a probability for ten countries
  4. 0:17The tokens pushing each model towards its prediction
  5. 0:23Test accuracy and macro-F1 with bootstrap 95% intervals and a paired McNemar test
  6. 0:33Reliability diagrams: Naive Bayes is overconfident, Logistic Regression much less so
  7. 0:38The model card on the methods page reports the same intervals, failure modes and limits
Try it yourself

1:22 · 2.8 MB · MP4 Download the video

Walkthrough 3

Hangman: n-grams vs an LLM

Watch the five n-gram guessers race on one word, then open the LLM evaluation harness, see how a bring-your-own-key works, and run it against a clearly labelled mock.

Mocked AI response for illustration. The LLM part of this recording uses a placeholder key and a mock model (mock-llm-for-illustration) answered by the recording script. No real model was called, so its numbers are not results.

  1. 1Five n-gram guessers race on the same secret word
  2. 2Select a guesser to see the letter scores behind each guess
  3. 3Open the LLM evaluation harness
  4. 4Bring your own key. It stays in this browser and is never sent to this site
  5. 5A placeholder key and a mock model: this run uses a mocked AI response for illustration
  6. 6The mocked LLM plays the same seeded words, under the same rules, one call per guess
  7. 7Mean mistakes with 95% intervals and the paired difference from the context n-gram
  8. 8You decide on the run, and every call is in the AI audit log, exportable as JSON or CSV
Transcript

Silent screen recording of hangman: n-grams vs an llm on the NLP Playground, captioned step by step. A highlighted circle shows the pointer.

  1. 0:00Five n-gram guessers race on the same secret word
  2. 0:17Select a guesser to see the letter scores behind each guess
  3. 0:23Open the LLM evaluation harness
  4. 0:29Bring your own key. It stays in this browser and is never sent to this site
  5. 0:36A placeholder key and a mock model: this run uses a mocked AI response for illustration
  6. 0:45The mocked LLM plays the same seeded words, under the same rules, one call per guess
  7. 1:00Mean mistakes with 95% intervals and the paired difference from the context n-gram
  8. 1:09You decide on the run, and every call is in the AI audit log, exportable as JSON or CSV
Try it yourself

Key features

Screenshots

Desktop screenshots are 1440 × 900. Phone screenshots are 390 × 844, rendered at twice the resolution and scaled down. Select one to see it full size.

On a phone

Reproducible

How these were made

One script

web/e2e/showcase.spec.ts drives Google Chrome through each journey with Playwright and doubles as an end-to-end test. Run it with pnpm showcase, and point BASE_URL at any deployment.

Fixed inputs

The hashtag, tweet and secret word are typed by the script, and the LLM harness plays the first five words of its default seeded sample, so a re-recording shows the same results.

No real key

The AI steps type a placeholder key and a mock model id. The script answers every call to the AI provider inside the browser, so nothing is sent to a provider, and every frame with a mocked reply carries the label "Mocked AI response for illustration".