Skip to content
NLP/Playground
All decision records

Decision record · DR-005

Optional LLM features run only on the visitor's own key, from the browser, with a local audit log

Status
Accepted
Date
2026-10
Applies to
/hangman/llm-eval, /geolocator, /ai-log, web/src/lib/ai

Decision in one line

The LLM features are optional. They call Anthropic or OpenAI directly from the visitor's browser with a key the visitor pastes in, every call is written to an audit log in the visitor's own browser, and the site works fully without a key.

Context

The site is a static Next.js app with no budget for AI calls. I wanted it to show how I evaluate a language model against a classical baseline and how I keep a record of what an AI system was asked, what it answered and what a person decided about it. A shared key held by the site would cost money with every visit and would need abuse protection, rate limiting and secret management.

Decision

  • The visitor chooses Anthropic (default, Claude Haiku 4.5, with Claude Sonnet 5.5 as an option) or OpenAI (model id is free text) and pastes their own key in AI settings.
  • The key is kept in sessionStorage, so it is gone when the tab closes. It moves to localStorage only if the visitor ticks "remember on this device", and "Forget keys" clears both.
  • Requests go from the browser straight to api.anthropic.com or api.openai.com. This site has no server code for AI and never receives the key. A Content-Security-Policy only allows network requests to this site and to those two APIs.
  • Every call is appended to an audit log in IndexedDB on the visitor's device. An entry holds the time, feature, provider, model, the full prompt, the raw and parsed output, latency, token usage and the human decision. Entries are checked for the key before they are stored. The log can be viewed, filtered and exported as JSON or CSV at /ai-log.
  • Every AI output on the site carries an "AI-generated" label.
  • Structured output is requested with a JSON schema and validated again with zod. A reply that does not validate is recorded as invalid, never silently repaired.
  • The LLM Hangman guesser is scored by the same game loop, mistake budget and seeded words as the n-gram guessers, with paired bootstrap intervals against the context guesser.
  • The words a run plays, and the n-gram mistakes on them, are frozen when the run starts. Results and exports are computed from that snapshot only, and a guard refuses to pair a game with a word other than the one planned at its position.
  • Words, not calls, are the independent unit. Invalid-reply and repeat rates are ratio estimators (events over calls) with a cluster bootstrap that resamples words. Words solved has a Wilson interval.
  • A reply cut off at the token limit (max_tokens or length) is scored as truncated and a provider refusal as refused. Each costs one mistake, so every turn moves the game on, but both are counted apart from invalid replies, and every audit entry stores the provider's stop reason. A refusal does not end the run. Claude Haiku 4.5 gets 1,024 output tokens per call; Claude Sonnet 5.5 and OpenAI reasoning models get 4,096, because their thinking counts against the limit.
  • A run stops when the visitor presses Stop or leaves the page, and the visitor's decision on a run is stored on every call that run made, including calls for an unfinished word.

Options considered

  1. A server route with my key. Every visit would cost me money, and the route would need authentication and rate limits.
  2. A server route with the visitor's key. The key would pass through a server I run, which asks visitors to trust me with it.
  3. Browser-direct calls with the visitor's key (chosen).
  4. No AI features. Simplest, but it leaves out the evaluation and audit work the project is meant to show.

Why

Option 3 costs nothing to host, keeps the key between the visitor and their provider, and keeps the audit record on the visitor's device, where they can inspect and export it.

What happened

  • Anthropic only allows browser calls when the request carries the anthropic-dangerous-direct-browser-access header. The name is a fair warning. Any script running on the page, including a browser extension, could read a key held in the page. The site loads no third-party scripts, keeps the key in sessionStorage by default and limits network requests with the Content-Security-Policy, but it cannot protect a key from the visitor's own extensions.
  • The adapters, the validation, the retry logic and the audit log are covered by unit tests with mocked fetch responses. No paid run has been recorded in this repository, because running one needs a key and spends money. The evaluation page reports results only for runs made in the visitor's browser.
  • The harness counts invalid replies and repeated letters as mistakes, the same way the notebook's game loop counts a repeated guess, so a model cannot hide formatting failures.
  • A review before merging found four problems in the first version, all fixed before release. Results were recomputed from the live seed and size inputs, so editing them after a run silently re-paired the LLM with n-gram results on other words, and lowering the size crashed the page. A run kept calling the provider after the visitor navigated away. The invalid and repeat rates had Wilson intervals over calls, which treated correlated calls as independent: with 12 invalid replies all in one word out of 236 calls, Wilson gave 2.9% to 8.7%, where resampling words gives 0% to about 14%. And the cost note assumed 15 calls per word and 12 output tokens per call for every model, which understates a weak model's calls and leaves out Sonnet 5.5's thinking tokens. The note now derives calls from the sample, labels reasoning models' estimates as lower bounds, and shows the running cost.

What I'd change

  • Add a second prompt variant and several runs per word, so prompt sensitivity and run-to-run variation are measured as well as the mean.
  • Retry a truncated reply once with a larger token limit before scoring it, and report how often the retry was needed.
  • Record the provider's request id in the audit entry, so a visitor can match an entry to their provider's usage logs.
  • Offer a hosted demo run with a capped, rate-limited key if the project ever has a budget, recorded in a new decision record that supersedes this one.