Decision record · DR-006
Call the providers with plain fetch, and keep refusal fallbacks off so one model is evaluated
- Status
- Accepted
- Date
- 2026-10
- Applies to
- web/src/lib/ai, /hangman/llm-eval, /geolocator, /ai-log
Decision in one line
The browser calls Anthropic and OpenAI with plain fetch through two small adapters instead of the providers' SDKs, and the site does not turn on Anthropic's server-side refusal fallbacks, so every result comes from the model the visitor chose and a refusal is counted as a refusal.
Context
DR-005 put the optional LLM features in the visitor's browser, on the visitor's own key. Two pieces of Anthropic's current guidance for the Claude API pull against how that harness works.
- The guidance is to call Claude through the official SDK (
@anthropic-ai/sdkfor TypeScript) and to use raw HTTP only when there is no SDK or when it is asked for. - For Claude Sonnet 5.5, one of the two models the site offers, it recommends turning on
server-side fallbacks by default (
fallbacks: "default"with theserver-side-fallback-2026-07-01beta header). When a safety classifier declines a request, the API then re-runs it on another model inside the same call.
The Hangman page is an evaluation harness. Its numbers only mean something if every call was answered by the model named in the results, and if the visitor can see exactly what was sent to whom.
Decision
- Provider calls use the browser's
fetchthrough one adapter per provider (web/src/lib/ai/providers.ts). Retries, back-off and error classification live inweb/src/lib/ai/client.ts. Tests inject a mockedfetch, so they check the exact request each provider receives without touching the network. - The site never sends the
fallbacksparameter or the fallback beta header. A request the model declines is scored as refused. It costs one mistake, like any other turn, and is counted apart from invalid replies, with its own rate and interval. - A refusal is kept in the audit log with the serving model, the provider's stop reason, the refusal category when Anthropic gives one, and any token usage the provider reports, which also goes into the running cost.
Options considered
- The Anthropic and OpenAI SDKs in the browser. Typed responses and the providers' own retry logic. They add two dependencies to the page bundle, the Anthropic SDK needs an explicit opt-in for browser use, and the tests would still have to mock the network underneath them.
- Plain
fetchadapters (chosen). About 200 lines that send one JSON request and read one JSON response. The cost is that retries, error types and new API fields are mine to maintain. - Fallbacks on, with each call labelled by the model that served it. A run would then be a mixture of models reported under one name unless every result was split by the serving model, and the comparison with the n-gram guessers would no longer be one model against one baseline. For Sonnet 5.5 the fallback model is Claude Sonnet 5, and only for two refusal categories, so the mixture would also depend on why each request was declined.
- Fallbacks off, refusals scored and counted (chosen).
Why
Option 2 keeps the part of the site that holds a visitor's key small enough to read in one
sitting, and it lets the tests check the wire format the providers actually receive. The
Content-Security-Policy already limits the page to the two API hosts, so plain fetch
gives up no protection that an SDK would add.
Option 4 keeps the thing being measured fixed. A model that declines a Hangman turn has behaved in a way the evaluation should report, not hide by handing the turn to a different model. Counting refusals separately also keeps them from being mistaken for formatting failures.
What happened
- A review before release found that the Anthropic adapter raised the refusal before it read the response's token usage, so refused calls reached the audit log with no tokens and were left out of the running cost. Whether a refusal that arrives before any output is billed depends on its category, so the honest default is to record whatever usage the provider reports. Refusals now carry their usage, serving model and category through to the audit entry, and a unit test covers it for both providers.
- No paid run has been recorded (see DR-005), so the refusal rate on this task is unknown. The prompt holds only a word pattern, the letters tried and a mistake count, so refusals are not expected, but the harness reports them rather than assuming there are none.
- The adapters needed their own handling for things an SDK would have done: classifying
HTTP errors, honouring
retry-after, and telling an aborted request from a network failure. Each has a unit test with a mocked response.
What I'd change
- If the site ever evaluates a fallback chain as a product, add it as a separate, clearly labelled configuration, and use the response's per-attempt usage to attribute every call to the model that served it.
- Store the refusal category in its own audit field instead of inside the error message, so refusals can be filtered and counted by category.
- Revisit the SDK if the adapters grow to cover streaming or tool use, where the SDK's helpers would save real code.