Switch

Early (0.x) · Apache-2.0 · Free

Snapshot tests for
AI model upgrades.

Before you switch models, for example to a cheaper or newer one, run your real prompts on both, several times. Your current model’s behavior is the baseline, so you write no expected answers. Switch reports only consistent changes in tool calls and marks random variation as flaky.

Quickstart

npx github:inseat-labs/switch init
npx github:inseat-labs/switch compare

Real output from a run on 30 September 2026.

$ npx github:inseat-labs/switch compare
Switch  claude-cli:sonnet -> claude-cli:haiku   6 cases x 2 samples
? FLAKY       search docs              arg search_docs.limit (2/2 -> 1/2)
✓ SAME        search with limit
✓ SAME        schedule with details
✓ SAME        schedule missing date
✓ SAME        email with recipient
✓ SAME        email missing recipient
calls: 24 live, 0 cached, 0 failed
Result: 0 regression, 0 change, 5 same, 1 flaky

What happened: Could we move this assistant from Sonnet to the cheaper Haiku? Switch ran 6 everyday tasks (search docs, schedule a meeting, send an email, including requests missing a date or a recipient) twice on each model. Haiku picked the same tools and asked for the missing details too. One difference showed up only once, so it’s marked flaky, not a regression.

What a regression looks like

Example of a failing result. Illustrative, not from a real run.

✗ REGRESSION  email missing recipient   acted-instead-of-asking: send_email (0/3 -> 3/3)

In plain words: asked to send an email with no recipient, the new model sent the email without asking who to send it to, in 3 of 3 tries. The old model never did (0 of 3). Because it happens every time, it counts as a regression and the command returns exit code 1, so a CI job fails.

What it checks

Tool droppedThe new model stopped calling a tool the old one called.

Wrong toolIt called a different tool for the same request.

Invented argumentIt filled in an argument the old model left out, such as a detail nobody gave it.

Acted instead of askingSomething was missing, like a date or a recipient, and it went ahead instead of asking.

Broken schemaIts tool call no longer fits the input format the tool declares.

FlakyA difference that doesn’t repeat across tries. Listed separately, never counted as a regression.

Only consistent changes count. Responses are cached, so a re-run reuses earlier answers. You get an HTML and a JSON report, and exit code 1 on a regression.

How it compares

Switch next to general eval tools
ToolWhat you writeWhat it focuses on
promptfoo, DeepEvalAssertions for each test caseBroad evaluation of model answers, with many kinds of checks
SwitchNo assertions. The old model is the snapshot.Tool-calling drift between two models, with repeated tries to separate real changes from noise

Use them together: an eval tool for answer quality, Switch before a model change.

Limitations

Help shape it

Issues, real regressions you have seen, and pull requests are welcome.