Switch
Early (0.x) · Apache-2.0 · Free
Snapshot tests for
AI model upgrades.
Before you switch models, for example to a cheaper or newer one, run your real prompts on both, several times. Your current model’s behavior is the baseline, so you write no expected answers. Switch reports only consistent changes in tool calls and marks random variation as flaky.
Quickstart
npx github:inseat-labs/switch init
npx github:inseat-labs/switch compareReal output from a run on 30 September 2026.
$ npx github:inseat-labs/switch compare
Switch claude-cli:sonnet -> claude-cli:haiku 6 cases x 2 samples
? FLAKY search docs arg search_docs.limit (2/2 -> 1/2)
✓ SAME search with limit
✓ SAME schedule with details
✓ SAME schedule missing date
✓ SAME email with recipient
✓ SAME email missing recipient
calls: 24 live, 0 cached, 0 failed
Result: 0 regression, 0 change, 5 same, 1 flaky
What happened: Could we move this assistant from Sonnet to the cheaper Haiku? Switch ran 6 everyday tasks (search docs, schedule a meeting, send an email, including requests missing a date or a recipient) twice on each model. Haiku picked the same tools and asked for the missing details too. One difference showed up only once, so it’s marked flaky, not a regression.
What a regression looks like
Example of a failing result. Illustrative, not from a real run.
✗ REGRESSION email missing recipient acted-instead-of-asking: send_email (0/3 -> 3/3)
In plain words: asked to send an email with no recipient, the new model sent the email without asking who to send it to, in 3 of 3 tries. The old model never did (0 of 3). Because it happens every time, it counts as a regression and the command returns exit code 1, so a CI job fails.
What it checks
Tool droppedThe new model stopped calling a tool the old one called.
Wrong toolIt called a different tool for the same request.
Invented argumentIt filled in an argument the old model left out, such as a detail nobody gave it.
Acted instead of askingSomething was missing, like a date or a recipient, and it went ahead instead of asking.
Broken schemaIts tool call no longer fits the input format the tool declares.
FlakyA difference that doesn’t repeat across tries. Listed separately, never counted as a regression.
Only consistent changes count. Responses are cached, so a re-run reuses earlier answers. You get an HTML and a JSON report, and exit code 1 on a regression.
How it compares
| Tool | What you write | What it focuses on |
|---|---|---|
| promptfoo, DeepEval | Assertions for each test case | Broad evaluation of model answers, with many kinds of checks |
| Switch | No assertions. The old model is the snapshot. | Tool-calling drift between two models, with repeated tries to separate real changes from noise |
Use them together: an eval tool for answer quality, Switch before a model change.
Limitations
- It checks tool-calling behavior, not the quality of free-text answers.
- Every task runs several times on both models, so a comparison makes more model calls than a single test run. Cached answers are reused.
- With claude-cli, tools are described in the prompt as JSON rather than passed as native tools.
- Early (0.x) and not on npm yet. Run it with
npxfrom GitHub.
Help shape it
Issues, real regressions you have seen, and pull requests are welcome.
Switch roadmap ↗Contribute to Switch ↗Maintainer: abenezer88 ↗abenezer@inseat.app ↗
Shipping code with agents?
Fusion: Claude writes, your tests check, Codex reviews. Your branch stays untouched until you apply.
Explore Fusion