EvalShift is a local-first CLI. Runs, scoring, statistics, and reports all happen on your machine under .evalshift/ — no account, no sign-up, nothing leaves your laptop. The only network calls are the model API calls you ask for. EvalShift Cloud exists for teams that want shared runs, diffs, and PR comments — it is strictly opt-in, and everything below works without it.
## 1. Install the CLI
EvalShift is a Python package (Python 3.11+). Install it with uv or pip:
# uv (recommended) uv pip install evalshift # or pip pip install evalshift evalshift --version
That also installs the capture SDK (evalshift-sdk, import name evalshift): the CLI depends on it, so the same environment can instrument your agent. A production agent that only records captures installs the SDK alone.
## 2. Initialise your project
cd your-project evalshift init # minimal, capture-first evalshift.yaml
init writes a minimal, capture-first evalshift.yaml: a passthrough replay prompt, advisory semantic + LLM-judge evaluators, an empty managed suites: block for capture sync to fill, and a migration policy. --provider picks whose model ids the scaffold uses, --profile picks pre-tuned policy budgets, --ci also scaffolds the GitHub Action workflow, and --wire-agents (on by default) points your AI coding agents at the right reference docs.
## 3. Capture real traffic
The golden suite is the crux, so the capture SDK is the recommended way to build one: it records what your agent actually does — model calls, tool calls, outputs — as capture files on disk:
# the CLI install above already brought the SDK along; # a production agent that only records captures installs it alone: uv add evalshift-sdk # decorate the entry point (or wrap the provider client), then run with the gate on: EVALSHIFT_CAPTURE=1 python run_agent.py # -> .evalshift/captures/<suite>/cap_<id>.json
Captures are off unless EVALSHIFT_CAPTURE=1 is set, so the decorator can stay in production code. If the agent calls OpenAI, Anthropic or Google GenAI directly, wrap the client once — wrap_openai(OpenAI()), wrap_anthropic, wrap_genai (SDK 0.4.0+) — and every model call is recorded with no further code. See Framework adapters.
No captures to work from? Write a golden.jsonl by hand — see Golden suite. Hand-written suites are fully supported; they are just more work to keep honest.
## 4. Promote and run
export GEMINI_API_KEY=... # or OPENAI_API_KEY / ANTHROPIC_API_KEY evalshift capture sync # captures -> golden suite + wired config evalshift compare --suite-name <suite> --to <candidate-model> --open
capture sync promotes the recorded captures into a golden suite — inputs, expected tools, expected output, history, each example’s recorded toolset — and wires the managed suites: block. evalshift compare then drives doctor → run → evaluate → analyze → report under one live progress display and opens report.html: a single-file HTML report with per-prompt and per-slice comparisons, effect sizes with 95% CIs, and a migration-policy verdict panel.
run/compare estimate worst-case cost up front and prompt for confirmation above $10 (skip with --yes). defaults.samples_per_example multiplies the call count — and the estimate — by N. Live responses are cached in SQLite, so re-running an identical evaluation is nearly free. Any model LiteLLM supports works — see CLI overview.## Where to next
- +Recommended workflow — SDK capture → CLI evaluation → Action gating, the loop we suggest for real projects.
- +Capture-first example ↗ — the loop above checked in end to end: an instrumented agent, its captures, and the suite
capture syncderived from them. - +CLI overview — the five-stage pipeline, run artefacts, cache, and resume.
- +Golden suite — the example format, expected tools, toolsets, and multi-turn history.
- +Cloud setup — optional: shared runs, baselines, and diffs for teams.
- +GitHub Action — optional: the golden suite as a merge gate on every PR.
