// cli · reference
Command reference
every command & flag
Common conventions: -c/--config defaults to ./evalshift.yaml; run artefacts live under .evalshift/runs/; exit code 1 on handled errors.
## Pipeline
| Command | What it does | Flags |
|---|---|---|
| evalshift init | Scaffold a minimal capture-first evalshift.yaml. Without --ci, warns after writing when an existing workflow under .github/workflows/ pins an older evalshift-version than this CLI, or none (CI pin drift; advisory). | -f/--force · -d/--directory · --ci · --wire-agents/--no-wire-agents · --provider gemini|openai|anthropic · --profile |
| evalshift doctor | Environment/config check. An evalshift-sdk row reports which package the evalshift import name resolves to (warn when missing or shadowed). Reports each suite's toolset and flags a suite whose examples carry more than one. Adds a ci pin row when a workflow uses the GitHub Action (warn on pin drift) and a judge family row when an llm_judge judge shares a provider with a configured arm (warn). Exit 1 only on an invalid existing config. | — |
| evalshift run | Paired evaluation run (costs money — calls real models). | -f/--from · -t/--to · -c/--config · -s/--suite · --suite-name · --resume · -y/--yes |
| evalshift evaluate <run-id> | Score all pairs → scores.jsonl. Prints a red “broken eval harness” row when the source model fails the recorded tool ground truth on half or more of the conformance rows. | -c/--config |
| evalshift analyze <run-id> | Paired stats → analysis.json (+ migration_decision.json). | -c/--config · --gate <severities> · --policy-gate |
| evalshift report <run-id> | Render report.html + report.json (+ insights.json — the machine-written run narrative). | -c/--config · --open · --insights/--no-insights |
| evalshift compare | Full pipeline under one live display, over one suite: two models, one comparison. Auto-selects the suite when exactly one is wired; with several, names them and prints a ready-to-run command per suite. Formerly `all`, which stays registered as a hidden alias and keeps working. | every run flag, plus --gate · --policy-gate · --open · --push · --insights/--no-insights |
## Cloud (opt-in)
Nothing leaves your machine unless you run these. See Cloud setup.
| Command | What it does | Flags |
|---|---|---|
| evalshift login | Device-code browser flow, or paste a token. | --token es_... · --host · --no-browser · --timeout |
| evalshift logout | Remove stored credentials. | — |
| evalshift whoami | Show the authenticated identity. | --host · --token |
| evalshift bundle <run-id> | Build the upload artefact (run_bundle.json.gz) without uploading. | -c/--config · -s/--suite · --suite-name · -o/--output · --project |
| evalshift push [run-id] | Upload a run bundle to EvalShift Cloud. Idempotent per project on the local run id; the Cloud run gets its own server-minted id. | --bundle · --project · --host · --token · --create-project/--no-create-project · -c/--config · -s/--suite · --suite-name |
## Captures
| Command | What it does | Flags |
|---|---|---|
| evalshift capture list [suite] | Table of recorded captures. | --json |
| evalshift capture promote <id> | Promote one capture into a golden case. Exits 1 on a capture whose trace recorded an error. | --as · --suite · --input-var · --tag · --strict-args · --names-only · --tool-count · --rounds first|all · --allow-errored · -f/--force |
| evalshift capture sync | Promote ALL captures → suites + wire the managed suites: block in config. Content-duplicate and errored captures are skipped. After writing (or printing the block), warns when a CI workflow pins an older evalshift-version than this CLI (CI pin drift; advisory, exit code unchanged). | --suite · --input-var · --tag · --strict-args · --names-only · --tool-count · --rounds first|all · --allow-errored · -c/--config · -f/--force · --write/--print · --keep-duplicates |
| evalshift capture clean [suite] | Delete promoted capture files + sweep orphaned toolset sidecars. A sidecar a promoted suite still uses is never touched. | --promoted (default) · --all · -y/--yes |
| evalshift capture diff <a> <b> | Compare two capture tool traces. | — |
## Traces, debugging
| Command | What it does | Flags |
|---|---|---|
| evalshift traces import <run-id> | Attach external agent timelines to a run. | --source (required) · --target (required) · --strict |
| evalshift inspect <run-id> [case <example>] | Inspect a run or a single example. | --failed |
| evalshift diff case <run-id> <example> | Side-by-side trace diff when traces exist, text diff otherwise. | — |
| evalshift replay case <run-id> <example> | Re-run one example live. | --model source|target · --trace |
## Housekeeping & hidden
| Command | What it does | Flags |
|---|---|---|
| evalshift runs clean | Prune run history on demand. | --keep · --older-than · --suite · --dry-run · -y/--yes · --config |
| evalshift cache clear | Wipe the response cache. | — |
| evalshift validate | Load config + suite + prompts, cross-check compatibility (hidden). After the success line prints the CI pin drift and judge family warnings, if any (advisory, exit code unchanged). | -s/--suite · -c/--config |
| evalshift test-call | One live smoke-test call (hidden). | -m/--model (required) · -p/--prompt · -t/--temperature · --max-tokens · --tools |
## Environment variables
| Variable | Meaning |
|---|---|
| GEMINI_API_KEY / GOOGLE_API_KEY | Google auth (either works) |
| OPENAI_API_KEY | OpenAI auth (also the default semantic embedding model) |
| ANTHROPIC_API_KEY | Anthropic auth |
| EVALSHIFT_NONINTERACTIVE | Non-empty → skip the cost-confirmation prompt (implied --yes); set in scaffolded CI |
| EVALSHIFT_MAX_RUNS | Override retention.max_runs_per_suite; 0/none/unlimited/off disables count pruning |
| EVALSHIFT_DIR | Base dir for SDK captures the capture commands read (default .evalshift) |
| EVALSHIFT_HOST | Cloud API base URL (default https://api.evalshift.dev) |
| EVALSHIFT_TOKEN | Cloud token (beats the credentials file, loses to --token) |
| EVALSHIFT_CREDENTIALS_PATH | Credentials file override (default ~/.evalshift/credentials) |
| GITHUB_STEP_SUMMARY | When set, analyze appends a markdown results table |
Keys are consumed by LiteLLM at call time; EvalShift itself never stores or transmits them.
