The SDK records real runs; the evalshift capture command group turns them into golden suite cases that run can score. The two never call each other — the SDK writes capture files, the CLI reads them.
Captures live under <base>/captures/<suite>/ and promoted cases under <base>/suites/<suite>/, where base follows the SDK convention ($EVALSHIFT_DIR, else .evalshift).
## capture list
See what the SDK has recorded: a table of capture id, suite, created-at, tool-call and event counts, and whether each was already promoted. --json emits the same rows plus code_version and input_hash for tooling.
# everything the SDK has recorded (optionally for one suite) evalshift capture list evalshift capture list support_agent # machine-readable (adds code_version + input_hash) evalshift capture list --json
## capture promote
Turn one capture into a golden case. By default the case id is the capture id (--as renames it), and the recorded tool calls become the expected trace — the model’s requested calls when the capture recorded them, the executed ones otherwise (see below). The capture’s recorded toolset (toolset_ref on its first model call) is carried verbatim onto the example — a capture whose first model call recorded no toolset blocks promotion unconditionally, naming the capture id; re-capture with a current SDK. The command prints the suites: block to paste into evalshift.yaml and the exact run invocation.
evalshift capture promote cap_a1b2c3 --as refund_happy_path # it writes the golden case and prints exactly how to wire it: # suites: # support_agent: # source: captured # path: .evalshift/suites/support_agent/golden.jsonl # # then: evalshift run --suite-name support_agent
### Matching strictness
- +
--strict-args— require exact tool-argument matches (default is a looser compare). - +
--names-only— match tool names only; ignore arguments. - +
--tool-count— also pin the expected number of tool calls (under--rounds all, the total over the rounds the replay reaches). - +
--input-var— template-variable name for a bare-string model input (defaultinput). - +
--rounds first(default) —expected_toolsis the capture’s first agent round andrunreplays it single-shot: one call, no tool results fed back.--rounds all— also carries the recorded tool results astool_result_fixtures, andrunreplays every covered round teacher-forced. Under both settingsexpected_toolsisexpected_tool_rounds[0]; nothing is flattened. Details under capture sync. - +
--tag— attach an extra tag (repeatable);--suiterestricts the capture search;--force/-foverwrites an existing case.
### Requested vs executed calls
A capture records three things that are easy to conflate: the tools offered to the model (toolset_ref), the calls the model requested in its response (model_call.requested_tool_calls — written by the SDK 0.4.0 client wrappers, or by passing requested_tool_calls=extract_requested_tool_calls(response) to record_model_call), and the calls the app executed (the tool_call / tool_result events). Promotion prefers requested: executed calls have already passed through the application — its filtering, retries, re-ordering and its own function signatures — so they show what the app did, while a golden case has to state what a model should produce. The promoted case records which yardstick it used as promotion_source: "requested" | "executed".
- +Every
model_callin the capture must carry the field. Pass[]for a round in which the model requested no tools —nullmeans not recorded, not nothing requested. A capture where no call carries it falls back to the executed calls silently; one where only some do falls back as a whole, with a warning, rather than mixing yardsticks. - +On the requested path each
model_callis one round (rounds that requested nothing are dropped) and arguments are carried verbatim — the wrapper unwrapping below only runs on executed calls. - +When requested and executed calls disagree, the requested ones win and promotion warns, naming the tools on both sides.
### What the recorded run cost
Each promoted case file carries cost_usd — the capture’s model_call events summed — and cost_source: recorded when every non-zero part came from your own instrumentation (a recorded cost is never re-estimated), estimated when at least one call recorded tokens but no cost and promotion priced it from litellm’s table for that call’s own model_id. The SDK never prices anything — the provider client wrappers record tokens but leave cost at 0 by design. A model litellm does not price (local, self-hosted) stays at 0.0 with no tag and no warning. The figure is provenance of the capture: it lives on the case file only, and the golden.jsonl example never carries it.
## capture sync
The bulk version, and the one the recommended workflow uses. sync promotes every capture: it groups them by conversation_id, orders them by turn_index, writes one example per turn to .evalshift/suites/<suite>/golden.jsonl, and rewrites the managed suites: block in your evalshift.yaml so run --suite-name just works.
# promote every capture, write the suite, wire the config evalshift capture sync # preview the managed suites: block instead of writing it evalshift capture sync --print # also carry every round's recorded tool results, so run replays the # whole agent loop teacher-forced (default: round 1 only, single-shot) evalshift capture sync --rounds all evalshift run --suite-name support_agent
- +Duplicates are skipped. Two captures with the same replayed content would double that example’s weight in every paired statistic. The check is seeded from the cases already in the suite directory, so it holds across sync runs;
--keep-duplicatesopts out. - +Errored captures are skipped — a turn whose trace carries an error event is not ground truth.
--allow-erroredpromotes them anyway (promoteexits 1 on one instead). - +Agent rounds. Tool calls are grouped into rounds, split at each recorded model call. Every round lands in
expected_tool_rounds, andexpected_toolsis always round 1 (expected_tool_rounds[0]). Under the default--rounds first,runmakes one call per example and feeds no tool result back, so round 1 is the only round scored; a multi-round capture warns, naming the calls it will not replay.--rounds allalso writes the recorded tool results astool_result_fixtures— one inner list per covered round, positionally aligned withexpected_tool_rounds, each entry{tool_name, result, error}, paired with the round’stool_resultevents bycall_idfirst and by tool name within the round second — andrunreplays the example teacher-forced: round k is sent the prompt (plus anyhistory) followed by the recorded rounds 1..k−1 as assistant tool calls andtoolresults, never the candidate’s own calls, for every covered round plus the answer round after it. Coverage stops at the first round with a call that has no recorded result, with a warning naming the rounds the replay will cover; if round 1 itself is uncovered the replay stays single-shot. Scoring and reporting are per round — see Agent rounds. - +Expected text. A
final_outputevent wins; otherwise promotion falls back to the last model call with non-empty string output — the reply the user actually saw, after the tool round-trips. Only the SDK’s LangChain adapter emitsfinal_output, so without that fallback a manually instrumented project promotesexpected: nullevery time. - +Wrapper arguments are unwrapped when the capture’s own recorded toolset schema confirms it — an SDK decorating a Python function records that function’s parameters, so an agent whose tools take one dict would otherwise pin ground truth no model can produce. With no resolvable toolset sidecar the recording is left untouched.
- +CI pin check. After writing,
sync(likeinit) parses.github/workflows/*.ymlforbabaliauskas/evalshift-actionsteps and warns when theevalshift-versionpin is older than the local CLI or absent, printing the exact line to set — advisory only; it never edits a workflow or changes the exit code.
## capture clean
Prune capture files once you’re done with them. By default it removes only promoted captures (the case already lives in the suite); --all removes every capture for the scope. --yes / -y skips the confirm.
# prune only captures you've already promoted (the default) evalshift capture clean --promoted # wipe everything for a suite, promoted or not evalshift capture clean support_agent --all --yes
After deleting captures, clean sweeps <base>/toolsets/: any sidecar referenced by neither a surviving capture nor a promoted suite example is deleted and reported. The refcount spans both captures/ and suites/, so a sidecar a promoted golden.jsonl still uses is never swept — even with --all.
clean only ever deletes capture files and unreferenced toolset sidecars. The golden cases it produced under suites/ are never touched.## capture diff
Compare two captures’ tool-call traces — which tools fired, with what args, in what order. Useful before promoting to confirm a capture is the run you mean to pin.
# compare two captures' tool-call traces evalshift capture diff cap_a1b2c3 cap_d4e5f6
## The end-to-end loop
Capture a real run with the SDK, promote it into a captured suite, then run migrations against it like any other suite — verdicts, baselines, and policy gating all apply. Your golden cases now come from production behavior instead of hand-authored JSONL.
