An evaluator scores the (source_output, target_output) pair for one example. Every evaluator returns a PairedScore with both halves and a delta = target_score − source_score. Negative deltas mean the target regressed; positive deltas mean it improved.
EvalShift ships five families:
## Structural (deterministic, free)
Fast checks that depend only on the output shape — no API calls.
- +
json_schema— does the output parse as JSON and validate against a schema you provide? 1.0 / 0.0. - +
regex— does the output match a pattern? 1.0 / 0.0. - +
length— is the output within[min_chars, max_chars]? 1.0 inside, distance-decayed outside.
Use these whenever you can — they’re free, deterministic, and cover most real regressions.
## Semantic (cheap)
- +
semantic.cosine— embed both outputs with a configurable embedding model, compute cosine similarity, then frame the result as a target-preservation score: source = 1.0, target = similarity. Delta < 0 means the target drifted in meaning from the source.
Use when:
- +You don’t have a clean structural check.
- +You want to detect “wandered off” outputs that still look fine syntactically.
Don’t use when:
- +The target is intentionally meant to differ from the source (e.g. you’re migrating from a verbose model to a terse one). The similarity will look low and you’ll get a confusing “regression” signal.
## LLM-as-judge (most expensive)
- +
llm_judge.<criterion>— ask a strong model “which output better satisfies this criterion?” with random A/B ordering to reduce positional bias. Verdict maps cleanly to (source, target) scores.
Use when you can articulate the difference you care about as a sentence (“which output preserves more factual detail?”). Multiple llm_judge entries are allowed — each becomes its own evaluator. Tool-only turns (both outputs empty) are skipped without spending a judge call.
Pick the judge from a third model family. A judge prefers output that reads like its own (self-preference bias), and A/B randomisation does nothing against that. doctor and validate warn when an llm_judge judge_model resolves to the same provider as defaults.source_model or target_model; the report repeats the note above the verdict for judges that actually contributed rows, and report.json carries it as judge_family_overlap. Advisory only, never a failure — init scaffolds a same-provider judge on purpose so a first run needs one API key. defaults.judge_model is exempt: it only seeds the run-insights model, never a pairwise verdict.
## Tool-call evaluators (agent migrations)
For examples whose toolset is non-empty (toolset_ref or inline tools), EvalShift parses each model’s response into a provider-agnostic ToolTrace and scores three orthogonal dimensions:
- +
tool_selection— which tools fire? Two independent axes, one record each.conformancegrades each side absolutely against the example’s ground truth:expected(default;expected_toolsmatched in order),expected_set(order-insensitive multiset recall, for parallel fan-outs),off.divergencegrades the target against the source:set(default; Jaccard on tool names, so reordered identical calls don’t read as drift),exact,first,off. A conformance row both sides missed is taggedTOOL_GROUND_TRUTH_MISS— a broken harness, not a migration finding — and leaves the policy rates. Configureseverity_floor: highso a regression here can never be downgraded. - +
tool_arguments— what did the model pass? Greedy match by(tool_name, sequence_index), then per-field strategies (exact/subset/numeric/semantic). Use when arg drift matters (e.g. the model still callsissue_refundbut the amount is wrong).against: source(default) measures drift from the current model, which pins its own score at 1.0 by construction;against: expectedscores both models against the recorded ground truth instead. - +
tool_trace_structure— how did it sequence them? Sub-scores: call count, parallelism, refusal alignment, expected count. Refusal mismatches forceseverity_floor: high.
Per-round scoring. When an example carries tool_result_fixtures (a suite promoted with --rounds all), run replays it teacher-forced and each of the three scores per round: tool_selection conformance grades round k against expected_tool_rounds[k] (“called nothing” for the answer round) and divergence compares round k of the target to round k of the source; tool_arguments pairs calls within a round, so a right call in the wrong round is a miss; tool_trace_structure counts calls over the whole trace and compares parallelism round by round, recording details.rounds_replayed. Each record is the mean over replayed rounds with the per-round names and scores under metadata.rounds, so max_tool_divergence counts an example as diverged when any round diverged; a round with no ground truth in which neither side called anything does not enter the mean. Single-shot pairs score exactly as before and carry no rounds key. See Agent rounds.
Each agent failure mode maps to one of the three: dropped tool / wrong tool → tool_selection; argument drift → tool_arguments; parallel↔serial flip, loop divergence, refusal regression → tool_trace_structure. Start with tool_selection; turn the others on once you’ve decided you care.
## Imported agent traces
If your agent runs its loop outside EvalShift — LangChain, a custom orchestrator, another language — attach the timelines to a completed run with evalshift traces import and score them with the agent_trace evaluator: tool-order similarity (LCS normalised), per-field argument equality, and missing-verification checks, with verification_tools and dangerous_tools naming the calls that matter. Extra dangerous calls on the target flag DANGEROUS_ACTION_DRIFT; a dangerous call with no preceding verification flags MISSING_VERIFICATION_STEP.
## Blocking vs advisory
Every evaluator takes blocking: bool (default true). Advisory evaluators still score and still appear in reports, but never move the migration verdict — which is what a fresh evalshift init writes, so a new suite cannot fail a build on evaluators nobody has calibrated yet. If every evaluator is advisory the verdict is inconclusive, with the exception of the cost and latency budgets: those read the run’s calls rather than evaluator records, so a conclusive breach still fails.
A turn that ends in tool calls has no text. semantic and llm_judge mark such rows skipped rather than scoring them a 0.5/0.5 tie, and skipped rows are excluded from the paired test, the slice aggregates and the policy metrics. A comparison with nothing left reports as unknown (“Nothing measured”), never as equivalent.
## Mixing evaluators
You can configure several at once. They all run on every (prompt, example) pair, and the analysis layer treats each as a separate comparison (so BH correction adjusts for the multiple-test count correctly).
A typical migration uses:
- +1–2 structural evaluators (cheap baseline checks)
- +1
semantic.cosine(catches semantic drift) - +1
llm_judgeper criterion the team cares about
## Cost considerations
Per (prompt, example) pair, each evaluator means:
| Evaluator | Cost |
|---|---|
structural.* | $0 (no calls) |
semantic | 2 embedding calls |
llm_judge | 1 judge model completion |
tool_selection | $0 (compares parsed traces only) |
tool_trace_structure | $0 (compares parsed traces only) |
agent_trace | $0 (compares imported timelines only) |
tool_arguments | $0 normally; embedding calls per semantic-strategy field if you opt in |
A 100-example suite with 1 prompt and 4 evaluators (2 structural + 1 semantic + 1 judge) is:
- +Run: 200 model calls (100 × 2 models)
- +Evaluate: 200 embedding calls + 100 judge calls
LiteLLM’s pricing data drives the pre-flight estimate; the local SQLite cache absorbs identical re-runs.
Want the statistical machinery behind delta-aggregation? See methodology.
