A migration policy is the set of regression budgets a run is judged against. It is written once, as the migration_policy block of your evalshift.yaml, and it travels with every run you push — so the local compare --policy-gate, the run’s Policy tab, the Pull Requests view, and the GitHub Action’s commit status are all reading the same numbers. The web app displays that policy; it does not edit it.
thresholds block in the same file. Thresholds are free-form, gate nothing, and are overwritten in the project by evalshift push; the policy is a fixed set of budgets and is what decides whether a run blocks. Settings shows your config thresholds read-only, for context.## One policy, in evalshift.yaml
migration_policy:
max_overall_regression_rate: 0.03 # 3% of blocking records may regress
max_critical_regressions: 0 # a count, not a ratio
min_equivalence_rate: 0.95
max_tool_argument_drift: 0.05
max_tool_divergence: 0.10
tool_argument_drift_floor: 0.9 # below this score a call counts as drifted
max_cost_increase: 0.20 # the target may cost 20% more
max_latency_increase: 0.30
fail_on_dropped_params: false
slices:
security:
max_overall_regression_rate: 0.0 # zero tolerance on this sliceRatio fields are decimal fractions (0.20 = 20%), max_critical_regressions is a count, fail_on_dropped_params is a boolean, and slices holds per-slice overrides — each slice may override any budget and inherits the top-level value where it doesn’t.
The CLI resolves the block into a complete policy (every omitted field filled with its default), computes the verdict against it locally, and writes both into the bundle as decision.policy. The server stores that policy with the run. A pushed run therefore carries the budgets that produced its verdict, which is what lets EvalShift Cloud answer for a pull request with the same decision the CLI made on the developer’s machine rather than a second opinion.
run:create — it is evidence about the run, like the rest of the decision block. There is no separate permission for configuring a policy, because there is nothing left to configure over HTTP. A branch can loosen its own gate by editing the yaml; that shows up in the diff, the way every CI config does.## The nine budgets
Defaults are the ones a config that omits a field inherits. For the full field reference — bounds, the Wilson-interval rules, and the --profile presets — see migration_policyin the config reference.
| Budget | Default | Meaning |
|---|---|---|
max_overall_regression_rate | 0.30 | Max share of blocking records allowed to regress. |
max_critical_regressions | 1 | Max count of critical-severity regressions. An integer, not a ratio. |
min_equivalence_rate | 0.75 | Floor on the non-regression rate — equivalent or improved both count. |
max_tool_argument_drift | 0.20 | Max share of tool-argument rows allowed to drift. |
max_tool_divergence | 0.20 | Max share of divergence rows where the target called different tools than the source. |
tool_argument_drift_floor | 0.9 | Not a budget — the score below which a call counts as drifted at all. |
max_cost_increase | 0.30 | Max relative increase in mean per-call cost, target vs source. |
max_latency_increase | 0.30 | Max relative increase in mean per-call latency. |
fail_on_dropped_params | false | Fail when an arm could not honour a generation parameter the capture recorded. A boolean, and top level only. |
The Cloud stores all nine and can re-derive six of them: the last three in the table above — max_tool_divergence, tool_argument_drift_floor and fail_on_dropped_params — have no counterpart in the aggregate metrics a run uploads. That asymmetry is exactly why a run that carries its own policy is answered with the CLI’s stored decision instead of a re-derivation: the CLI checked all nine, and a six-budget re-check would disagree precisely where the other three matter.
Two things no budget can decide. A run whose blocking evaluators scored no records at all is inconclusive however the budgets are set — its quality numbers are clean by absence, not by evidence — though the cost and latency budgets still apply, because those are derived from the calls the run really made. And a rate breach the sample cannot confirm is inconclusive rather than a fail. A slice budget, by contrast, blocks on exactly the terms an overall one does: breach it and the run fails. See Verdicts & gating.
## A run is gated on the policy it was pushed with
The policy stored with a run is a snapshot, and the run is judged against that snapshot forever. Tightening max_cost_increase in evalshift.yaml changes what the next run is held to; it never rewrites the history of runs already pushed, because the verdict on a merged PR is a record of what was known at merge time.
A run’s Policy tab shows a binary status (PASS / FAIL / INCONCLUSIVE), the reason, each budget’s result, and one line naming which policy answered — the one pushed with this run, a legacy web-configured policy, or none at all. The gate is binary even though the badge is four-valued.
## What the web app shows
The Settings tab on a project renders the project’s policy read-only, names the run it came from, and offers it back as a migration_policy: block to copy. There is no editor: PATCH of a project’s policy is rejected, and a member sees exactly what an owner sees.
“The project’s policy” is resolved per request rather than stored, in this order:
- +the policy of the latest available run on the project’s default branch — a feature branch may loosen its own gate for one PR, which is an exception rather than the project’s policy;
- +else the policy of the latest available run on any branch;
- +else a legacy web-configured policy, if the project still has one;
- +else none, and the project is not gated.
With no policy anywhere, the card shows the starter template as a paste-ready migration_policy: block. Those numbers are a suggestion to copy into the file, not budgets anything is held to — EvalShift does not invent a limit nobody chose.
## A run pushed with no policy is not gated
This is a real state rather than a missing one, and it is loud rather than silent. The hosted gate answers inconclusive with no budget rows and nothing blocked, evalshift push warns before the upload that the run carries no policy, and the GitHub Action posts a warning annotation and a commit status that says the gate is off instead of quietly reporting a clear check.
migration_policy block to evalshift.yaml and push again. The CLI behaves identically until you do: with no policy configured it emits inconclusive with a reason and no budget results, so the two never describe the same run differently.## Projects with a web-configured policy
Projects gated through the web app before the policy moved into evalshift.yaml keep that policy as a legacy, read-only value. Runs pushed without a policy of their own are still checked against it — six budgets, re-derived from the run’s stored aggregates, so it can answer differently from the verdict frozen at push time. Runs that carry a policy ignore it entirely.
Move it into the file: the Settings card renders it as a migration_policy: block to copy, and push prints the same block when your config has no policy of its own. Only the budgets the web app actually set are printed — writing the CLI’s other defaults into the file would pin values meant to move with the CLI. Once evalshift.yaml has a policy, the hint stops appearing and the legacy value stops mattering.
## Per-PR enforcement
The Pull Requests tab folds every run that was pushed with a pr_number (e.g. from the GitHub Action) into one row per PR, showing its latest verdict and policy status. PRs whose latest run failed its policy get a Blocked marker. Each row links to the run history filtered to that PR.
- +The status on a row is the latest run’s own decision, so which PRs are blocked changes when a new run is pushed — not when someone edits a setting.
- +The GitHub Action gates on the same policy from the other end: it reads the pushed run’s decision for its commit status, and this view is the server’s record for the web UI and non-GitHub CI.
