EvalShift
compareevalshift-vs-langsmith

EvalShift vs LangSmith

LangSmith is LangChain's framework-agnostic platform for tracing, evaluating, deploying and monitoring agents. EvalShift is a migration and regression testing tool: it runs one golden suite through two models and decides whether the candidate is safe to ship.

claims checked 2026-10-01LangSmith ↗

// 01
where each one sits
the category ladder, broad to narrow
  1. general LLM evaluation
  2. agent engineering platforms — trace, evaluate, deploy, monitor
  3. LLM migration testing — whether a change is safe to ship
  4. migration testing for agents — tool calls, arguments, ordering
// 02
what each is built for
two jobs, not one job done twice

+ what EvalShift is built around

  • The SDK records what your agent actually did — model calls, tool calls, retrievals — and evalshift capture sync promotes those captures into a golden suite instead of hand-written fixtures.
  • One run sends every example through both models, source and target, so each example yields a paired delta rather than two independent averages.
  • Tool selection, argument correctness, call ordering, parallelism and refusals are built-in scored evaluators, with severity floors per evaluator.
  • Deltas go through a Shapiro-Wilk screen into a paired t-test or Wilcoxon, with Cohen's d, 95% CIs and Benjamini-Hochberg FDR across every comparison in the run.
  • A migration policy collapses that into one verdict — budgets for regression rate, critical count, equivalence, tool-argument drift, cost and latency, with per-slice overrides.
  • The GitHub Action runs the suite on the PR, keeps one comment updated, and fails the check when the policy says the candidate is not safe to ship.

· what LangSmith is for

  • It traces runs and multi-turn threads from LangChain and LangGraph, the OpenAI, Anthropic and Vercel AI SDKs, other frameworks or plain code, and accepts OpenTelemetry spans.
  • Every tracing project gets a dashboard for traces, latency, errors, cost, tokens and tool calls, with threshold alerts routed to Slack, PagerDuty or a webhook.
  • Datasets run through evaluate() as experiments, scored by code, LLM-as-a-judge, human, summary or pairwise evaluators, and a comparison view marks each example that improved or regressed against a baseline experiment.
  • The open-source openevals and agentevals packages cover agent trajectories — strict, unordered, subset or superset match with tool-argument modes — plus multi-turn user simulation.
  • Online evaluators score live production runs and threads; automation rules sample traffic into datasets or annotation queues for human review.
  • Versioned prompts, a playground and LangSmith Deployment — a runtime for agent workloads — sit in the same platform, on LangSmith Cloud or self-hosted on the Enterprise plan.
// 03
side by side
every row checked against the sources below
Statistical significance
evalshiftShapiro-Wilk screen, paired t-test or Wilcoxon, Cohen's d with 95% CI, Benjamini-Hochberg FDR at α=0.05.
LangSmithCounts of examples that improved or regressed per score. Its docs describe no significance test, confidence interval or multiple-comparison correction.
Ship / no-ship verdict
evalshiftA migration policy in evalshift.yaml — nine budgets plus per-slice overrides — that gates locally and rides into every pushed run.
LangSmithAssertions you write in pytest, Vitest or Jest tests; a failing assertion fails the CI job.
PR gate
evalshiftThe Action pushes the run, keeps one PR comment updated, and fails the check on the policy verdict.
LangSmithNo official GitHub Action is documented; the CI example runs the pytest suite inside your own GitHub Actions workflow.
Where raw data lives
evalshiftRun, scoring and HTML report stay under .evalshift/ on your machine; only evalshift push uploads a finalized bundle.
LangSmithTraces go to LangSmith Cloud (US, EU or APAC) or a self-hosted Enterprise deployment; inputs and outputs can be hidden or masked first.
Baseline vs candidate
evalshiftOne paired run puts every example through both models; the unit of analysis is the per-example delta.
LangSmithTwo experiments on one dataset, compared example by example against the experiment you pick as baseline, or scored head to head by a pairwise evaluator.
Agent tool-call scoring
evalshiftTool selection, arguments, call count, parallelism and refusals ship as evaluators with severity levels.
LangSmithagentevals matches a trajectory against a reference you supply, or has an LLM judge it; you wire it in as an evaluator.
Where test cases come from
evalshiftCaptured production runs, promoted turn by turn into golden.jsonl, deduplicated across syncs.
LangSmithDataset examples from traces, annotation queues, the playground, CSV or JSONL import, LLM-generated synthetic data, or the SDK.
LLM-as-a-judge
evalshiftPairwise A/B between the two outputs, order-randomized, with an explicit tie score.
LangSmithJudges on experiments and on live production runs, plus pairwise evaluation with optional order randomization.
Production monitoring
evalshiftNone. EvalShift runs offline suites before a change ships and does not watch live traffic.
LangSmithPrebuilt and custom dashboards for latency, errors, cost and tokens, with threshold alerts to Slack, PagerDuty or a webhook.
Human review loop
evalshiftNone. Judgement is encoded in evaluators and policy budgets, not collected from reviewers.
LangSmithAnnotation queues for runs and threads, plus feedback logged from end users, annotators or evaluators through the SDK.
// 04
which one to reach for
honest routing — both answers are real

reach for evalshift when

  • You are swapping models or providers and need evidence the candidate holds before it reaches users.
  • You want “did it regress” answered with a significance test, an effect size and FDR correction, not a tally of examples that moved.
  • What breaks is agent behaviour — a tool that stops being called, an argument that drifts, an extra round — and text-only scoring reads green.
  • You want the merge blocked by a stated budget, not by an assertion threshold someone has to pick per test.
  • Raw prompts and responses should stay on your machine, with only finalized run data pushed anywhere.

reach for LangSmith when

  • You want tracing, evaluation, prompt versions and agent deployment on one platform, especially if you already build on LangChain or LangGraph.
  • You need to see what production is doing right now — latency, errors, cost — and get paged when a threshold is crossed.
  • Your evals belong in the pytest, Vitest or Jest suite the team already runs, scored against reference outputs you curate.
  • Your quality loop is human — annotation queues, labelled examples, end-user feedback — rather than a comparison of two model versions.

They cover different stretches of the same loop, and running both is normal. LangSmith traces production, scores live runs and routes the ones worth a look into datasets and annotation queues; EvalShift takes the question those cases raise — does this model or prompt change make them worse — and answers it before the change ships. A workable split: keep LangSmith tracing in production for monitoring, online evaluation and human review, run the EvalShift SDK alongside it so the agent's model and tool calls also land in .evalshift/captures/, promote those into a golden suite, and let the EvalShift Action gate the PR that changes the model. The two SDKs are independent — LangSmith sends traces to its server, EvalShift writes JSON captures to disk — so running both is a configuration question, not an integration.

// 05
questions
the ones people ask
Can I use EvalShift and LangSmith together?
Yes. LangSmith covers production tracing, dashboards, online evaluation and human review; EvalShift covers the pre-merge question of whether a model or prompt change regresses your golden suite.
Is LangSmith an alternative to EvalShift?
They overlap on datasets, experiment comparison and CI testing, but they answer different questions. LangSmith is a platform for building, running and observing agents; EvalShift is built around one decision — whether a candidate model is safe to ship.
Does LangSmith test whether a difference between experiments is significant?
Its documentation does not describe significance tests, confidence intervals or multiple-comparison correction; the comparison view counts which examples improved or regressed on each score. EvalShift runs a paired t-test or Wilcoxon per comparison, reports Cohen's d with a 95% CI, and applies Benjamini-Hochberg FDR across the run.
Does LangSmith only work with LangChain?
No. LangChain's docs describe it as framework-agnostic: it traces the OpenAI, Anthropic and Vercel AI SDKs, other frameworks and plain code, and accepts OpenTelemetry. EvalShift is framework-agnostic too — the capture SDK records model and tool calls, and models run through LiteLLM.
Does EvalShift send my prompts and responses to a server?
Only when you run evalshift push. The run, the scoring and the HTML report stay under .evalshift/ on your machine; the push uploads a finalized bundle containing example inputs, both models' outputs, scores and the analysis.

sources

run your own migration diff

the CLI is open source and runs offline — no account, no upload, nothing leaves your machine until you push a run.