EvalShift vs LangSmith
LangSmith is LangChain's framework-agnostic platform for tracing, evaluating, deploying and monitoring agents. EvalShift is a migration and regression testing tool: it runs one golden suite through two models and decides whether the candidate is safe to ship.
claims checked 2026-10-01LangSmith ↗
- general LLM evaluation
- agent engineering platforms — trace, evaluate, deploy, monitor
- LLM migration testing — whether a change is safe to ship
- migration testing for agents — tool calls, arguments, ordering
+ what EvalShift is built around
- The SDK records what your agent actually did — model calls, tool calls, retrievals — and
evalshift capture syncpromotes those captures into a golden suite instead of hand-written fixtures. - One run sends every example through both models, source and target, so each example yields a paired delta rather than two independent averages.
- Tool selection, argument correctness, call ordering, parallelism and refusals are built-in scored evaluators, with severity floors per evaluator.
- Deltas go through a Shapiro-Wilk screen into a paired t-test or Wilcoxon, with Cohen's d, 95% CIs and Benjamini-Hochberg FDR across every comparison in the run.
- A migration policy collapses that into one verdict — budgets for regression rate, critical count, equivalence, tool-argument drift, cost and latency, with per-slice overrides.
- The GitHub Action runs the suite on the PR, keeps one comment updated, and fails the check when the policy says the candidate is not safe to ship.
· what LangSmith is for
- It traces runs and multi-turn threads from LangChain and LangGraph, the OpenAI, Anthropic and Vercel AI SDKs, other frameworks or plain code, and accepts OpenTelemetry spans.
- Every tracing project gets a dashboard for traces, latency, errors, cost, tokens and tool calls, with threshold alerts routed to Slack, PagerDuty or a webhook.
- Datasets run through
evaluate()as experiments, scored by code, LLM-as-a-judge, human, summary or pairwise evaluators, and a comparison view marks each example that improved or regressed against a baseline experiment. - The open-source openevals and agentevals packages cover agent trajectories — strict, unordered, subset or superset match with tool-argument modes — plus multi-turn user simulation.
- Online evaluators score live production runs and threads; automation rules sample traffic into datasets or annotation queues for human review.
- Versioned prompts, a playground and LangSmith Deployment — a runtime for agent workloads — sit in the same platform, on LangSmith Cloud or self-hosted on the Enterprise plan.
.evalshift/ on your machine; only evalshift push uploads a finalized bundle.golden.jsonl, deduplicated across syncs.reach for evalshift when
- You are swapping models or providers and need evidence the candidate holds before it reaches users.
- You want “did it regress” answered with a significance test, an effect size and FDR correction, not a tally of examples that moved.
- What breaks is agent behaviour — a tool that stops being called, an argument that drifts, an extra round — and text-only scoring reads green.
- You want the merge blocked by a stated budget, not by an assertion threshold someone has to pick per test.
- Raw prompts and responses should stay on your machine, with only finalized run data pushed anywhere.
reach for LangSmith when
- You want tracing, evaluation, prompt versions and agent deployment on one platform, especially if you already build on LangChain or LangGraph.
- You need to see what production is doing right now — latency, errors, cost — and get paged when a threshold is crossed.
- Your evals belong in the pytest, Vitest or Jest suite the team already runs, scored against reference outputs you curate.
- Your quality loop is human — annotation queues, labelled examples, end-user feedback — rather than a comparison of two model versions.
They cover different stretches of the same loop, and running both is normal. LangSmith traces production, scores live runs and routes the ones worth a look into datasets and annotation queues; EvalShift takes the question those cases raise — does this model or prompt change make them worse — and answers it before the change ships. A workable split: keep LangSmith tracing in production for monitoring, online evaluation and human review, run the EvalShift SDK alongside it so the agent's model and tool calls also land in .evalshift/captures/, promote those into a golden suite, and let the EvalShift Action gate the PR that changes the model. The two SDKs are independent — LangSmith sends traces to its server, EvalShift writes JSON captures to disk — so running both is a configuration question, not an integration.
- Can I use EvalShift and LangSmith together?
- Yes. LangSmith covers production tracing, dashboards, online evaluation and human review; EvalShift covers the pre-merge question of whether a model or prompt change regresses your golden suite.
- Is LangSmith an alternative to EvalShift?
- They overlap on datasets, experiment comparison and CI testing, but they answer different questions. LangSmith is a platform for building, running and observing agents; EvalShift is built around one decision — whether a candidate model is safe to ship.
- Does LangSmith test whether a difference between experiments is significant?
- Its documentation does not describe significance tests, confidence intervals or multiple-comparison correction; the comparison view counts which examples improved or regressed on each score. EvalShift runs a paired t-test or Wilcoxon per comparison, reports Cohen's d with a 95% CI, and applies Benjamini-Hochberg FDR across the run.
- Does LangSmith only work with LangChain?
- No. LangChain's docs describe it as framework-agnostic: it traces the OpenAI, Anthropic and Vercel AI SDKs, other frameworks and plain code, and accepts OpenTelemetry. EvalShift is framework-agnostic too — the capture SDK records model and tool calls, and models run through LiteLLM.
- Does EvalShift send my prompts and responses to a server?
- Only when you run
evalshift push. The run, the scoring and the HTML report stay under.evalshift/on your machine; the push uploads a finalized bundle containing example inputs, both models' outputs, scores and the analysis.
sources
- LangSmith — product overview ↗
- LangSmith — observability & tracing ↗
- LangSmith — trace with OpenTelemetry ↗
- LangSmith — dashboards ↗
- LangSmith — alerts ↗
- LangSmith — manage datasets ↗
- LangSmith — evaluation types ↗
- LangSmith — compare experiment results ↗
- LangSmith — pairwise evaluation ↗
- LangSmith — trajectory evals (agentevals) ↗
- LangSmith — online LLM-as-a-judge evaluation ↗
- LangSmith — annotation queues ↗
- LangSmith — pytest integration ↗
- LangSmith — CI/CD pipeline example ↗
- LangSmith — cloud regions ↗
- LangSmith — self-hosted (Enterprise) ↗
- LangSmith — mask inputs & outputs ↗
- LangSmith — deployment ↗
- openevals on GitHub ↗
run your own migration diff
the CLI is open source and runs offline — no account, no upload, nothing leaves your machine until you push a run.
