← All Articles
Evals & TestingCI/CDMLOps

Why We Run LLM Evals in CI Like Tests, Not as a One-Off Notebook

Christian Chukwuka··5 min read
Why We Run LLM Evals in CI Like Tests, Not as a One-Off Notebook
TL;DR

A one-off eval notebook answers "did this look good once"; a CI-integrated eval suite answers "did this change make things worse, every time, automatically." Structure evals like tests — a fixed question set, LLM-as-judge grading, and a pass/fail gate — so a regression blocks a merge instead of reaching production.

Most teams building with LLMs have run an eval at some point — a notebook, a spreadsheet of test prompts, a manual pass through twenty examples before a demo. Almost none of them run that same eval automatically on every prompt change, model swap, or retrieval update. That gap is where regressions live.

The notebook trap

A notebook run once tells you the system was good at one point in time, against one version of the prompt. It says nothing about whether the change someone is about to merge makes things better or worse, because nobody re-runs the notebook on every change — it's manual, slow, and easy to skip under deadline pressure. This is exactly the failure mode unit tests were invented to prevent for regular code, and it applies just as directly here.

What an eval suite needs to look like a test suite

  • A fixed, versioned set of test questions or scenarios — not ad-hoc examples someone thinks of that morning.
  • An automated grader for each response — LLM-as-judge, exact-match, or a rules-based check, depending on the task.
  • A pass/fail threshold that blocks a merge or deploy, not just a dashboard someone might look at later.
  • Execution on every relevant change — prompt edits, model version bumps, retrieval strategy changes — via CI, the same trigger as unit tests.

LLM-as-judge grading: useful, with caveats

Using a strong model to grade another model's output works well for open-ended quality checks — faithfulness, completeness, tone — where exact-match scoring doesn't apply. The caveat is that the judge model needs its own fixed rubric and, ideally, periodic calibration against human-graded examples, or grading drift becomes its own silent regression.

Regression detection is the actual point

The value isn't the score on any single run — it's the diff between runs. A retrieval strategy change that drops accuracy on a specific category of question, a model swap that quietly increases hallucination rate, a system-prompt edit that improves one metric while degrading another: none of these show up from eyeballing a few examples, but they show up immediately in a scored, versioned eval run compared against the previous one.

What this catches that manual review doesn't

Manual review catches the failure someone happens to test for. An automated, comprehensive question set catches the failure nobody thought to check — the edge case in a category the reviewer wasn't focused on that day. At scale, that's the difference between a regression caught in CI and one reported by a customer.

The dataset matters as much as the pipeline

A CI-integrated eval suite is only as good as what's in the question set it runs. A dataset that's only ever tested the easy, obvious cases will happily pass a genuinely worse model — the pipeline mechanics being correct doesn't help if the content underneath it never covers the failure modes that actually matter. Structuring that dataset by domain rather than as one flat list also makes a huge difference in practice: an aggregate score dropping 3% is vague, but "the payments-support slice dropped 12%, everything else is flat" is immediately actionable.

The baseline question

One design decision that matters more than it first appears: what exactly is a run being compared against? A rolling "recent average" drifts every time a dataset or config changes, which makes "did this get worse" a fuzzier question than it should be. Pinning an explicit baseline — a specific run someone actually reviewed and approved, by ID — and comparing every future run against exactly that one, until a new baseline is deliberately approved, is a stricter and more honest signal. It's a small structural choice with an outsized effect on whether the CI gate can be trusted.

Severity, not just pass/fail

Not every regression should block a deploy the same way. A severe drop in a critical category should stop a merge cold; a minor dip in an edge case probably shouldn't block a hotfix. A binary pass/fail gate forces every regression into the same bucket, which either makes the gate too strict to be usable or too loose to catch what matters. Severity-aware thresholds — configurable per pipeline — let a team decide that bar explicitly instead of accepting whatever a single fixed cutoff gives them.

Putting this into practice

This piece is the philosophy; the mechanics are a separate, more concrete problem. For the actual dataset-building work, see setting up golden datasets for LLM regression testing. For the CI wiring itself — explicit baselines, severity-aware exit codes, gating a merge rather than just reporting a score — see how to add LLM regression testing to a CI/CD pipeline, which is also where EvalCI, the tool we eventually built around this exact philosophy, actually lives.

Frequently asked questions

How is a CI-integrated eval suite different from just running evals manually before a release?

Consistency and coverage. A manual pre-release check depends on someone remembering to run it and choosing what to test that day — it catches what the reviewer happens to think of. A CI-integrated suite runs the same fixed, comprehensive question set automatically on every relevant change, which catches regressions nobody was specifically looking for.

Does this replace manual review entirely?

No — it changes what manual review is for. Automated evals catch known failure modes at scale and on every change; manual review is still valuable for genuinely novel behavior an existing dataset wasn't built to catch, and for periodically checking whether the dataset itself needs updating.

The takeaway

Treat evals as a CI gate, not a pre-demo ritual. A fixed question set, automated grading, and a merge-blocking threshold turn "does this still work" from a hopeful guess into a number checked on every change — which is the entire point of having tests in the first place.

Christian Chukwuka
Christian Chukwuka
Founder & AI Systems Engineer

Have a similar challenge?

Book a free 30-minute architecture call and we'll tell you honestly whether and how we can help.

Book Free Discovery Call →