# The Eval-Driven Development Loop — Research Notes — Continual Labs Limited

Continual Labs · Research Notes

# The Eval-Driven Development Loop.

Evaluations are the CI/CD of AI — the feedback loop that keeps every model change honest.

02 /

## Why Regression Suites Beat Vibes.

When models change daily, judgement calls become unreviewable. A suite turns "I think it is better" into "the suite says it is".

Every lab that works with models eventually confronts the same moment: the model changed last night, and nobody can say with certainty whether the system got better.

TL;DR. A regression suite produces a diff, not a mood — and doubles as the institutional memory of every failure that reached a user.

Prompts get edited. Fine-tunes land. Providers ship new model versions. In that world, the sentence "it feels better" is a liability, because feelings are not reproducible and not reversible. A regression suite is the alternative: a set of cases with known-good behaviour that runs on every change and produces a diff, not a mood. That is the core of what we describe in eval-driven development .

The suite is also a memory. Every failure that reaches a user becomes a case that must never pass silently again. Over time the suite grows into an institutional record of everything the system is supposed to be — exactly what a test suite is for software, and exactly what models have been missing.

03 /

## Drift: The Second Law of Eval Systems.

Eval suites decay. Cases age out, thresholds stop meaning what they used to, and the suite quietly stops measuring the product.

We call it eval drift, and it is the reason an eval suite is a product, not a project.

TL;DR. A suite that is not maintained is worse than no suite, because it produces false confidence.

Products change; users change; model capabilities change. A case written six months ago may now be testing behaviour nobody ships anymore, or passing trivially for reasons that have nothing to do with quality. The graders themselves drift — a rubric that was decisive last quarter may be ambiguous this one.

The remedy is maintenance discipline: review the suite on a schedule, retire stale cases, add a case for every regression, and re-tune thresholds when the graders or the models change. The second law of eval systems is that entropy always wins, and the only defence is scheduled attention.

04 /

## Evals as Requirements.

The best eval cases are written before the work, as product requirements in test form.

We write the eval before we write the improvement. What must the system do? What must it refuse? What did it get wrong last week that it must never get wrong again?

TL;DR. Green means ship, red means fix — and the only interesting conversation is whether the suite measures the right thing.

Those questions, in case form, are the requirements. The improvement is whatever makes the suite pass without breaking the rest — the same discipline our self-improving agents apply to their own outputs.

That ordering changes the economics of the whole loop. Building stops being open-ended. Reviews stop being arguments over taste. Shipping stops being a leap of faith and becomes a build artefact: green means ship, red means fix, and the only interesting conversation is whether the suite measures the right thing — which is a much better argument to have.

05 /

## Start a Mission.

Bring a problem in this area — we will scope it with you.

Start a Mission Last updated: 13 August 2026
