Continual Labs·Research
Eval-Driven Development.
Benchmarks, regression suites, and behavioural evals for every model change — the CI/CD of AI.
01 /
What Is Eval-Driven Development?
Evaluations treated the way software engineering treats tests: no change ships until the suite passes.
Eval-driven development treats evaluations the way software engineering treats tests: a model, prompt, or agent change is not shipped until it passes a suite of evals that encodes what good behaviour means.
TL;DR. EvDD makes model behaviour a build artefact: every change runs the eval suite, every regression fails the build, and shipping decisions are made on evidence, not vibes.
A benchmark is a static score. An eval suite is a product: curated cases, golden answers, and graders that check whether the system did the right thing in situations your users will actually meet. The suite is versioned with the code, run on every change, and treated as a contract between the lab and the people who depend on the system.
The discipline is borrowed from continuous integration. Tests in software pin down what code must keep doing; evals pin down what a model must keep doing — while the model itself changes underneath. Every prompt edit, fine-tune, or model swap is a candidate regression until the suite says otherwise.
02 /
The Eval Suite.
Unit evals for single skills, regression suites for what must not break, behavioural evals for what is hard to measure.
A serious suite is layered: unit evals for single skills, regression suites for things that must not break, and behavioural evals for the qualities that are hard to measure.
TL;DR. Unit evals test one capability in isolation; regression suites freeze past failures; behavioural evals measure how the system acts under pressure.
Unit evals isolate one skill — summarising a contract, extracting a date, refusing a request that is out of scope. They are fast, deterministic, and run first. Regression suites are the memory of past failures: every bug that reached a user becomes a case that must never pass silently again.
Behavioural evals take on the fuzzy properties — honesty, safety, tone, refusal boundaries — using rubric graders and, where judgement is needed, model-graded checks sampled by humans. They are the layer that keeps an optimised model from becoming a technically correct but wrong-behaving one.
03 /
The Loop.
Not a release gate at the end — the loop the lab runs every day.
Evals are not a release gate at the end; they are the loop the lab runs every day.
TL;DR. Propose a change, run the suite, read the diffs, fix or ship. The same loop applies to prompts, tools, and models alike.
Every mission gets an eval suite before it gets an optimisation budget. That ordering matters: you cannot improve what you cannot measure, and you cannot ship what you cannot regress. The suite also decays — eval drift is real. Cases age out as behaviour shifts, and the suite itself is a product that needs maintenance.
The payoff is organisational: decisions stop being arguments. When a prompt change scores better on the suite, it ships; when it scores worse, it does not — and the conversation moves to whether the suite measures the right thing, which is a much better argument to have.
05 /
Questions, Answered.
Direct answers to the questions people actually ask.
What is an LLM eval?
A test that checks an AI system's behaviour against a defined case — input, expected outcome, and a grader that decides pass or fail. Evals turn model quality into something measurable and repeatable.
What is a regression suite for a model?
A set of eval cases drawn from past failures and core requirements, run on every model or prompt change, so improvements cannot silently break behaviour that already worked.
What is eval drift?
The slow decay of an eval suite's usefulness as products, users, and model capabilities change. Cases go stale, thresholds stop meaning what they used to, and the suite needs the same maintenance as any other code.
How is eval-driven development different from benchmarking?
Benchmarking measures a model against a public, general test set. Eval-driven development runs private, product-specific suites on every change to gate shipping — benchmarks inform, evals decide.
06 /
Start a Mission.
Bring a problem in this area — we will scope it with you.
Start a Mission Last updated: 13 August 2026