Continual Labs·Research

Self-Improving Agents.

Agentic systems that evaluate their own outputs, refine their own prompts and tools, and improve across every run — bounded by guardrails.

What Is a Self-Improving Agent?

A self-improving agent is an agentic system that evaluates its own outputs and uses the results to refine its prompts, tools, and memory across runs.

TL;DR. A self-improving agent closes a loop over itself: it acts, evaluates what it produced, and uses the evaluation to change how it acts next time. The improvement lives in its prompts, tools, and memory — and is measured across runs, not assumed.

Most agents are static: the prompt, the tools, and the strategy are written once by engineers, and every run executes the same plan. A self-improving agent makes the agent itself the optimisation target. After each run it scores its own outputs against an objective, identifies where it failed, and updates something concrete — a prompt, a tool, a memory entry, a planning strategy — so the next run on a similar task does better. Improvement is demonstrated across runs, not promised in a launch post.

Self-reflection is the capability that makes this possible, and it is now a research result rather than a metaphor. Reflexion, for instance, lets a language-model agent store verbal reflections on failed attempts — what went wrong and what to try next — in an episodic memory buffer, then feed those reflections into the next attempt, lifting success rates on coding and reasoning benchmarks (Shinn et al., 2023). The key claim: an agent's own critique of its own output is usable training signal.

The loop we standardise on is eval → refine → act. Evaluate: score the output against explicit criteria, not a gut feeling. Refine: change the smallest thing that explains the failure — the wording of a prompt, the arguments of a tool, the contents of a memory. Act: run again with the refined setup. Repeating that loop is what separates a self-improving agent from a clever one-off.

The Improvement Loop.

Self-evaluation sets the standard, refinement changes the machinery, and memory carries what worked.

TL;DR. Self-improvement is a stack, not a single trick: self-evaluation decides what good looks like on this run, prompt and tool refinement changes how the agent works, and memory carries what worked into the next run.

The loop starts with self-evaluation, and self-evaluation starts with criteria. An agent that judges its own work needs an explicit rubric — did the output compile, did it answer the question, did it stay within budget — plus automated tests wherever tests exist. Vague self-praise is useless. We want the agent to produce a score it can be wrong about, so a later comparison can show whether its judgement itself improved.

Once a failure is located, the agent changes its own machinery. Prompt refinement rewrites the instructions the agent gives itself — adding a missing constraint, tightening a format, inlining the case that failed. Tool refinement changes the interface layer: adding a parameter, swapping a brittle tool for a better one, or writing a small new tool that composes existing ones. Both are cheap, inspectable, and reversible, which is why they are the right first home for self-improvement.

Memory closes the loop across sessions. Which approach passed review, which tactic wasted two hours, which fix worked on the last incident — stored as episodic entries and retrieved when a similar situation appears, so the agent does not relearn the same lesson every run. Memory is also where regressions are cheapest to fix: retract a bad entry and behaviour changes immediately, with no retraining involved.

Guardrails.

Bounded autonomy: the agent may improve itself freely inside a sandbox, and not at all at the gates.

TL;DR. A system that can rewrite itself can also break itself. We bound self-improvement with sandboxes, versioned changes, and human gates — and we keep the evaluation standard fixed outside the loop, so the agent can improve its methods but never its marks.

Unbounded self-modification is the failure mode that keeps this area honest. An agent that can rewrite any part of itself — its prompts, its tools, its evaluator — can optimise away the very mechanism that keeps it correct: it can soften its own rubric, disable a failing check, or memorise a hack that passes today's tests and breaks tomorrow's. Improvement without an external, fixed standard is just drift. The standard must sit outside the loop.

We bound the loop at three levels. Sandboxes: all self-modification happens in an isolated environment where the agent cannot touch production traffic, credentials, or the evaluator it is being scored against. Versioning: every self-change is a diff with a rollback. Gates: changes that pass the sandbox still wait for a human review or a fixed automated check before promotion.

Human review concentrates on the boundaries: which prompts the agent may edit, which tools it may create or modify, which memories it may write, and which actions always require approval. The default is narrow write permission — a short allow-list of things the agent may change, with everything else read-only. When the agent wants more scope, it must earn it with evidence across many runs, and the expansion is reviewed.

Questions, Answered.

Direct answers to the questions people actually ask.

What is a self-improving agent?

A self-improving agent is an agentic system that evaluates its own outputs, uses those evaluations to refine its prompts, tools, and memory, and gets measurably better across runs. The improvement loop is explicit — eval, refine, act — and is bounded by guardrails so the agent cannot silently rewrite the standards it is judged against.

How does an agent evaluate its own work?

Against explicit criteria rather than a general impression: a rubric that scores completeness, correctness, and constraint adherence, plus automated tests where the task allows them. The agent produces a score, cites the evidence for it, and stores both, so later runs can check whether its judgement was right.

What stops an agent from rewriting itself into failure?

Guardrails: self-modification happens in a sandbox, every change is versioned and reversible, and a fixed evaluator outside the loop — plus human gates on anything risky — decides what gets promoted. The agent can change its prompts, tools, and memory, but not the standard it is measured against.

How is this different from agents that just call tools?

A tool-calling agent executes a fixed strategy written by engineers: the same prompt, the same tools, every run. A self-improving agent treats the strategy itself as mutable — after each run it evaluates what failed and changes its prompts, tools, or memory accordingly, so each new run is better than the last.

Start a Mission.

Bring a problem in this area — we will scope it with you.

Start a Mission Last updated: 13 August 2026