# Model Optimization — Research — Continual Labs Limited

Continual Labs · Research

# Model Optimization.

Quantization, distillation, speculative decoding, and adaptive serving — frontier capability at commodity cost.

01 /

## What Is Model Optimization?

Making inference faster and cheaper while keeping the behaviour the eval suite certifies.

Model optimization is the work of making inference faster and cheaper while keeping the behaviour the eval suite certifies. It is how frontier capability reaches commodity cost.

TL;DR. Optimization attacks latency and cost from both ends: compress the model with quantization and distillation, and accelerate how it runs with speculative decoding and adaptive serving.

Raw model capability is only half the product; the other half is the cost of delivering it per request. An inference pipeline that costs ten times what it should is a research demo, not a system. Optimization is where a model becomes something you can afford to run at scale.

The constraint that makes the discipline interesting: every optimisation is a trade against quality. The eval suite is what makes the trade legible — a quantised model either still passes the behavioural evals or it does not. Optimisation without evals is just guessing with extra steps.

02 /

## The Techniques.

Four techniques cover most of the ground: quantization, distillation, speculative decoding, adaptive serving.

Four techniques cover most of the ground: quantization shrinks the weights, distillation shrinks the model, speculative decoding accelerates generation, and adaptive serving routes work to the right model.

TL;DR. Quantization reduces numerical precision; distillation transfers capability into a smaller model; speculative decoding drafts and verifies tokens in parallel; adaptive serving matches each request to the cheapest model that passes its bar.

Quantization stores weights at lower precision — 8-bit, 4-bit, and below — shrinking memory use and speeding up matrix arithmetic. Distillation (Hinton, Vinyals &amp; Dean, 2015) trains a small student model on the behaviour of a larger teacher, preserving the capabilities that matter while dropping the bulk.

Speculative decoding (Leviathan, Kalman &amp; Matias, 2023) runs a small draft model ahead and lets the large model verify its guesses in parallel — faster output, identical distribution. Adaptive serving takes the portfolio view: route easy requests to small models, escalate hard ones, and cut over between backends as load and cost change.

03 /

## Cost-Engineering.

Optimization decisions are budget decisions: latency, cost, and quality — chosen together, enforced by the suite.

Optimization decisions are budget decisions: a latency target, a cost ceiling, and a quality floor — chosen together, enforced by the eval suite.

TL;DR. Pick the constraint that binds first — latency, cost, or quality — then optimise against it and gate every change on the suite.

The binding constraint changes by mission. A real-time agent has a latency budget; a batch analytics pipeline has a cost ceiling; a customer-facing assistant has a quality floor it must never cross. The right optimisation is the one that buys the most of the scarce resource per point of quality spent.

We treat optimisation as a pipeline, not an event: quantise, distill, measure; route, cache, measure again. Each stage is reversible and each stage is gated. When the numbers stop moving in the right direction, the optimisation is done — and the evals say so.

04 /

## Related Research.

Follow the loop into the next area.

Continual Learning · Self-Improving Agents · Eval-Driven Development · Agentic Orchestration · What We Mean by Self-Improving Systems

05 /

## Questions, Answered.

Direct answers to the questions people actually ask.

**What is quantization? +**

Storing a model's weights at lower numerical precision — such as 8-bit or 4-bit instead of 16-bit — to reduce memory use and speed up computation, at a measured cost in quality that evals gate.

**What is knowledge distillation? +**

Training a smaller "student" model to reproduce the behaviour of a larger "teacher" model, so most of the capability ships in a fraction of the size and cost.

**What is speculative decoding? +**

A decoding strategy where a small draft model proposes tokens and the large model verifies them in parallel, producing the same output distribution faster.

**When does optimization hurt model quality? +**

When compression or routing decisions are not gated by an eval suite. Aggressive quantization, careless distillation, and over-eager routing to small models can each degrade behaviour; the suite is what catches it before users do.

06 /

## Start a Mission.

Bring a problem in this area — we will scope it with you.

Start a Mission Last updated: 13 August 2026
