Approach
Measurable, inspectable, and under human control.
The Delta Method keeps every AI engagement honest: agree the change, build the simplest system that can produce it, prove it, and keep proving it.
Stages
Stage 1 of 5
What change should this system produce, and how will we measure it?
We map the workflow, the people in it and the data behind it, then agree a baseline and a success metric before any model is chosen.
Leaves behind
Outcome brief with baseline metric
Release gates
What has to be true before anything ships.
These gates are set per engagement. The numbers come from your baseline, never from a generic benchmark.
| Gate | How it is measured |
|---|---|
| Task quality | Graded evaluation set built from your real cases, compared with the agreed baseline. |
| Groundedness | Share of answers fully supported by cited sources; unsupported claims block release. |
| Latency | p50 and p95 end-to-end response time against the agreed budget. |
| Cost per task | Model, retrieval and infrastructure cost per completed task, tracked per release. |
| Safety | Red-team prompts, permission checks on tools and human approval for consequential actions. |
| Operability | Traces, alerts and a runbook in place, with a named owner, before go-live. |
Budgets
Quality, latency and cost are design inputs.
Drag the marker, or choose a priority below.
- Model
- Mid-tier model with routing to a larger model for hard cases
- Retrieval
- Hybrid retrieval (≈25 candidates) with reranking
- Caching
- Prompt caching on shared context
- Oversight
- Sampled human review with drift alerts
Illustrative design heuristics. Real choices are set by your evaluation results and budgets.
Principles
How we make decisions.
Measured, not assumed
Every engagement starts with a baseline and ends with a comparison against it.
People stay in control
Consequential actions pause for human approval. Automation earns autonomy gradually.
Explicit over clever
State machines, typed contracts and written decisions make systems easier to trust and to change.
Budgets are features
Quality, latency and cost targets are agreed up front and tracked like any other requirement.
Toolbox
Tools we work with.
Chosen per project against your constraints. We are not resellers or certified partners of any vendor listed.
- LangGraph
- Model Context Protocol
- pgvector
- Hybrid search + rerankers
- LangSmith
- Docker
- Kubernetes
- GitHub Actions
- LoRA / QLoRA fine-tuning
- Next.js
- Connect RPC
- XState
Tell us the change you need. We will tell you how we would measure it.
A short first conversation, a written summary afterwards, no obligation.