Close the loop with verified improvement
Turn production evidence into requirements, tests, controlled changes, and measured outcomes. Define what self-improving software can responsibly mean.
Published by TaigaHow we write
What you will learn
- Connect an operational observation to a verifiable engineering change.
- Separate runtime recovery, workflow improvement, and model training.
- Measure a claimed improvement without weakening its evaluation.
Define the loop you want to close
Software produces evidence during use: errors, delays, support requests, incidents, maintenance findings, and repeated manual work. A complete lifecycle returns that evidence to engineering decisions.
Self-improving software can mean that automation helps identify, propose, implement, and verify changes. It does not necessarily mean that a model trains itself. State which part changes: application code, configuration, tests, instructions, workflow, or model parameters.
Self-healing restores a known operating condition. Self-improvement changes the system to produce a better outcome in future. The second claim requires a comparison and protection against regressions.
Follow one observation through the lifecycle
The sequence below is a proposed engineering method. It is not a claim that any product performs every step autonomously.
| Stage | Required output | Fictional export example |
|---|---|---|
| Observe | Versioned evidence with scope and uncertainty | Worker memory rises during large exports |
| Diagnose | Testable cause and competing explanations | Retained row buffers may explain memory growth |
| Specify | Desired outcome and constraints | Stream rows without changing permissions or output |
| Reproduce | A test that exposes the original failure | A representative large synthetic export exceeds the limit |
| Change | A reviewable correction | Release completed row buffers during streaming |
| Evaluate | Old failure addressed; other requirements preserved | Memory test, output comparison, authorization, and retry checks pass |
| Release | Controlled exposure with recovery criteria | Limited rollout of an identified artifact |
| Verify | Comparable production evidence and an owner | Memory stabilizes while correctness and latency remain acceptable |
Keep links between these outputs. A postmortem action that says “improve monitoring” is difficult to verify. A defined signal, owner, threshold, and tested response make completion observable.
Keep evaluation independent from the proposal
An agent can create a patch and propose tests. The team must still inspect whether those tests detect the original problem. Preserve a versioned evaluation set that the change cannot quietly weaken.
For the fictional memory leak, compare equivalent workloads and versions. Include large exports, cancellation, retry, and access-denial cases. Use synthetic data that represents the relevant shapes without exposing customer records.
Reject a faster export if it drops records, bypasses authorization, or exceeds the allowed cost. Define these constraints before optimization. Otherwise, the system can improve the chosen metric while making the service worse.
If you change an agent’s instructions or model, evaluate its behavior on representative tasks and known failures. Keep the previous version available. Updating instructions is not evidence that the underlying model learned from an incident.
Release and measure the result
A canary release exposes a limited population to a candidate version. Compare candidate and control signals, and define when to expand or stop. Sparse traffic or different workloads can make the comparison inconclusive. Canary guidance.
The fictional team records a baseline from a fixed synthetic workload. It tests the correction, releases within an approved boundary, and checks comparable production periods. If evidence remains insufficient, it records uncertainty rather than declaring a gain.
Measure repeated manual work as well. Automation can reduce toil, but it also needs maintenance and failure handling. Include these costs when judging the result. Toil guidance.
Make the feedback record usable
Use these fields for the exercise: observation and version; baseline; proposed cause; acceptance criteria; regression checks; change and review; release boundary; measured result; owner and next review.
Taiga Maintaining connects repository findings to remediation work. Initiatives connect an intended change to planning and delivery. These provide parts of an evidence chain. Your service owner must still verify deployment and the operational result. Maintaining, Initiatives.
A mature software factory connects this work across products. Keep the decision rights and evaluation criteria visible as automation increases. The final proof is a better verified service, not a larger count of generated changes.
Do the exercise
Complete the feedback record in this lesson for the fictional memory leak. Define a baseline, acceptance test, regression checks, release boundary, production measurement, and owner. Add a rule for rejecting a faster but less correct export.
Download worksheet (Markdown)Check your understanding
Sources & further reading
- Google SRE: Postmortem Culture ↗
- Google SRE: Canarying Releases ↗
- Google SRE: Eliminating Toil ↗
- Taiga docs: Maintaining ↗
- Taiga docs: Initiatives ↗
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.