Path 05Lesson 8 / 8

Close the loop with verified improvement

Turn production evidence into requirements, tests, controlled changes, and measured outcomes. Define what self-improving software can responsibly mean.

Advanced12 minReviewed

Published by How we write

What you will learn

  • Connect an operational observation to a verifiable engineering change.
  • Separate runtime recovery, workflow improvement, and model training.
  • Measure a claimed improvement without weakening its evaluation.

Define the loop you want to close

Software produces evidence during use: errors, delays, support requests, incidents, maintenance findings, and repeated manual work. A complete lifecycle returns that evidence to engineering decisions.

Self-improving software can mean that automation helps identify, propose, implement, and verify changes. It does not necessarily mean that a model trains itself. State which part changes: application code, configuration, tests, instructions, workflow, or model parameters.

Self-healing restores a known operating condition. Self-improvement changes the system to produce a better outcome in future. The second claim requires a comparison and protection against regressions.

Follow one observation through the lifecycle

The sequence below is a proposed engineering method. It is not a claim that any product performs every step autonomously.

StageRequired outputFictional export example
ObserveVersioned evidence with scope and uncertaintyWorker memory rises during large exports
DiagnoseTestable cause and competing explanationsRetained row buffers may explain memory growth
SpecifyDesired outcome and constraintsStream rows without changing permissions or output
ReproduceA test that exposes the original failureA representative large synthetic export exceeds the limit
ChangeA reviewable correctionRelease completed row buffers during streaming
EvaluateOld failure addressed; other requirements preservedMemory test, output comparison, authorization, and retry checks pass
ReleaseControlled exposure with recovery criteriaLimited rollout of an identified artifact
VerifyComparable production evidence and an ownerMemory stabilizes while correctness and latency remain acceptable

Keep links between these outputs. A postmortem action that says “improve monitoring” is difficult to verify. A defined signal, owner, threshold, and tested response make completion observable.

Keep evaluation independent from the proposal

An agent can create a patch and propose tests. The team must still inspect whether those tests detect the original problem. Preserve a versioned evaluation set that the change cannot quietly weaken.

For the fictional memory leak, compare equivalent workloads and versions. Include large exports, cancellation, retry, and access-denial cases. Use synthetic data that represents the relevant shapes without exposing customer records.

Reject a faster export if it drops records, bypasses authorization, or exceeds the allowed cost. Define these constraints before optimization. Otherwise, the system can improve the chosen metric while making the service worse.

If you change an agent’s instructions or model, evaluate its behavior on representative tasks and known failures. Keep the previous version available. Updating instructions is not evidence that the underlying model learned from an incident.

Release and measure the result

A canary release exposes a limited population to a candidate version. Compare candidate and control signals, and define when to expand or stop. Sparse traffic or different workloads can make the comparison inconclusive. Canary guidance.

The fictional team records a baseline from a fixed synthetic workload. It tests the correction, releases within an approved boundary, and checks comparable production periods. If evidence remains insufficient, it records uncertainty rather than declaring a gain.

Measure repeated manual work as well. Automation can reduce toil, but it also needs maintenance and failure handling. Include these costs when judging the result. Toil guidance.

Make the feedback record usable

Use these fields for the exercise: observation and version; baseline; proposed cause; acceptance criteria; regression checks; change and review; release boundary; measured result; owner and next review.

Taiga Maintaining connects repository findings to remediation work. Initiatives connect an intended change to planning and delivery. These provide parts of an evidence chain. Your service owner must still verify deployment and the operational result. Maintaining, Initiatives.

A mature software factory connects this work across products. Keep the decision rights and evaluation criteria visible as automation increases. The final proof is a better verified service, not a larger count of generated changes.

Do the exercise

Complete the feedback record in this lesson for the fictional memory leak. Define a baseline, acceptance test, regression checks, release boundary, production measurement, and owner. Add a rule for rejecting a faster but less correct export.

Download worksheet (Markdown)

Check your understanding

An agent lowers export latency by omitting authorization checks. The speed metric improves. Has the system improved?

Sources & further reading

Related reading from Taiga