Path 05Lesson 7 / 8

Set safe boundaries for self-healing

Automate known recovery actions with explicit authority, verification, and stop conditions. Separate runtime recovery from changing software.

Advanced12 minReviewed

Published by How we write

What you will learn

  • Distinguish self-healing from a permanent software correction.
  • Define a bounded recovery policy and independent success checks.
  • Recognize when automation must stop and escalate.

Recover a known condition

Self-healing automatically detects a defined failure and attempts an authorized recovery action. Restarting a failed process or replacing an unhealthy instance can be examples. The action must fit the failure and the service’s state model.

Kubernetes can replace failed workload instances and reconcile declared state. This does not correct faulty application logic or every storage failure. Infrastructure recovery and software correctness need different checks. Kubernetes self-healing.

Define the objective before the mechanism. Restoring an export means that eligible work completes correctly. A running container is only one precondition.

Separate three kinds of change

ChangeExampleRequired decision
Runtime recoveryReplace one failed stateless workerA preapproved recovery policy can authorize this
Software correctionFix the memory leak that stops the workerReview, tests, release controls, and production verification
Policy changeIncrease the permitted restart rate or access scopeExplicit approval from the policy owner

An agent may propose a correction after recovery. That proposal is a new software change. It must not inherit unlimited authority from the recovery controller.

The controller also must not edit its own success criteria when a check fails. Otherwise, the system can report improvement without improving the service.

Write the recovery policy before enabling it

The following policy is fictional. Its numbers illustrate design choices; they are not recommended defaults.

Policy fieldFictional export-worker rule
TriggerWorker heartbeat is absent for 90 seconds and queued work exists
PreconditionsAnother worker is healthy; dependency checks pass; no suspected compromise or integrity failure
Allowed actionReplace one worker using the currently approved artifact
State protectionJobs use durable storage and a verified idempotency key
LimitAt most two replacements in 15 minutes; never more than one at a time
CooldownWait five minutes after replacement before another attempt
SuccessA synthetic job completes correctly and the affected queue begins to drain
Stop and escalateAny precondition fails, the limit is reached, or success cannot be verified

Use a least-privileged identity. Log the policy version, trigger evidence, action, resource, and result. Provide an independent way to disable the controller. Define the human owner who receives an escalation.

Test failure paths as well as successful recovery

A retry can repeat a side effect. A worker might store a file and stop before acknowledging the job. Verify idempotency before allowing another execution. See the cloud native failure example.

Retries can also amplify an overloaded dependency. Use bounded attempts, timeouts, and appropriate backoff. Avoid synchronized retries across the fleet. AWS explains why backoff and jitter help reduce this amplification. Retry guidance.

Test the fictional policy against three cases. A single stopped worker should recover. A database outage should prevent repeated replacement. An uncertain integrity failure should stop automation and request a response decision.

Check missing telemetry too. The absence of a heartbeat can mean a failed worker or a failed collection path. The controller needs evidence sufficient for its action, not confidence in an AI explanation.

Measure whether the policy helps

Record verified recoveries, unsuccessful attempts, escalations, duplicate work, and time spent with user impact. Compare these with the previous operating method under similar conditions.

Retain the underlying defect as engineering work. Repeatedly restarting a leaking process can reduce immediate impact while the leak continues. Continue with self-improvement to connect the observation to a durable correction.

Do the exercise

Design a recovery policy for the fictional export worker in this lesson. Specify its trigger, exclusions, allowed action, retry limit, cooldown, success check, and escalation owner. Test it against a database outage and an unknown data-integrity failure.

Download worksheet (Markdown)

Check your understanding

The controller has restarted a worker twice. The queue continues to grow, and the database is unreachable. What should the policy do?

Sources & further reading

Related reading from Taiga