Set safe boundaries for self-healing
Automate known recovery actions with explicit authority, verification, and stop conditions. Separate runtime recovery from changing software.
Published by TaigaHow we write
What you will learn
- Distinguish self-healing from a permanent software correction.
- Define a bounded recovery policy and independent success checks.
- Recognize when automation must stop and escalate.
Recover a known condition
Self-healing automatically detects a defined failure and attempts an authorized recovery action. Restarting a failed process or replacing an unhealthy instance can be examples. The action must fit the failure and the service’s state model.
Kubernetes can replace failed workload instances and reconcile declared state. This does not correct faulty application logic or every storage failure. Infrastructure recovery and software correctness need different checks. Kubernetes self-healing.
Define the objective before the mechanism. Restoring an export means that eligible work completes correctly. A running container is only one precondition.
Separate three kinds of change
| Change | Example | Required decision |
|---|---|---|
| Runtime recovery | Replace one failed stateless worker | A preapproved recovery policy can authorize this |
| Software correction | Fix the memory leak that stops the worker | Review, tests, release controls, and production verification |
| Policy change | Increase the permitted restart rate or access scope | Explicit approval from the policy owner |
An agent may propose a correction after recovery. That proposal is a new software change. It must not inherit unlimited authority from the recovery controller.
The controller also must not edit its own success criteria when a check fails. Otherwise, the system can report improvement without improving the service.
Write the recovery policy before enabling it
The following policy is fictional. Its numbers illustrate design choices; they are not recommended defaults.
| Policy field | Fictional export-worker rule |
|---|---|
| Trigger | Worker heartbeat is absent for 90 seconds and queued work exists |
| Preconditions | Another worker is healthy; dependency checks pass; no suspected compromise or integrity failure |
| Allowed action | Replace one worker using the currently approved artifact |
| State protection | Jobs use durable storage and a verified idempotency key |
| Limit | At most two replacements in 15 minutes; never more than one at a time |
| Cooldown | Wait five minutes after replacement before another attempt |
| Success | A synthetic job completes correctly and the affected queue begins to drain |
| Stop and escalate | Any precondition fails, the limit is reached, or success cannot be verified |
Use a least-privileged identity. Log the policy version, trigger evidence, action, resource, and result. Provide an independent way to disable the controller. Define the human owner who receives an escalation.
Test failure paths as well as successful recovery
A retry can repeat a side effect. A worker might store a file and stop before acknowledging the job. Verify idempotency before allowing another execution. See the cloud native failure example.
Retries can also amplify an overloaded dependency. Use bounded attempts, timeouts, and appropriate backoff. Avoid synchronized retries across the fleet. AWS explains why backoff and jitter help reduce this amplification. Retry guidance.
Test the fictional policy against three cases. A single stopped worker should recover. A database outage should prevent repeated replacement. An uncertain integrity failure should stop automation and request a response decision.
Check missing telemetry too. The absence of a heartbeat can mean a failed worker or a failed collection path. The controller needs evidence sufficient for its action, not confidence in an AI explanation.
Measure whether the policy helps
Record verified recoveries, unsuccessful attempts, escalations, duplicate work, and time spent with user impact. Compare these with the previous operating method under similar conditions.
Retain the underlying defect as engineering work. Repeatedly restarting a leaking process can reduce immediate impact while the leak continues. Continue with self-improvement to connect the observation to a durable correction.
Do the exercise
Design a recovery policy for the fictional export worker in this lesson. Specify its trigger, exclusions, allowed action, retry limit, cooldown, success check, and escalation owner. Test it against a database outage and an unknown data-integrity failure.
Download worksheet (Markdown)Check your understanding
Sources & further reading
- Kubernetes: Self-Healing ↗
- AWS Builders’ Library: Timeouts, retries, and backoff with jitter ↗
- Google SRE: Automation at Google ↗
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.