# Set safe boundaries for self-healing

Taiga Learning · Worksheet
https://taiga.training/en/lessons/self-healing/

Use fictional or approved information. Do not put secrets in this worksheet.

## Learning objectives
- Distinguish self-healing from a permanent software correction.
- Define a bounded recovery policy and independent success checks.
- Recognize when automation must stop and escalate.

## Exercise
Design a recovery policy for the fictional export worker in this lesson. Specify its trigger, exclusions, allowed action, retry limit, cooldown, success check, and escalation owner. Test it against a database outage and an unknown data-integrity failure.

## Your response
- Scenario and scope:
- Assumptions and open questions:
- Proposed answer or decision, with reasons:

## Verify your response
| Claim or criterion | Evidence or test | Result or gap | Owner |
| --- | --- | --- | --- |
| | | | |
| | | | |
| | | | |

## Next action
- Action, owner, and date:
- When will you review this response?

## Principle to retain
Self-healing needs a defined failure, authorized action, measurable result, and stop condition. Repeating an action without recovery is another failure.

## Sources
- [Kubernetes: Self-Healing](https://kubernetes.io/docs/concepts/architecture/self-healing/)
- [AWS Builders’ Library: Timeouts, retries, and backoff with jitter](https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/)
- [Google SRE: Automation at Google](https://sre.google/sre-book/automation-at-google/)

This worksheet supports learning. Completing it does not itself authorize a production change.
