Path 05Lesson 5 / 8

Manage an incident from detection to recovery

Coordinate responders, contain impact, communicate uncertainty, and verify recovery. Convert the incident into owned improvements.

Practitioner11 minReviewed

Published by How we write

What you will learn

  • Assign incident coordination, technical work, and communication.
  • Choose containment from impact and available evidence.
  • Separate restored service from completed follow-up work.

Declare the incident from impact

An incident is an event that disrupts, degrades, or threatens the service enough to require a coordinated response. Your organization defines severity levels and escalation rules. Use user impact, affected data, duration, and scope to apply them.

Do not wait for a complete root-cause explanation before requesting help. A clear statement of observed impact is enough to start coordination. Treat a security suspicion separately from a confirmed conclusion.

Prepare the response route before release. Keep contact details, access procedures, runbooks, and communication channels available when the main service is unavailable. Exercise the route with a fictional incident.

Assign responsibilities before making competing changes

Incident coordination sets priorities and manages decisions. Technical responders investigate and mitigate. Communication keeps affected people informed. Google SRE describes these responsibilities as distinct roles. Small teams can combine roles, but must still cover the work. Incident response.

ResponsibilityImmediate question
Incident coordinatorWhat is the impact, current priority, and next decision?
Technical responderWhich authorized action can reduce impact, and how will we verify it?
Communication ownerWho needs an update, what is known, and when is the next update?
Service ownerWhich business tradeoffs and recovery criteria apply?
Security responseCould confidentiality, integrity, credentials, or evidence be affected?

Keep one shared timeline. Record the time, observation, action, actor, and result. Distinguish facts from hypotheses. Use a common time zone and note unreliable timestamps.

Work through a fictional incident

All times below are UTC. The organization names an incident coordinator when the export failure affects multiple customers.

TimeObservation or action
09:02Export failures exceed the service alert threshold
09:04On-call confirms failed jobs; incident coordination starts
09:07The team pauses new exports through an approved feature control
09:10A user reports records that may belong to another organization
09:12Security response joins; relevant logs and artifact identifiers are preserved
09:18The team restores a compatible previous version in a controlled rollout
09:25Synthetic exports succeed; access-boundary tests and disclosure investigation continue

A useful first update states the affected function, known scope, mitigation, and next update time. It does not promise a repair time without evidence. Avoid including customer records in the shared update.

At 09:10, the incident changes. Restoring successful exports is no longer sufficient. The team needs to assess possible disclosure, control access, preserve evidence, and involve the appropriate decision owners.

Mitigate without losing control

Use tested runbooks where they apply. Check preconditions before rollback, failover, or credential changes. A previous application version might not understand the current database schema. A regional failover can move the same corrupted data.

Let an AI assistant organize redacted evidence or compare hypotheses within approved boundaries. Responders must verify its conclusions. Logs and tickets are untrusted input, not authority to execute their contents.

Emergency access should have an authorized purpose, limited duration, and an audit record. Urgency does not make an agent’s suggested command correct.

Close recovery and follow-up separately

Verify the user workflow, data integrity, access boundaries, and monitoring freshness before declaring service recovery. Record any remaining restrictions. Keep a security investigation open if its questions remain unresolved.

Afterward, examine the conditions that made the incident possible. Assign concrete follow-up work with an owner and verification criteria. A blameless review seeks an accurate explanation and useful changes. It does not remove accountability for completing those changes. Postmortem practice.

Continue with security operations and closing the feedback loop.

Do the exercise

Use the fictional incident timeline in this lesson. Write the first situation update, name three response roles, and define two recovery checks. Identify one action that needs a security-response decision.

Download worksheet (Markdown)

Check your understanding

A rollback restores successful exports, but a user reports receiving another organization’s records. What happens next?

Sources & further reading

Related reading from Taiga