Path 05Lesson 1 / 8

Own the service after deployment

Define useful service signals, incident decisions, recovery, and maintenance. Keep operational responsibility visible after code generation ends.

Practitioner10 minReviewed

Published by How we write

What you will learn

  • Define a service signal from the user’s perspective.
  • Separate incident coordination from technical investigation.
  • Plan maintenance and recovery as continuing responsibilities.

Define the service that users depend on

Deployment makes software available. Operation keeps it useful as users, dependencies, traffic, and requirements change. A code generator does not remove this continuing work.

For a fictional customer export, users need more than a reachable page. They need the permitted records in the required format within an acceptable time. They also need the service to prevent access to another organization’s data.

Name the owner before release. Record who responds outside normal working hours if that is part of the service commitment. A supplier can perform some work, but the organization still needs a clear route for decisions and communication.

Choose signals that support action

A service-level indicator, or SLI, measures a defined property of service behavior. A service-level objective, or SLO, sets a target for that indicator over a stated period. Choose the target from user needs and operational capability.

Google’s SRE guidance explains this approach and the use of an error budget for reliability decisions. Do not copy another service’s target without checking its meaning. SLO guidance, example error-budget policy.

For the export, define what counts as a successful eligible request. Separate expected denials from system failures. Document exclusions so that a metric cannot improve merely by hiding difficult requests.

SignalWhat it helps detectImportant limit
Public availability checkService cannot be reachedDoes not verify a signed-in workflow
Export completion and latencyEligible requests fail or take too longRequires a precise success definition
Authorization denial checksA critical boundary regressesCovers the tested conditions
Resource and dependency signalsA likely internal causeDoes not alone describe user impact

Avoid logging complete exports to improve visibility. Collect the minimum information needed to diagnose the problem and protect access to it.

Prepare the incident response

Decide who coordinates, who investigates, and who communicates. These roles can be combined in a small team, but the responsibilities must remain clear. Keep a record of observations and actions.

Google’s incident-response guidance emphasizes coordination and communication alongside technical mitigation. A technically correct fix can still leave users uninformed or multiple responders making conflicting changes. Incident response.

An agent can summarize logs or compare hypotheses within approved data boundaries. It should not gain unrestricted production authority because the incident is urgent. Use a defined escalation route for exceptional access.

Exercise recovery and fund maintenance

Test the recovery procedure with representative fictional data. Identify what code rollback cannot undo, including deleted records or messages already sent. Record the time and information required to restore the service.

Assign ongoing work: dependency updates, access reviews, certificate renewal where applicable, capacity changes, and documentation corrections. A service without maintenance capacity accumulates obligations after its launch budget ends.

After an incident, select improvements that address the observed causes. Connect them to implementation and verification. This closes the lifecycle: operational evidence changes what the team specifies and builds next.

Do the exercise

Write a one-page operating note for the fictional customer export. Include one user-facing signal, its target, an alert recipient, a safe first response, a recovery limit, and a maintenance owner. State what the monitoring cannot detect.

Download worksheet (Markdown)

Check your understanding

An uptime check returns HTTP 200, but exports contain no records because authorization is broken. What does this show?

Sources & further reading

Related reading from Taiga