Observe the service and its users
Connect metrics, logs, and traces to service objectives. Design alerts, data boundaries, and checks for missing telemetry.
Published by TaigaHow we write
What you will learn
- Choose telemetry that answers a specific operational question.
- Distinguish a service symptom from an internal cause.
- Protect telemetry and detect missing or stale evidence.
Start with the question
Monitoring checks known conditions. Observability helps you investigate system behavior, including failures you did not predict. More dashboards do not automatically provide better answers.
For a fictional export service, start with a user question: can an authorized user receive the correct export within the agreed time? Then choose signals that support that question and help explain failures.
OpenTelemetry provides instrumentation and standards for telemetry. It can send signals to compatible backends. You still need storage, queries, access controls, retention, and people who act on the evidence. Observability primer.
Connect different forms of evidence
A metric measures a quantity over time. A log records an event. A trace connects related operations as a request moves through a system. A span represents one operation within a trace.
| Operational question | Example evidence | Limit to remember |
|---|---|---|
| How many eligible exports fail? | Failure count and eligible request count | A wrong denominator gives a misleading rate |
| What happened to one export? | Structured log with job ID, outcome, and version | Missing events leave gaps |
| Where did time go? | Trace across API, queue, worker, and database | Sampling and broken context propagation can hide work |
| What changed before the symptom? | Deployment and configuration records | Timing alone does not establish a cause |
For asynchronous work, preserve a safe correlation between the submitted job and worker execution. An HTTP 202 response can mean that work was accepted. It does not prove that the export finished.
Alert when action is needed
Define the SLI and its denominator before setting the SLO. For the example, count eligible exports completed correctly within the agreed duration. Define how long-running and abandoned jobs enter the measurement.
An error budget describes the permitted failure within the SLO window. A burn rate describes how quickly failures consume that budget. Google’s guidance uses multiple windows to balance timely detection and alert noise. Alerting on SLOs.
Page someone when the condition requires a timely action. Send lower-urgency work to a queue. Every alert needs an owner, impact description, investigation link, and response instruction. Review alerts that repeatedly lead to no action.
Do not use one generic threshold for every service. User impact, traffic, business hours, and response capacity affect the decision.
Protect the telemetry pipeline
Telemetry can contain personal data, tokens, request parameters, and confidential documents. Define allowed fields before collection. Restrict access and retention. Redact secrets before export to an external backend. Sensitive telemetry.
Do not use a customer email or unique job ID as a metric label. Unbounded labels increase the number of time series and can expose identifiers. Use controlled dimensions for metrics. Put approved correlation identifiers in access-controlled logs or traces.
Measure the pipeline itself. Check ingestion failures, dropped data, and the age of the latest observation. A flat error chart can mean no errors or no incoming telemetry. Show that distinction.
Investigate a concrete failure
The fictional service reports HTTP 202 for every request. Its queue age rises from seconds to 15 minutes. Worker logs show repeated database timeouts. Sampled traces place most worker time in database calls.
This evidence supports a focused investigation. It does not establish whether the cause is a query change, exhausted connections, or database capacity. Compare these hypotheses with the debugging method.
Taiga Monitoring provides a product-health view with availability and browser experience signals. It complements infrastructure and application observability; it does not replace those systems. Monitoring.
Do the exercise
For the fictional export in this lesson, define one SLI, one actionable alert, three allowed telemetry fields, and two prohibited fields. State how you would detect a broken telemetry pipeline.
Download worksheet (Markdown)Check your understanding
Sources & further reading
- OpenTelemetry: Observability primer ↗
- OpenTelemetry: Handling sensitive data ↗
- Google SRE: Alerting on SLOs ↗
- Taiga docs: Monitoring ↗
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.