Path 04Lesson 10 / 10

Measure the delivery system

Combine delivery flow, instability, service outcomes, and effort. Use explicit definitions when evaluating AI’s effect.

Practitioner10 minReviewed

Published by How we write

What you will learn

  • Distinguish delivery performance from code-generation activity.
  • Interpret a metric using its event definitions and scope.
  • Use measurements to select an improvement rather than rank individuals.

Begin with the decision you need to make

A team wants to know whether AI improves delivery. Counting generated lines answers a different question. Define the useful outcome and the quality conditions before selecting a metric.

For a fictional export service, the desired outcome is reliable delivery of accepted changes with less total effort. Record preparation, implementation, review, correction, and waiting. Include changes that failed or were abandoned.

Use one service with a clear boundary. Combining an experimental website and a critical payment service can produce a number that explains neither. Describe the context before comparing periods or teams.

Use current definitions

DORA’s current delivery model contains five metrics. Their scope is delivery performance, not the value of every feature or an individual’s contribution. DORA metric definitions.

MetricMeasurement focus
Change lead timeCommit to production
Deployment frequencyProduction deployment rate
Failed deployment recovery timeRecovery after a failed deployment
Change fail rateDeployments requiring immediate intervention
Deployment rework rateUnplanned deployments caused by production incidents

A dashboard may use another definition. Read it before interpreting the result. Taiga’s current deployment documentation describes four reported metrics derived from provider deployment records. Its recovery measure uses a subsequent successful deployment. That is not a complete record of every production incident. Taiga definitions.

Inspect a fictional change sequence

Suppose a service makes twelve deployments in a month. Eight deliver planned changes. Four repair problems from earlier releases. The count is twelve, but the composition matters.

In the next month, the team makes ten deployments: nine planned changes and one repair. Fewer deployments can coexist with more useful work. These figures illustrate interpretation; they are not a performance benchmark.

Also inspect the distribution. One long review wait can disappear inside an average. A recovery measure from a single failure is weak evidence for future reliability. Report the number of observations and material exceptions.

Pair flow with consequences

Use service signals to check whether delivery changes affect users. A faster pipeline is not sufficient if exports fail more often. Use an appropriate SLO or another clearly defined outcome measure. SLO guidance.

Review effort and rework help explain the result. If AI shortens implementation but produces large diffs, review may become the constraint. If environments take days to obtain, faster coding may have little effect on elapsed delivery time.

Choose one improvement that addresses the observed constraint. For example, provide a supported test environment or reduce the size of a change. Define a quality balancing metric so the team can detect an apparent speed gain caused by weaker checks.

Keep measurement useful

Avoid individual rankings based on PR counts or generated code. These measures can reward splitting work artificially, avoiding difficult maintenance, or shifting review effort to colleagues.

Review the result with the people responsible for the complete service. Record what changed in the tool, work mix, team, and environment. Treat a before-and-after comparison as evidence with limitations, not automatic proof of causation.

The purpose is a better next decision. A small, trustworthy measurement that leads to a verified improvement is more useful than a large dashboard without agreed meaning.

Practice with ten changes

This separate fictional dataset records ten planned changes. All times are UTC on the date shown. An empty correction field means no correction was recorded in this dataset.

Change / dateWork startsCode readyReview startsAcceptedReleasedCorrected
C01 · 2026-09-1408:0008:4509:1509:3010:00—
C02 · 2026-09-1409:0009:3012:0012:2013:00—
C03 · 2026-09-1508:0009:0009:1509:4010:0015:00
C04 · 2026-09-1510:0010:3010:4511:0011:15—
C05 · 2026-09-1608:0009:0013:0013:3014:00—
C06 · 2026-09-1610:0011:0011:3012:0012:15—
C07 · 2026-09-1708:0008:3009:0009:2009:30—
C08 · 2026-09-1710:0010:4511:0011:3014:30—
C09 · 2026-09-1808:0008:3009:0009:3010:00—
C10 · 2026-09-1809:0009:3010:0010:3011:0014:00

Compare the time from code ready to review start, then from acceptance to release. Identify the longest visible wait. Investigate its cause before calling it avoidable. These timestamps do not measure active effort or identify when an incident began. A correction release alone cannot establish failed deployment recovery time.

Download the fictional dataset (CSV)

Check the waiting times

Check your interpretation: C05 waits four hours for review. C08 waits three hours after acceptance before release. The dataset does not explain those waits. Ask about capacity, working hours, release policy, and dependencies.

Do the exercise

Use the ten-change dataset in this lesson. Define a deployment, a failed change, and a recovery event. Find the longest visible wait and state what would establish its cause. Propose an improvement and a measure that would reveal worse quality.

Download worksheet (Markdown)

Check your understanding

Deployment frequency rises after introducing AI, while unplanned repair deployments also rise. What should you conclude?

Sources & further reading

Related reading from Taiga