Measure useful progress
Measure completed work, review effort, and rework. Do not use generated code volume as a measure of value.
Published by TaigaHow we write
What you will learn
- Distinguish activity from a useful outcome.
- Include preparation, review, and correction in a time comparison.
- Recognize the limits of a productivity claim.
Define the outcome before the metric
An AI tool can produce code quickly. The useful outcome is a change that meets a user need at the required quality level. These are different measurements.
Generated lines, accepted suggestions, and agent runs describe activity. They can help you understand tool use. They do not establish that a service improved or that the team delivered useful work sooner.
Start with one question. For example: “Does this workflow reduce the total effort needed to complete a small maintenance task?” Define completion before you collect results. Include the required tests, review, and documentation.
Count the whole task
Consider a fictional change to a report filter. Without AI, implementation takes 60 minutes and review takes 10 minutes. With AI, implementation takes 30 minutes and review takes 45 minutes.
Implementation is faster. The measured effort for these phases rises from 70 to 75 minutes. Neither result includes preparation, later corrections, or defects after release. Keep those limits visible.
| Phase | Example without AI | Example with AI |
|---|---|---|
| Implementation | 60 minutes | 30 minutes |
| Review | 10 minutes | 45 minutes |
| Measured total | 70 minutes | 75 minutes |
These figures illustrate a calculation. They are not research results or a forecast for your team. The review increase might reflect a larger diff, unfamiliar code, or a missing requirement. Investigate the cause before changing the tool policy.
Separate effort from elapsed time
Effort measures the time people spend on work. Elapsed time includes waiting. An agent can execute checks while a developer does another task. Do not count the same human time twice. Also record how long the change waits for review or an environment.
A workflow can reduce effort without reducing delivery time. This can happen when an approval queue determines the completion date. The saved effort may still have value, but the organization needs a separate decision about how to use it.
Ask developers whether the workflow helps them understand the system and maintain focus. Treat these responses as experience data. Do not convert a feeling of speed into a verified percentage improvement.
Read research within its limits
METR reported a slowdown in a specific early-2025 study of experienced open-source developers. The study did not establish an effect for all developers or tasks. Its February 2026 update described selection effects and measurement problems in a later experiment.
The useful lesson is about measurement. Tool versions, task selection, quality requirements, and participant behavior can change the result. Do not use one historical percentage as a permanent rule for AI development.
DORA’s 2025 research also directs attention to the organization around the tools. A team needs effective development practices to convert tool capability into useful delivery outcomes.
Make a small, repeatable comparison
Use representative tasks and the same completion criteria. Record the model and tool versions. Include unsuccessful attempts and review effort. Compare several tasks rather than selecting the best demonstration.
Report the range and the main limitations. If a change reduces effort but increases defects, investigate before expanding it. If the result is mixed, narrow the recommendation to the task types that have useful evidence.
A good measurement supports a specific next decision. It does not need to prove that AI is universally good or bad.
Do the exercise
Select five comparable completed tasks. Record preparation, implementation, review, correction, and waiting time. Record defects separately. Compare the total effort and elapsed time. Note differences in task difficulty, people, and tool versions before drawing a conclusion.
Download worksheet (Markdown)Check your understanding
Sources & further reading
- METR: Early-2025 developer productivity study ↗
- METR: February 2026 study update and measurement limitations ↗
- DORA: 2025 research report ↗
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.