Path 02Lesson 3 / 6

Use tests as evidence

Choose checks that can reject the wrong behavior. Review generated tests as carefully as generated implementation.

Practitioner11 minReviewed

Published by How we write

What you will learn

  • Connect each important requirement to a meaningful check.
  • Distinguish unit, integration, and end-to-end evidence.
  • Detect a test that repeats the same incorrect assumption as the implementation.

Start with the requirement

Tests are evidence for specific claims. A successful test run does not establish every property of the software. Before requesting tests, identify the behavior that matters and the defect each check should detect.

For a fictional organization export, the main requirement is data isolation. A user in organization A must not receive records from organization B. A test that checks only a successful download does not establish this requirement.

Ask the agent to explain the relationship between the requirement and the assertion. This makes missing cases easier to detect before the test suite becomes large.

Select the appropriate test scope

A unit test can check a small transformation quickly. An integration test can check how components work together. An end-to-end test can check an important user sequence through the deployed or representative application.

Use the narrowest scope that supplies the required evidence. A formatter does not need a full browser test for every input. An authorization boundary may need a real route and data access path. A critical browser interaction needs evidence about the rendered interface.

ClaimExample evidence
CSV output escapes a quote correctlyUnit test with a quote in a field
Another organization cannot read the exportIntegration test through real authorization
A keyboard user can start the exportBrowser test and manual keyboard review
A failed export gives a useful errorFailure-path check at the relevant interface

No fixed percentage of test types fits every system. Choose based on the failure you need to detect and the cost of maintaining the check.

Avoid a shared incorrect assumption

An agent can write implementation and tests from the same misunderstanding. Both can agree while the requirement remains unmet.

Suppose the implementation filters records by the organization ID supplied in the request. The test uses the same ID for the signed-in user and the request. It passes. The missing case is a user who requests a different organization’s ID.

Add that case through the actual trusted identity and authorization path. A mock that always returns “allowed” cannot establish tenant isolation. It only establishes behavior after authorization succeeds.

Verify that the test can fail

For a known defect, run the new regression test against the defective version in an isolated branch. Confirm that it fails for the intended reason. Then apply the correction and rerun it.

A test that fails because a fixture cannot load is not yet evidence about the business behavior. Inspect the failure, not only the exit code.

For broader changes, mutation testing can help assess whether selected code changes cause test failures. It has a cost and does not replace requirement review. Use it where the additional evidence supports a consequential decision.

Keep the evidence connected to the change

Run the relevant checks on the final revision. Record skipped checks and their reasons. A result from an earlier commit may no longer apply after a review correction.

Keep tests understandable. Prefer an explicit setup and assertion to a large helper that hides the important condition. Remove redundant checks when they add maintenance cost without detecting a different failure.

The reviewer should be able to state what the tests establish and what remains uncertain. That explanation is more useful than a large test count.

Do the exercise

Select one generated test. State the requirement it checks. Temporarily introduce the relevant defect in an isolated branch. Confirm that the test fails for the intended reason, then restore the code. Record what the test still does not cover.

Download worksheet (Markdown)

Check your understanding

A generated test mocks the authorization function to always allow access. What does a passing result establish?

Sources & further reading

Related reading from Taiga