Path 01Lesson 6 / 6

Choose a model with evidence

Compare models on representative tasks, acceptance criteria, cost, and the operating constraints of your team.

Practitioner10 minReviewed

Published by How we write

What you will learn

  • Create a small evaluation set from real task types.
  • Separate model quality from the effect of tools and context.
  • Record the conditions that require a new evaluation.

Define the decision

A model comparison needs a specific use. A model that explains a small function well may not handle a large repository change equally well. A lower-cost model may meet the quality requirement for routine transformations. A difficult investigation may need greater reasoning capability.

Write the task and constraints first. Include permitted data, required tools, response time, and the maximum acceptable cost. Some constraints are mandatory. Do not average away a data handling restriction because another score is high.

Use representative tasks

Build a small evaluation set from work your team actually does. Remove sensitive data unless the evaluation environment is approved for that data. Include simple tasks, difficult tasks, and tasks where the correct response is to request missing information.

For a fictional reporting service, use a known date defect, a small filter feature, and an explanation of an authorization rule. Prepare the expected result before you run the comparison. Include a negative test that rejects the known defect.

Keep some tasks separate from prompt development. If you repeatedly tune the prompt on every example, the final score can overstate general performance. A separate set helps reveal whether the improved prompt works beyond the examples used to create it.

Keep the comparison fair

Record the exact model version, prompt, supplied context, tools, and permissions. Use equivalent starting states. If one model receives a complete repository and another receives one file, the result compares workflows as well as models.

Workflow comparisons can be useful. Label them correctly. An agent product includes more than a model: context selection, tools, execution limits, and recovery behavior can affect the result.

Use repeated runs when output variation matters. Record failed attempts rather than reporting only the best result. For subjective criteria, use a written rubric and more than one reviewer where practical.

Score quality before speed

First, check mandatory acceptance criteria. Does the change meet the requirement? Does it preserve access controls? Do the relevant tests pass? Can a reviewer understand the diff?

Then compare effort, elapsed time, and cost for acceptable results. Include retries and human review. A cheap response that needs repeated correction can be expensive at task level.

Evaluation fieldWhat to record
Task resultWhich acceptance criteria passed or failed
ScopeUnrequested changes or missing requirements
Human effortPreparation, review, and correction time
Execution costModel and tool costs, including retries
EvidenceVersion, input, output, checks, and reviewer notes

Public benchmarks can help identify candidates. They use specific task sets and scoring methods. Do not treat a benchmark score as a direct measurement of your team’s productivity.

Record the decision and its expiry conditions

The result can be a narrow recommendation. For example: “Use this model for small test additions in this repository, with the existing review requirements.” You do not need one model for every task.

State what would require another evaluation. Examples include a model version change, a different tool configuration, a new data category, or a persistent failure pattern. Keep a fallback for tasks that exceed the selected model’s capability.

The purpose of evaluation is to reduce uncertainty about a real decision. Avoid a permanent model competition that consumes more effort than the work it supports.

Do the exercise

Create an evaluation sheet for three task types: a known defect, a small feature, and a repository explanation. Define acceptance criteria before comparing models. Include one failure case per task. Record the model version, context, tool permissions, attempts, cost, and review effort.

Download worksheet (Markdown)

Check your understanding

Model A passes more public benchmark tasks. Model B performs better on your representative repository tasks. Which result should guide the decision?

Sources & further reading

Related reading from Taiga