Choose a model with evidence
Compare models on representative tasks, acceptance criteria, cost, and the operating constraints of your team.
Published by TaigaHow we write
What you will learn
- Create a small evaluation set from real task types.
- Separate model quality from the effect of tools and context.
- Record the conditions that require a new evaluation.
Define the decision
A model comparison needs a specific use. A model that explains a small function well may not handle a large repository change equally well. A lower-cost model may meet the quality requirement for routine transformations. A difficult investigation may need greater reasoning capability.
Write the task and constraints first. Include permitted data, required tools, response time, and the maximum acceptable cost. Some constraints are mandatory. Do not average away a data handling restriction because another score is high.
Use representative tasks
Build a small evaluation set from work your team actually does. Remove sensitive data unless the evaluation environment is approved for that data. Include simple tasks, difficult tasks, and tasks where the correct response is to request missing information.
For a fictional reporting service, use a known date defect, a small filter feature, and an explanation of an authorization rule. Prepare the expected result before you run the comparison. Include a negative test that rejects the known defect.
Keep some tasks separate from prompt development. If you repeatedly tune the prompt on every example, the final score can overstate general performance. A separate set helps reveal whether the improved prompt works beyond the examples used to create it.
Keep the comparison fair
Record the exact model version, prompt, supplied context, tools, and permissions. Use equivalent starting states. If one model receives a complete repository and another receives one file, the result compares workflows as well as models.
Workflow comparisons can be useful. Label them correctly. An agent product includes more than a model: context selection, tools, execution limits, and recovery behavior can affect the result.
Use repeated runs when output variation matters. Record failed attempts rather than reporting only the best result. For subjective criteria, use a written rubric and more than one reviewer where practical.
Score quality before speed
First, check mandatory acceptance criteria. Does the change meet the requirement? Does it preserve access controls? Do the relevant tests pass? Can a reviewer understand the diff?
Then compare effort, elapsed time, and cost for acceptable results. Include retries and human review. A cheap response that needs repeated correction can be expensive at task level.
| Evaluation field | What to record |
|---|---|
| Task result | Which acceptance criteria passed or failed |
| Scope | Unrequested changes or missing requirements |
| Human effort | Preparation, review, and correction time |
| Execution cost | Model and tool costs, including retries |
| Evidence | Version, input, output, checks, and reviewer notes |
Public benchmarks can help identify candidates. They use specific task sets and scoring methods. Do not treat a benchmark score as a direct measurement of your team’s productivity.
Record the decision and its expiry conditions
The result can be a narrow recommendation. For example: “Use this model for small test additions in this repository, with the existing review requirements.” You do not need one model for every task.
State what would require another evaluation. Examples include a model version change, a different tool configuration, a new data category, or a persistent failure pattern. Keep a fallback for tasks that exceed the selected model’s capability.
The purpose of evaluation is to reduce uncertainty about a real decision. Avoid a permanent model competition that consumes more effort than the work it supports.
Do the exercise
Create an evaluation sheet for three task types: a known defect, a small feature, and a repository explanation. Define acceptance criteria before comparing models. Include one failure case per task. Record the model version, context, tool permissions, attempts, cost, and review effort.
Download worksheet (Markdown)Check your understanding
Sources & further reading
Related reading from Taiga
Clearing this selection deletes all progress saved in this browser.
Progress stays in this browser. No account, no tracking.