# Choose a model with evidence

Taiga Learning · Worksheet
https://taiga.training/en/lessons/evaluate-models/

Use fictional or approved information. Do not put secrets in this worksheet.

## Learning objectives
- Create a small evaluation set from real task types.
- Separate model quality from the effect of tools and context.
- Record the conditions that require a new evaluation.

## Exercise
Create an evaluation sheet for three task types: a known defect, a small feature, and a repository explanation. Define acceptance criteria before comparing models. Include one failure case per task. Record the model version, context, tool permissions, attempts, cost, and review effort.

## Your response
- Scenario and scope:
- Assumptions and open questions:
- Proposed answer or decision, with reasons:

## Verify your response
| Claim or criterion | Evidence or test | Result or gap | Owner |
| --- | --- | --- | --- |
| | | | |
| | | | |
| | | | |

## Next action
- Action, owner, and date:
- When will you review this response?

## Principle to retain
Select a model for a defined task and environment. Recheck the decision when the model, tools, or requirements change.

## Sources
- [NIST: Generative AI Profile](https://doi.org/10.6028/NIST.AI.600-1)
- [METR: Developer productivity study methodology](https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/)

This worksheet supports learning. Completing it does not itself authorize a production change.
