Also called Evals.
Evals, short for evaluations, are a test set for an AI system. You pick real questions, record the right answer, and run them whenever the prompt, the documents, or the model changes. The score tells you whether this week's change helped or quietly hurt.
A demo is one lucky question. An eval is the list you wish you had run before the last incident. Include the awkward cases: missing data, angry users, and questions the system should refuse. Agree in advance what good enough means. A team that picks the score after seeing it will always pass.
Think of it this way: Evals are the quality control station at the end of the production line. Without them, you are shipping product and hoping customers find the defects instead of you.
A team builds evals using 200 real customer queries with expected answers. After every model change or prompt update, the eval suite runs automatically and no update ships if it drops below the quality threshold.
A support team keeps forty real tickets with the reply a lead marked correct. Every Friday the new bot version answers all forty. When a model update drops the score on billing tickets, they hold the update. The live customers never see that version.
Before launch, after every significant change, and on a recurring schedule as the model or data drifts. Evals turn 'it seems to be working' into a measurable, defensible fact. Evals only catch what they were designed to test. A strong eval suite for one task provides no coverage for a new task the system has been extended to handle. Expand evals when the system expands.
RaftLabs treats this as part of the build: a source on the answer, a test set, and a record of what the system did. The launch is the start of that work, not the end. The related work on our side is AI consulting.
This sits with the other reliability & risk terms on the glossary. Why a confident answer can still be wrong, and how you catch it. Worth reading next: Hallucination, Guardrails, and Model Drift.