AI evals: how to test an AI system on your own data before and after every change
An AI eval, short for evaluation, is a repeatable test of an AI system: a set of real past jobs with the right answers, run through the system and scored against pass rules. You run it before launch and again after every change to the prompt, the model or a connected system, so you can see whether the change helped or broke something. This guide shows how a smaller manufacturer or distributor can build one from its own records, which grading methods to use, and which tools exist.

What an AI eval is
People use the terms AI evals, LLM evaluation and AI evaluation for the same practice: checking what an AI system produces against criteria you set, on inputs you choose. An LLM, or large language model, is the kind of AI behind ChatGPT and Claude. An eval has four parts:
- A test set is a list of real inputs, such as past purchase orders or customer emails, each with the answer your team gave.
- Pass rules, also called graders, decide whether each output is right. Some are code, some use a second AI model, and some are a person.
- A run sends every case through the system and records each output and its grade.
- A comparison sets the new run against the last one, case by case, to catch regressions. A regression is a case that used to pass and now fails.
An eval on your own data answers one question: whether the system gets your jobs right. Anthropic’s guide to building evaluations puts this first among its design principles: “Design evals that mirror your real-world task distribution.”
Step 1: Decide what a right answer means
Anthropic’s guide starts with success criteria that are specific, measurable, achievable and relevant. Its example of a weak criterion is “good performance,” and it suggests naming the exact result you want instead. For a system at a plant or a distributor, that means naming the fields and judgments that must be right:
- Order entry from email: the customer, every part number, quantity, unit price and ship date match what your team keyed.
- Quote drafting: the material, quantity and finish match the request, and the draft names the past jobs it drew on.
- Answers from documents: each answer cites the document it came from, and the cited passage supports it. RAG explains how that kind of system finds its sources.
Decide at the same time which mistakes are unacceptable. A wrong unit price on a customer order is a different kind of error from an awkward sentence in a draft email, and the pass rules should treat it differently.
Step 2: Build the test set from real past jobs
Pull 30 to 50 recent cases to start, each with the answer your team gave, and grow the set to 100 or more as the system goes into production, the range How to build an AI agent uses. Anthropic’s guide favours more cases over fewer: “More questions with slightly lower signal automated grading is better than fewer questions with high-quality human hand-graded evals.”
Include the awkward cases on purpose. Anthropic lists edge cases such as “Irrelevant or nonexistent input data.” At a manufacturer or distributor, that means scanned faxes, handwritten changes, two orders in one email, a part number that has been replaced, and a price that does not match the quote. Keep the set in files your company owns, give each version a date, and remove personal information the test does not need.
One test case, laid out
This generic example shows the fields a single case in an order entry test set might carry.
| Field | What it holds | Example |
|---|---|---|
| Input | The file or message the system receives | A customer purchase order as a two-page PDF |
| Expected output | The answer your team gave | Customer account, four order lines with part number, quantity, price and ship date |
| Pass rule | How each field is graded | Part numbers, quantities and prices must match exactly; a missing ship date must be flagged |
| Tags | Labels for sorting results | Scanned, handwritten change, replaced part number |
| Source and date | Where the case came from | Order mailbox, second quarter, reviewed by the order desk |
Step 3: Write the pass rules
Anthropic’s guide describes three ways to grade, and says to choose “the fastest, most reliable, most scalable method” that fits:
| Grading method | How Anthropic describes it | Where it fits at a plant or distributor |
|---|---|---|
| Code-based grading | “Fastest and most reliable, extremely scalable,” but lacks nuance for complex judgments | Part numbers, quantities, prices, dates and totals |
| LLM-based grading | “Fast and flexible, scalable and suitable for complex judgment”; test it for reliability before you scale it | Whether a draft email answers the question, or whether an answer is supported by its source |
| Human grading | “Most flexible and high quality, but slow and expensive” | A weekly sample of live work, and cases the other methods disagree on |
LLM-based grading means a second model reads the output and grades it against a rubric, a written list of what a right answer contains. Anthropic’s tips are to write detailed rubrics, to ask for a specific verdict such as correct or incorrect or a score from 1 to 5, and to have the grader reason before it scores. Its code samples note that it is “generally best practice to use a different model to evaluate than the model used to generate the evaluated output.”
Set a pass mark for each kind of case, and a stricter rule for fields that cost money when they are wrong. A system can pass most of its cases and still fail on the one field that matters, so report results by field and by tag as well as overall.
Build the first test set together
Tell Derik which job you want to test and where the past examples live. He will tell you how many cases you need and how to grade them.
Start a conversationStep 4: Rerun the eval after every change
Rerun the full test set whenever one of these changes:
- The prompt, meaning the instructions the system follows.
- The model. Providers publish retirement dates for their models, as Anthropic does on its model deprecations page, and a retired model has to be replaced.
- A connected system, such as an ERP upgrade, a new field or a new connector.
- The inputs, for example a large customer that switches to a new purchase order layout.
Compare the new run with the last one case by case. An overall score can hold steady while ten cases that used to pass now fail and ten others start passing, so look at each regression before you ship the change. Run the set on a schedule as well, so a change nobody announced, such as a customer’s new purchase order layout, shows up in the results.
Step 5: Review a sample of live work
A test set covers the cases you already know about, and live work brings new ones. OpenAI’s practical guide to building agents says human intervention is “especially important early in deployment, helping identify failures, uncover edge cases.” Have a person review a sample of live outputs each week, and add every case a reviewer corrects to the test set with the corrected answer.
When a person already approves each action, as human in the loop describes, every edit they make is a new test case. Record what changed and why.
Tools for AI evals
You can run a small eval with a spreadsheet and a script. The tools below add test runners, graders and result views. Each description comes from the vendor’s own page as it stood on September 28, 2026.
| Tool | What the vendor says it is | Where it runs |
|---|---|---|
| promptfoo | “An open-source CLI and library for evaluating and red-teaming LLM apps,” with test cases defined in a configuration file (promptfoo) | On your own machine or in your build pipeline; promptfoo says “The evals run on your machine and talk directly with the LLM” |
| Anthropic’s evaluation guide | Guidance on success criteria, eval design and grading, with code samples (Anthropic) | Wherever you run the code |
| Amazon Bedrock AgentCore Evaluations | 13 built-in evaluators on common quality dimensions, plus custom evaluators (AWS) | Your AWS account |
| Google Model Evaluation | A service in the Gemini Enterprise Agent Platform for “objective, data-driven assessment of generative AI models” (Google) | Your Google Cloud project |
| OpenAI Evals | OpenAI’s hosted evals product, which OpenAI is retiring (OpenAI) | OpenAI’s platform, until the shutdown |
OpenAI announced the deprecation of its Evals platform on June 3, 2026. Existing evals become read-only on October 31, 2026, and the Evals dashboard and API are scheduled to shut down on November 30, 2026. OpenAI’s cookbook entry, Moving from OpenAI Evals to Promptfoo, says OpenAI “is winding down the Evals product and recommends Promptfoo” for continuing that work. If you have evals there, export the test data before the shutdown.
Tracing tools record every model call and tool call in a live system, which helps you find the cases to add. How to build an AI agent names several.
The test set is the part of an eval that took your team’s time to build. Store the cases, expected answers and pass rules in plain files your company owns, so they move with you when a tool or a model changes.
Questions people ask
What are AI evals?
What is the difference between LLM evaluation and AI evaluation?
How many test cases does an AI eval need?
Can an AI model grade another AI model's work?
Is OpenAI Evals being shut down?
How often should we run AI evals?
How ThriveAI helps
ThriveAI is an AI engineering company in Ottawa. It builds private AI systems on the client’s own data for manufacturers and distributors in Ontario and Quebec. Derik Lawlis, the founder, leads every project and stays close to the build. In a ThriveAI build, a test set of labelled past cases measures the system from the first pilot and runs on every prompt or model change, as How to build an AI agent describes.
The platform is designed to keep each client’s data on its own server in Canada. You choose the model: one that runs on that server, or a hosted model under a written zero data retention agreement, under which the provider keeps no copy of a request or its answer. A hosted model may process requests outside Canada, so the contract names the model. A named person at your company approves every action before anything is sent or saved. About ThriveAI covers the company.