AI evals: how to test an AI system on your own data before and after every change

An AI eval, short for evaluation, is a repeatable test of an AI system: a set of real past jobs with the right answers, run through the system and scored against pass rules. You run it before launch and again after every change to the prompt, the model or a connected system, so you can see whether the change helped or broke something. This guide shows how a smaller manufacturer or distributor can build one from its own records, which grading methods to use, and which tools exist.

Rows of machined brass blocks, each with a milled pocket and two bolt holes

What an AI eval is

People use the terms AI evals, LLM evaluation and AI evaluation for the same practice: checking what an AI system produces against criteria you set, on inputs you choose. An LLM, or large language model, is the kind of AI behind ChatGPT and Claude. An eval has four parts:

An eval on your own data answers one question: whether the system gets your jobs right. Anthropic’s guide to building evaluations puts this first among its design principles: “Design evals that mirror your real-world task distribution.”

Step 1: Decide what a right answer means

Anthropic’s guide starts with success criteria that are specific, measurable, achievable and relevant. Its example of a weak criterion is “good performance,” and it suggests naming the exact result you want instead. For a system at a plant or a distributor, that means naming the fields and judgments that must be right:

Decide at the same time which mistakes are unacceptable. A wrong unit price on a customer order is a different kind of error from an awkward sentence in a draft email, and the pass rules should treat it differently.

Step 2: Build the test set from real past jobs

Pull 30 to 50 recent cases to start, each with the answer your team gave, and grow the set to 100 or more as the system goes into production, the range How to build an AI agent uses. Anthropic’s guide favours more cases over fewer: “More questions with slightly lower signal automated grading is better than fewer questions with high-quality human hand-graded evals.”

Include the awkward cases on purpose. Anthropic lists edge cases such as “Irrelevant or nonexistent input data.” At a manufacturer or distributor, that means scanned faxes, handwritten changes, two orders in one email, a part number that has been replaced, and a price that does not match the quote. Keep the set in files your company owns, give each version a date, and remove personal information the test does not need.

One test case, laid out

This generic example shows the fields a single case in an order entry test set might carry.

FieldWhat it holdsExample
InputThe file or message the system receivesA customer purchase order as a two-page PDF
Expected outputThe answer your team gaveCustomer account, four order lines with part number, quantity, price and ship date
Pass ruleHow each field is gradedPart numbers, quantities and prices must match exactly; a missing ship date must be flagged
TagsLabels for sorting resultsScanned, handwritten change, replaced part number
Source and dateWhere the case came fromOrder mailbox, second quarter, reviewed by the order desk

Step 3: Write the pass rules

Anthropic’s guide describes three ways to grade, and says to choose “the fastest, most reliable, most scalable method” that fits:

Grading methodHow Anthropic describes itWhere it fits at a plant or distributor
Code-based grading“Fastest and most reliable, extremely scalable,” but lacks nuance for complex judgmentsPart numbers, quantities, prices, dates and totals
LLM-based grading“Fast and flexible, scalable and suitable for complex judgment”; test it for reliability before you scale itWhether a draft email answers the question, or whether an answer is supported by its source
Human grading“Most flexible and high quality, but slow and expensive”A weekly sample of live work, and cases the other methods disagree on

LLM-based grading means a second model reads the output and grades it against a rubric, a written list of what a right answer contains. Anthropic’s tips are to write detailed rubrics, to ask for a specific verdict such as correct or incorrect or a score from 1 to 5, and to have the grader reason before it scores. Its code samples note that it is “generally best practice to use a different model to evaluate than the model used to generate the evaluated output.”

Set a pass mark for each kind of case, and a stricter rule for fields that cost money when they are wrong. A system can pass most of its cases and still fail on the one field that matters, so report results by field and by tag as well as overall.

Build the first test set together

Tell Derik which job you want to test and where the past examples live. He will tell you how many cases you need and how to grade them.

Start a conversation

Step 4: Rerun the eval after every change

Rerun the full test set whenever one of these changes:

Compare the new run with the last one case by case. An overall score can hold steady while ten cases that used to pass now fail and ten others start passing, so look at each regression before you ship the change. Run the set on a schedule as well, so a change nobody announced, such as a customer’s new purchase order layout, shows up in the results.

Step 5: Review a sample of live work

A test set covers the cases you already know about, and live work brings new ones. OpenAI’s practical guide to building agents says human intervention is “especially important early in deployment, helping identify failures, uncover edge cases.” Have a person review a sample of live outputs each week, and add every case a reviewer corrects to the test set with the corrected answer.

When a person already approves each action, as human in the loop describes, every edit they make is a new test case. Record what changed and why.

Tools for AI evals

You can run a small eval with a spreadsheet and a script. The tools below add test runners, graders and result views. Each description comes from the vendor’s own page as it stood on September 28, 2026.

ToolWhat the vendor says it isWhere it runs
promptfoo“An open-source CLI and library for evaluating and red-teaming LLM apps,” with test cases defined in a configuration file (promptfoo)On your own machine or in your build pipeline; promptfoo says “The evals run on your machine and talk directly with the LLM”
Anthropic’s evaluation guideGuidance on success criteria, eval design and grading, with code samples (Anthropic)Wherever you run the code
Amazon Bedrock AgentCore Evaluations13 built-in evaluators on common quality dimensions, plus custom evaluators (AWS)Your AWS account
Google Model EvaluationA service in the Gemini Enterprise Agent Platform for “objective, data-driven assessment of generative AI models” (Google)Your Google Cloud project
OpenAI EvalsOpenAI’s hosted evals product, which OpenAI is retiring (OpenAI)OpenAI’s platform, until the shutdown

OpenAI announced the deprecation of its Evals platform on June 3, 2026. Existing evals become read-only on October 31, 2026, and the Evals dashboard and API are scheduled to shut down on November 30, 2026. OpenAI’s cookbook entry, Moving from OpenAI Evals to Promptfoo, says OpenAI “is winding down the Evals product and recommends Promptfoo” for continuing that work. If you have evals there, export the test data before the shutdown.

Tracing tools record every model call and tool call in a live system, which helps you find the cases to add. How to build an AI agent names several.

Keep the test set yours

The test set is the part of an eval that took your team’s time to build. Store the cases, expected answers and pass rules in plain files your company owns, so they move with you when a tool or a model changes.

Questions people ask

What are AI evals?
AI evals are repeatable tests of an AI system. You run a set of real past jobs with known answers through the system, grade each output against pass rules, and compare the results with the last run. Teams run evals before launch and after every change to the prompt, the model or a connected system.
What is the difference between LLM evaluation and AI evaluation?
The terms describe the same practice. LLM evaluation tests a system built on a large language model, the kind of AI behind ChatGPT and Claude, and AI evaluation is the broader term. In both cases you check the system's outputs against criteria you set, on inputs you choose.
How many test cases does an AI eval need?
Start with 30 to 50 real past cases, each with the answer your team gave, and grow the set to 100 or more as the system goes into production. Anthropic's guidance favours more cases with automated grading over fewer cases graded by hand. Include awkward cases on purpose, such as scanned documents and mismatched prices.
Can an AI model grade another AI model's work?
Yes. Anthropic describes LLM-based grading as fast, flexible and suitable for complex judgment, and says to test it for reliability before you scale it. Give the grading model a detailed rubric, ask for a specific verdict, and use a different model from the one that produced the output.
Is OpenAI Evals being shut down?
Yes. OpenAI announced the deprecation of its Evals platform on June 3, 2026. Existing evals become read-only on October 31, 2026, and the Evals dashboard and API are scheduled to shut down on November 30, 2026. OpenAI recommends promptfoo and has published a migration guide.
How often should we run AI evals?
Run the full test set before launch, after every change to the prompt, the model or a connected system, and on a regular schedule to catch changes in the inputs, such as a new form from a customer. After launch, review a sample of live work each week and add every corrected case to the test set.

How ThriveAI helps

ThriveAI is an AI engineering company in Ottawa. It builds private AI systems on the client’s own data for manufacturers and distributors in Ontario and Quebec. Derik Lawlis, the founder, leads every project and stays close to the build. In a ThriveAI build, a test set of labelled past cases measures the system from the first pilot and runs on every prompt or model change, as How to build an AI agent describes.

The platform is designed to keep each client’s data on its own server in Canada. You choose the model: one that runs on that server, or a hosted model under a written zero data retention agreement, under which the provider keeps no copy of a request or its answer. A hosted model may process requests outside Canada, so the contract names the model. A named person at your company approves every action before anything is sent or saved. About ThriveAI covers the company.

Contact

Test the AI you bought on your own past work

Tell Derik which AI system you run or plan to run and which job it does. He will tell you how to build a test set from your past work and how to grade it.

Prefer to talk? Book a meeting.

Your message goes to Derik Lawlis, the founder.