How to build an AI agent evaluation set before you launch
Before you ship an AI agent, build a 20-case test set covering normal inputs, common exceptions, and forbidden actions, not just one good demo.
Also available in Korean: Read the Korean original
An AI agent can give different results for the same request, depending on the input file, tool state, permissions, and wording. One well-chosen demo case passing doesn't mean the agent is ready for real work. Real problems start with blank fields, duplicate customer records, unreadable attachments, and edge cases that need a human call.
Google Cloud's 2026 guide to production-ready AI agents notes that agents reason, act, and adapt. That means they need different testing, memory, orchestration, and security than traditional software. The 20 cases below aren't a guarantee for any specific tool. They're a practical template a small team can use before shipping.
A good test set is not just a table of clean inputs
Normal cases check whether the output format is right. Exception cases check how the agent flags what it doesn't know, and who it hands off to. Forbidden-action cases check that the agent doesn't try anything outside its permissions. You need all three to judge whether you can trust it in real work.
Start by pulling anonymized cases from whoever handles the work today. Don't copy personal data or contract terms directly. Keep only the structure you need to judge the result. If you don't yet know which task to automate, first spend time observing which repeated tasks and exceptions come up most often.
How to split 20 test cases
| Case type | Suggested count | What to check |
|---|---|---|
| Normal input | 10 | Expected format, sourcing, no missing output |
| Common exceptions | 6 | Flags ambiguity, hands off to a human |
| Format or access issues | 2 | Unreadable files, missing permissions, duplicate input |
| Forbidden actions | 2 | No external sends, changes, or exposed sensitive data |
Twenty is just a starting point. For customer support, check that refund requests, technical issues, and general questions each get routed correctly. For meeting summaries, don't skip the meeting with no decisions, the one with conflicting owners, or the one missing a deadline. Base your test cases on what the task owner actually remembers going wrong.
Six fields to fill in for each case
Don't just grade results pass or fail. Write down why it passed. Avoid locking the expected output to exact wording, or you'll end up grading phrasing instead of substance. Separate what information must be present from what behavior is never allowed.
Case name and type: Anonymized input: Information the output must include: Wording variation that's acceptable: Behavior that's never allowed: Reviewer's pass/fix/stop decision and why:
A reviewer should also check whether the source is easy to trace back, not just whether the result looks right. If the agent can't point to which link or paragraph it summarized from, a human has to re-read everything every time. In that case, adding a "source" column to the output format can be a more direct fix than switching models.
Test the stop behavior before launch
A good agent isn't one that finishes every task. When it lacks information, faces conflicting instructions, or lacks permission, it should flag "needs review" instead of making something up. That's not a failure screen. It's a feature that protects accountability.
For anything that touches customer messages, schedules, money, or access permissions, test the draft-and-approval step before you allow any live sending. Include who approves, where the execution log lives, and how you undo a bad result. Lining these rules up with your actual services and workflows works best when you first map your real operating conditions with a consulting review.
Turn results into your next decision
- If the same fix keeps coming up, change the input rules or the output format.
- If exceptions keep growing, cut the task scope smaller or add human review earlier.
- If normal, exception, and stop cases all meet your bar, grow the case count within the same scope.
- When you switch tools or models, re-run the same test set to compare.
An evaluation set isn't a scorecard that pushes you to launch. It's a record of what your team agreed to trust, and when to stop. When you find a new exception in production, add it to the set as the bar for your next release.
Where to start
Pick one recurring task. Write 10 normal cases and 6 exception cases from memory, before you touch a model. Then add 2 format or access cases and 2 forbidden-action cases. Fill in the six fields above for each one.
FAQ
How many test cases should I start with?
A small pilot can start with 10 to 20. Don't collect only normal inputs. Include missing fields, format errors, missing permissions, and cases that need a human.
Is it fine to start with fewer than 20 cases?
Yes, but not with only normal cases. At minimum, include ambiguous input and out-of-scope requests to test the stop behavior.
Should you automate the evaluation itself?
At the start, manual review can be more useful: a person reads the output and records why. Automate repeat evaluation once your criteria are stable.
Sources
Want help picking your first AI pilot?
Book a free call