All English guides
AI agent evaluationAI agent testingAI agent checklistproduction AI agentsAI pilot checklist

How to build an AI agent evaluation set before you launch

Before you ship an AI agent, build a 20-case test set covering normal inputs, common exceptions, and forbidden actions, not just one good demo.

LeanX··4 min read

Also available in Korean: Read the Korean original

Short answer: Before you launch an AI agent, don't rely on one good demo. Build a test set of about 20 real cases that cover normal inputs, common exceptions, and actions the agent must never take. Write the expected result and the pass criteria for each one to reduce surprises after launch.

An AI agent can give different results for the same request, depending on the input file, tool state, permissions, and wording. One well-chosen demo case passing doesn't mean the agent is ready for real work. Real problems start with blank fields, duplicate customer records, unreadable attachments, and edge cases that need a human call.

Google Cloud's 2026 guide to production-ready AI agents notes that agents reason, act, and adapt. That means they need different testing, memory, orchestration, and security than traditional software. The 20 cases below aren't a guarantee for any specific tool. They're a practical template a small team can use before shipping.

A good test set is not just a table of clean inputs

Normal cases check whether the output format is right. Exception cases check how the agent flags what it doesn't know, and who it hands off to. Forbidden-action cases check that the agent doesn't try anything outside its permissions. You need all three to judge whether you can trust it in real work.

Start by pulling anonymized cases from whoever handles the work today. Don't copy personal data or contract terms directly. Keep only the structure you need to judge the result. If you don't yet know which task to automate, first spend time observing which repeated tasks and exceptions come up most often.

How to split 20 test cases

Case typeSuggested countWhat to check
Normal input10Expected format, sourcing, no missing output
Common exceptions6Flags ambiguity, hands off to a human
Format or access issues2Unreadable files, missing permissions, duplicate input
Forbidden actions2No external sends, changes, or exposed sensitive data

Twenty is just a starting point. For customer support, check that refund requests, technical issues, and general questions each get routed correctly. For meeting summaries, don't skip the meeting with no decisions, the one with conflicting owners, or the one missing a deadline. Base your test cases on what the task owner actually remembers going wrong.

Six fields to fill in for each case

Don't just grade results pass or fail. Write down why it passed. Avoid locking the expected output to exact wording, or you'll end up grading phrasing instead of substance. Separate what information must be present from what behavior is never allowed.

Case name and type:
Anonymized input:
Information the output must include:
Wording variation that's acceptable:
Behavior that's never allowed:
Reviewer's pass/fix/stop decision and why:

A reviewer should also check whether the source is easy to trace back, not just whether the result looks right. If the agent can't point to which link or paragraph it summarized from, a human has to re-read everything every time. In that case, adding a "source" column to the output format can be a more direct fix than switching models.

Test the stop behavior before launch

A good agent isn't one that finishes every task. When it lacks information, faces conflicting instructions, or lacks permission, it should flag "needs review" instead of making something up. That's not a failure screen. It's a feature that protects accountability.

For anything that touches customer messages, schedules, money, or access permissions, test the draft-and-approval step before you allow any live sending. Include who approves, where the execution log lives, and how you undo a bad result. Lining these rules up with your actual services and workflows works best when you first map your real operating conditions with a consulting review.

Turn results into your next decision

  • If the same fix keeps coming up, change the input rules or the output format.
  • If exceptions keep growing, cut the task scope smaller or add human review earlier.
  • If normal, exception, and stop cases all meet your bar, grow the case count within the same scope.
  • When you switch tools or models, re-run the same test set to compare.

An evaluation set isn't a scorecard that pushes you to launch. It's a record of what your team agreed to trust, and when to stop. When you find a new exception in production, add it to the set as the bar for your next release.

Where to start

Pick one recurring task. Write 10 normal cases and 6 exception cases from memory, before you touch a model. Then add 2 format or access cases and 2 forbidden-action cases. Fill in the six fields above for each one.

FAQ

How many test cases should I start with?

A small pilot can start with 10 to 20. Don't collect only normal inputs. Include missing fields, format errors, missing permissions, and cases that need a human.

Is it fine to start with fewer than 20 cases?

Yes, but not with only normal cases. At minimum, include ambiguous input and out-of-scope requests to test the stop behavior.

Should you automate the evaluation itself?

At the start, manual review can be more useful: a person reads the output and records why. Automate repeat evaluation once your criteria are stable.

Sources

  1. Google Cloud

Want help picking your first AI pilot?

Book a free call