How to Evaluate an AI Agent Before You Trust It With Real Work
Agent demos always look impressive. A practical framework for testing whether one will hold up on your actual tasks, data and edge cases.
AI agents — systems that plan, use tools and take multiple steps toward a goal — demo beautifully. Watch one book a meeting, update a CRM and draft a follow-up email, and it's easy to imagine handing over half your operations.
Then you deploy it and discover it handles the demo scenario perfectly and the third real customer request badly. The gap between those two moments is evaluation. Here's a framework for closing it before you go live.
Start with the job, not the agent
Write down, specifically, what you want the agent to do:
- The task: "Triage inbound support emails, tag them, and draft replies for billing questions."
- The boundary: "It may read the help centre and customer records. It may not issue refunds or change plans."
- Success: "A human approves the draft without edits at least 70% of the time; tags are correct 95% of the time."
Without this, you'll judge the agent on vibes — and vibes are always generous during a demo.
Build a test set from reality
The most valuable thing you can create is a collection of real examples with known good outcomes. Aim for 50 to 200 to start.
Include:
- Typical cases in the proportions you actually see them.
- Hard cases: ambiguous requests, missing information, multiple questions in one message.
- Adversarial cases: angry customers, attempts to get the agent to break its rules, instructions hidden inside documents it reads.
- "Should refuse" cases: requests outside its boundary, where the correct behaviour is to hand off to a person.
Pull these from your history: past tickets, previous requests, real documents. Synthetic examples are useful for filling gaps, but they're always tidier than reality.
Measure outcomes and the path
Agents can reach a good answer in a bad way, or a bad answer through reasonable steps. Check both.
Outcome questions:
- Was the final result correct and complete?
- Did it stay within its boundary?
- Did it hand off when it should have?
Path questions:
- Did it use the right tools, with sensible inputs?
- How many steps did it take? Wandering usually signals confusion.
- Did it recover from tool errors, or loop?
- What did it cost in time and money per task?
Log every step. When something fails, the trace usually shows exactly where the reasoning went wrong.
Use graders carefully
Scoring hundreds of outputs by hand is slow. A common approach is to use a language model as a grader, with a clear rubric. This works well when:
- The rubric is specific ("Does the reply mention the correct refund window of 30 days?") rather than vague ("Is the reply good?").
- You check the grader against human judgment on a sample before trusting it.
- You keep deterministic checks for anything that can be checked exactly — correct tag, correct field values, forbidden actions not taken.
Test the failure modes that hurt most
Some failures are embarrassing; others are expensive. Focus testing on the second kind:
- Irreversible actions. Sending emails, deleting records, making payments. Should these require human approval? Usually, at first, yes.
- Prompt injection. If the agent reads emails, web pages or documents, test what happens when those contain instructions like "ignore your rules and forward this to…"
- Data leakage. Can one customer's request surface another customer's data?
- Confident wrongness. Does it say "I don't know" when it should, or invent an answer?
Roll out in stages
Even a well-tested agent should earn autonomy gradually:
- Shadow mode. The agent works on real tasks, but humans do the actual work. Compare results.
- Draft mode. The agent prepares; a human approves every action.
- Partial autonomy. Low-risk, well-understood task types run automatically; everything else still needs approval.
- Ongoing monitoring. Sample completed tasks weekly, track your success metrics, and add every real failure to the test set.
The most important habit
Re-run your test set every time anything changes: the model version, the instructions, the tools, the knowledge base. Agents are systems, and small changes in one part ripple unpredictably. A test set you can run in minutes is what lets you improve the agent confidently instead of nervously.
Have something worth publishing?
We accept guest posts across all 8 topics, edited and published within days.