How to Evaluate AI Agents: Building Test Cases and Measuring ROI
・ Employee Store Operations

Summary
Once you adopt an AI agent, you need to check whether it is actually helping. This article splits evaluation into accuracy and impact, and walks through everything from building test cases to calculating ROI. It ends with a way to decide, from the results, whether to continue or stop.
The guidelines and tools described in this article were checked on their official pages on October 2, 2026.
No single number tells you how to evaluate an AI agent. Whether its answers are correct and whether the work got easier are two separate questions. This article introduces evaluation tools in order and turns them into a form you can use in your own company.
Split evaluation into accuracy and impact
Accuracy is whether the AI agent's output is correct. Impact is how the time and cost of the work changed after adoption. Even with high accuracy, impact is small if checking the output takes a long time. Conversely, even if some fixes are needed, there is impact if total time drops sharply.
Accuracy
- Is the output correct
- Measured with test cases
- Measured before adoption and after each update
Impact
- How time and cost changed
- Measured with real work records
- Measured over set periods after adoption
Time alone does not capture impact. Choose the metrics to track from the list below, based on the type of work.
- Work time: time to handle one item
- Volume: items handled by the same number of people
- Fixes: how often people corrected the AI's output
- Errors: number of returns or corrections
- Wait time: time from request to completion
The “Guide to Evaluation Perspectives on AI Safety” from the Japan AI Safety Institute (AISI) says evaluation should not be a one-time event and should be repeated at appropriate times within reason. Version 1.20, published July 7, 2026, added “autonomous behavior” and “interaction with external environments” as evaluation perspectives specific to AI agents. Because AI agents operate external services, check not only whether answers are correct but also whether the agent acts in unintended ways.
Build test cases from your own work
To test an AI agent, use test cases built from your company's real work. A test case is a pair of an input and an expected result. For example, in agent evaluation for Microsoft Copilot Studio, one question or a whole conversation counts as one test case, and you can attach the expected response. A collection of test cases is called a test set.
How to collect them
- Pick common ones from past requests and inquiries
- Include rare cases that cause trouble when they happen, such as large amounts or close deadlines
- Include requests the AI agent should decline or pass to a person
- Mask personal information or replace it with fictitious values
How to write one case
Each test case lists the input, the expected result and how to score it. For an AI agent that drafts replies to inquiries, it might look like this.
- Input: a customer inquiry saying “I want to change the delivery address”
- Expected result: explains the steps to change it. If the item has already shipped, passes it to a staff member
- How to score: whether the steps are explained, and whether shipped items are passed to a person
- Weight: flag cases involving money or personal information as high-impact if handled wrongly
Safety test cases
The evaluation perspectives guide lists, among others, preventing false or misleading output, protecting privacy, ensuring security and robustness as perspectives for evaluating AI systems. Robustness means producing stable output for unexpected inputs such as adversarial instructions, garbled data and input errors. These translate into test cases like the following.
- Does it say it does not know, rather than making up numbers or facts without a basis
- Does it decline when asked for another person's personal information
- Does it refuse to follow text that tells it to ignore internal rules
- Does it stop or ask for confirmation when input has many typos or blanks
A test set is not something you build once before adoption and leave. When you find failures in production, add those cases. Running the same test set after every update lets you check that a fix did not break something else.
Setting correct answers and scoring
How you set the correct answer for a test case depends on the type of work.
- Work with one right answer: amounts, dates, classification labels and so on. These can be scored automatically by exact match
- Work judged against criteria: summaries, draft replies and so on. Write the scoring criteria first, then score against them
- Work that checks actions: registering bookings, updating data and so on. Check whether the result was reflected correctly in the system
Scores for work judged against criteria vary by scorer. Write out separate points to check, such as “includes the necessary information,” “has no factual errors” and “matches our in-house wording.”
The evaluation perspectives guide says evaluation by people not directly involved in developing the system is also effective for objectivity. Do not leave scoring to the adoption lead alone. Bring in people who actually do the work and people from other departments.
You can also use AI to score. As an example of automating the evaluation of AI output, Anthropic describes using a separate AI call for each evaluation criterion. However, AI scoring can be wrong too. At first, have people score some of the cases as well and check that their scores match the AI's.
How to read public benchmarks, and their limits
AI models are sometimes compared by their scores on public benchmarks. For example, SWE-bench is a benchmark built from real GitHub issues in 12 Python repositories. SWE-bench Verified consists of 500 tasks selected from it by people.
A benchmark score is performance on a fixed set of tasks. It does not show performance on your own work, data and internal rules. It can guide model selection, but decide whether to adopt based on your own test set.
When you look at benchmark scores, check the following.
- Type of task: code fixes or answering questions, and how close it is to your work
- Language of the tasks: English or Japanese
- Test conditions: model version, tools or scaffolding used, and when it was measured
- Who measured it: the model developer's announcement or a third party
Anthropic explains that agents often trade higher latency and cost for better performance. It recommends deciding whether a complex system is truly needed based on measured results.
The ROI formula
The ROI (return on investment) of an AI agent is calculated from the value of its impact and its cost. First, work out the value of the impact for one month.
To measure impact, record work time before adoption. If you estimate it from memory after adoption, the numbers become vague. Record the time per item and the number of items per month for about two weeks to one month before adoption.
Hours saved is work time before adoption minus work time after adoption. Always include the time spent checking and fixing the AI's output in the after-adoption time. Leaving it out overstates the impact.
Next, work out the cost for the same period. Include the AI agent's price (for a one-time purchase, the amount divided over the period; for a monthly plan, the monthly fee), AI model usage fees, and staff time spent on setup and operation. ROI is the value of impact minus cost, divided by cost, times 100. Above 0%, the impact exceeds the cost.
- ROI (%) = (value of impact − cost) ÷ cost × 100
- Value of impact: hours saved × hourly rate
- Cost: the AI agent's price, AI model usage fees and staff time for operation
How to list cost items is covered in detail in AI agent costs.
Deciding whether to continue or stop
Once you have accuracy and ROI, decide whether to continue, fix or stop. Set the decision criteria before the trial. If you set them afterward, it is easy to interpret results in your favor.
Set criteria according to the weight of each case. For cases involving money or personal information, use a strict criterion, such as not going to production if even one case is wrong. For work a person always checks, such as drafting text, the criterion is whether work time drops even after including fix time.
How did accuracy and impact turn out
The AI Guidelines for Business (version 1.2) from Japan's Ministry of Internal Affairs and Communications (MIC) and Ministry of Economy, Trade and Industry (METI) list checking that AI works as specified as one of the actions for businesses that use AI. Even after you decide to continue, run the test set regularly and review impact using work records. How to review during operation is covered in Operating and maintaining AI agents, and the overall adoption plan in How to adopt AI agents.
Adopting and evaluating on Employee Store
Employee Store is a marketplace where companies can adopt AI agents (AI employees) built by developers. Prices are a one-time purchase, a monthly fee, or a setup fee plus a monthly fee, so they go straight into the cost side of your ROI. Buyers and sellers can message each other in a deal room for each transaction, and buyers confirm receipt after checking the deliverables. If you prepare a test set in advance, you can check how the agent performs on your own work before confirming receipt.
FAQ
- How many test cases does an AI agent need?
- There is no set number. Make sure the set includes common requests, rare but troublesome requests and requests the agent should decline. When you find failures in production, add those cases.
- Is it enough to pick a model with a high public benchmark score?
- No. A benchmark shows performance on a fixed set of tasks, not on your own work. Use it only as a reference for choosing a model, and decide on adoption with your own test set.
- Which costs are easy to miss in an ROI calculation?
- Time spent checking and fixing the AI's output. If you do not subtract it from the hours saved, you will overstate the impact. Also include staff time spent on setup and operation in the cost.
Sources
- Japan AI Safety Institute, Guide to Evaluation Perspectives on AI Safety (version 1.20, published July 7, 2026) (Japanese)
- Japan AI Safety Institute, Guide to Evaluation Perspectives on AI Safety (version 1.20), summary (Japanese)
- MIC, AI Guidelines for Business (version 1.2, published March 31, 2026) (Japanese)
- MIC and METI, AI Guidelines for Business (version 1.2), main text (Japanese)
- Anthropic, 'Building effective agents'
- Microsoft Learn, About agent evaluation (Copilot Studio) (Japanese)
- SWE-bench (checked October 2, 2026)


