Published on · updated on · Evaluation
Evaluating an agent before your customers use it
How to evaluate an AI assistant or agent before launch: what to measure, test conversations, a model as judge, attacks and testing with real people.


An AI agent almost always does well in the demo. The questions are the expected ones, the person testing it knows how to talk to it, and nobody tries to confuse it. Problems show up later, with customers who ask something else, make mistakes, change their minds or type in a hurry.
The research numbers aren’t very reassuring. In τ-bench, a benchmark that simulates customer service conversations with real tools and policies, the best agents of 2024 completed fewer than half the tasks. And when the same task was repeated eight times, in the online retail scenario, fewer than one in four succeeded all eight times (Yao et al., 2024). Models have improved since then, but the lesson stands: an agent that gets it right once won’t necessarily get it right every time.
These are the steps we follow to evaluate an assistant or agent before it reaches customers.
1. Define what doing it well means
It sounds obvious and it’s almost never written down. For each important task: what it has to achieve, which policies it must follow, which data it can use, what tone it should have and when it has to hand the conversation to a person. Without that, each reviewer judges by their own criteria and the results can’t be compared.
2. Gather realistic test conversations
The foundation of everything is a set of conversations to test each version with. The best ones come from real, anonymized conversations: customer service emails, chats and calls. You need to supplement them with what fails most:
- Ambiguous or incomplete requests.
- People who correct themselves or change their minds halfway through.
- Questions the agent shouldn’t answer.
- Data in odd formats: dates, addresses, order numbers.
- Cases where the right answer is to hand over to a person.
You can also use a language model to play the customer and generate lots of variations, as τ-bench does. It’s useful for expanding the set, but it doesn’t replace real conversations.
3. Repeat each test several times
A language model doesn’t always answer the same way. A test that passes once may fail the second time. That’s why each important conversation is run several times and you measure how many pass every time, not just whether any pass. The metric τ-bench proposes, which requires success on every attempt, is a good example.
4. Automate what can be checked with rules
Many things can be checked without human judgment: whether the answer has the right format, whether it exposes data it shouldn’t, whether it exceeds the maximum length, whether it called the right tool, whether the booking was saved correctly in the system. Hamel Husain, who has helped many teams evaluate AI products, recommends starting with these simple, cheap checks and running them with every change (Husain, 2024).
5. Use a model as judge, and keep an eye on it
For what can’t be checked with rules, like helpfulness or tone, you can ask another language model to rate the answer against a rubric. Its reliability varies by criterion, so you should validate it against human judgments and mitigate its biases. Whenever correctness can be verified (for example, against a database), it’s better to check it directly or give the judge access to that information.
A 2023 study found that GPT-4 as a judge agreed with human ratings in over 80% of cases, the same rate at which two people agree with each other. It also identified its biases: it prefers the first answer it reads, longer answers and answers that resemble its own (Zheng et al., 2023). That it works in a study doesn’t guarantee it works in your product: before relying on a model as judge (LLM-as-a-judge), check that it works for your use case and your domain.
The most reliable way to use it is to calibrate it: have several people on the team rate a sample, compare with what the judge says and adjust the criteria until its agreement with the people approaches the agreement among them. Then confirm it on a different sample, and check again every time the model or the type of conversations changes.
6. If it answers from your documents, measure two things separately
When an assistant answers by consulting your documentation (known as RAG, retrieval-augmented generation), it can fail in two phases. When searching, it may not retrieve the right passages, or bring them mixed with lots of irrelevant text. When answering, it may state something that isn’t in what it found, or go off on a tangent. It’s worth evaluating each phase separately, because they’re fixed differently: in one case you improve the search, in the other the instructions or the model that writes.
Ragas, an evaluation framework for this kind of system, proposes three metrics (Es et al., 2023). The first measures whether the retrieved context is focused on the question. The second, whether the answer is faithful to that context, that is, whether every claim is supported by it. The third, whether the answer actually addresses the question. Their appeal is that they don’t need hand-written reference answers, because another language model acts as the evaluator.
That said, it’s worth knowing their limits. A faithful answer isn’t necessarily correct: if the retrieved document is out of date, the assistant can repeat the error with complete faithfulness. These metrics also don’t detect well when the key passage is missing, because for that you’d need to know in advance what should have been found. And since the judge is a language model, its ratings aren’t infallible.
7. Try to break it
Someone is going to try to make the agent say what it shouldn’t, reveal its instructions or do something it isn’t allowed to. Prompt injection tops OWASP’s list of risks for language model applications (OWASP, 2025). Before launch, spend time attacking it on purpose. This is known as red teaming: a team puts itself in the attacker’s shoes to find the flaws before anyone else. It can be done with people, and also with language models that generate attacks automatically, a variant DeepMind researchers proposed in 2022 (Perez et al., 2022).
8. Test with people
No automated test replaces watching real people use the assistant. Requests nobody anticipated show up, along with different ways of speaking and moments of confusion that don’t appear in the data. If it’s a voice assistant, you also have to test it with noise, with different accents and with people who pause, as we explain in why doesn’t my voice assistant understand me and is it my turn to speak?.
9. Keep measuring after launch
Evaluation doesn’t end at launch. Review real conversations every week, look at where and why conversations get handed to a person, and add new failures to the test set. And every time you change models, even to a new version from the same provider, run all the tests again: behavior can change without warning.
Pre-launch checklist
- Is it written down what a good answer looks like for each task?
- Is there a set of realistic test conversations, including hard cases?
- Is each test repeated several times?
- Do the automated checks run with every change?
- Is the judge model calibrated against human ratings?
- Has someone tried to break it on purpose?
- Have real people used it?
- Is there a plan to keep measuring afterward?
It’s what we do in our assistant and agent evaluation service, and what we recommend to any team before putting an agent in front of its customers.
References
- Yao, S., Shinn, N., Razavi, P. and Narasimhan, K. (2024). τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv.
- Husain, H. (2024). Your AI Product Needs Evals.
- Zheng, L., Chiang, W.-L., Sheng, Y. et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
- Es, S., James, J., Espinosa-Anke, L. and Schockaert, S. (2023). Ragas: Automated Evaluation of Retrieval Augmented Generation. arXiv.
- OWASP GenAI Security Project (2025). Top 10 for LLM Applications.
- Perez, E., Huang, S., Song, F. et al. (2022). Red Teaming Language Models with Language Models. EMNLP 2022.
Share this article: