“We're worried about launching an assistant without knowing if it's reliable, treats users well or complies with regulation.”

Services

Assistant and agent evaluation

We analyze voice and chat assistants and agents against conversational quality, user experience and risk criteria. You get a clear diagnosis and an improvement plan, with special attention to Spanish.

Chatbot and AI agent evaluation, with a focus on Spanish

Evaluating a chatbot or AI agent means checking, with real or simulated conversations, what it understands, where it fails and what risks it carries before your users find out. We measure conversational quality, user experience and compliance with the EU AI Act, with particular attention to Spanish.

Conversation audit

We analyze real or simulated conversations to find misunderstandings, made-up answers, loops and drop-offs.

User testing

We watch real people use the assistant and see where they get frustrated.

Custom quality criteria

We define what working well means for your case and how to measure it continuously.

Custom evaluators

We build automated evaluators tuned to your criteria, so every new version gets checked without starting from scratch.

Conversation design proposal

Based on the diagnosis, we propose specific changes to dialogues, instructions and personality.

Risk and transparency

We check how it discloses it's an AI and whether it behaves in line with the EU AI Act, how it handles sensitive topics and when it hands over to a person.

How we work

01

Explore

We understand the challenge and review what today's technology can do for your case, your language and your data.

02

Prototype

We build just enough to validate the idea with real people.

03

Pilot

We try it out in your organization, measure and adjust.

04

Hand over

We document and train your team so they can keep evolving it.

What you get

  • Diagnosis with real examples
  • Set of test scenarios
  • Quality criteria and metrics
  • Prioritized improvement plan

Who it's for

  • Companies with assistants or agents in production
  • Teams about to launch an agent
  • Regulated sectors or services for vulnerable users

Related projects

All projects
Project

Evaluation and personality for a real estate portal's chatbots

Evaluating a real estate portal's two help chatbots, and a workshop to define their personality.

Project

Sanitas Mayores: Alexa in care homes

We supported Sanitas Mayores as Alexa arrived in its care homes: use cases, conversation design for older people and integration with its AI, SanIA.

Project

Clevergy: conversation design review

We reviewed the voice assistant of Clevergy, a platform that helps people understand their home electricity use, and proposed conversation design improvements.

In the lab

Ideas we're working on. If one fits a challenge of yours, we can pilot it together.

Chatbot and AI agent evaluation FAQ

Do you evaluate assistants you didn't build?

Yes. Fresh eyes catch what the team no longer sees.

Do you have an evaluation tool?

We're exploring a conversation evaluation platform. Meanwhile we offer it as a service and, if you're interested, we can pilot the platform with you.

What do you need to evaluate our assistant?

Access to the assistant, live or in testing, and ideally a sample of anonymized real conversations. With that we agree with you on what "working well" means for your case and build the test scenarios.

How long does an evaluation take?

It depends on how many assistants and channels (voice, text) there are and whether it includes user testing. For a real estate portal we evaluated two chatbots and closed with a personality workshop, a recommendations report and quality indicators.

Is automated evaluation with another AI model enough?

As support, yes: automated evaluators let you check each new version without starting over. But they have to be tuned to your criteria and checked against human review, because a model grading another model makes mistakes too.

Services

Got a voice or conversation challenge?

Tell us what you want to achieve. If we're not the right team, we'll say so.