Published on · updated on · AI agents

What it takes to call a chatbot an agent

What an AI agent is and how it differs from a chatbot: tools, memory, planning and autonomy, plus questions to find out what you're actually being offered.

Chatbot or agent

Over the last few months, almost anything that converses has been pitched as an agent. Assistants that answer questions about a catalog, customer service chatbots and systems that book, pay or write code all share the same word. They don’t do the same thing or carry the same risks.

A practical way to see it: a chatbot answers, while an agent also decides what to do and acts on other systems to reach a goal. There are many shades in between.

A definition that goes way back

The word isn’t new in artificial intelligence. Stuart Russell and Peter Norvig, in the textbook that trained many of us who work in AI, define an agent as anything that perceives its environment through sensors and acts on it through actuators (Russell and Norvig). By that definition, even a thermostat is an agent.

It’s useful as a starting point, but it doesn’t help tell products apart. A chatbot also perceives (it reads what you type) and acts (it answers you). A large language model (LLM) receives a prompt and acts by responding with text. To separate one from the other, you have to look at what else it can do.

The ingredients of an agent

Tools

An agent can use tools: search the web or your documents, query a database, call a booking system’s API, send an email. Without tools, a language model can only answer with what it learned during training. Andrew Ng listed tool use among the four design patterns of agentic systems, alongside reflection (reviewing its own work), planning and collaboration between several agents (Ng, 2024a).

A think-and-act loop

An agent doesn’t answer in one go. It decides on a step, carries it out, looks at the result and decides the next one. In 2022, a team from Princeton and Google proposed a way to do this with language models, ReAct, in which the model alternates reasoning and actions and uses what it gets from each action to keep going (Yao et al., 2022). It’s the basis of most current agents.

Memory

To complete a multi-step task, you have to remember what’s been done and what’s been found out. And to be useful over time, you also have to remember each person’s preferences. What gets stored, for how long and who can see it are design decisions, not just technical ones.

Autonomy

The ingredient that changes things most is who decides the next step. Anthropic explains it with a useful distinction: in a workflow, the language model and tools follow paths defined in code; in an agent, the model directs its own process and decides which tools to use (Anthropic, 2024). In the same guide they recommend starting with the simplest option and adding autonomy only when simple isn’t enough.

Agentic workflows

There’s a lot of room between a fixed flow and an agent that decides everything. Andrew Ng proposed we stop arguing about whether something is or isn’t an agent and talk about systems that are more or less agentic (Ng, 2024b). That’s where the idea of an agentic workflow comes from: the answer for hybrid approaches, where some steps need the model to reason and decide, others need the reliability and predictability of a step-by-step process, and others need a person to review before moving on (known as human in the loop, HITL).

Two-panel comic. In the first, one person points at a screen and says "It's an agent!"; another replies "No, it's not!". In the second, they shake hands and say "It's agentic!".Source: Andrew Ng, The Batch issue 253.

A scale, not a yes or no

To figure out how agentic a product is, it helps to place it on a scale:

  1. It answers with what the model knows.
  2. It answers by consulting your data, for example by searching your documents.
  3. It carries out a specific action when you ask, such as booking an appointment, and asks you to confirm.
  4. It chains several steps on its own to reach a goal, within a defined process.
  5. It decides how to reach the goal and acts without supervision.

Many useful products sit at levels 2 and 3. That’s not a flaw: for most customer service cases, a well-designed flow is more reliable and cheaper than an autonomous agent.

Moving up the scale has a cost. Margaret Mitchell and her colleagues at Hugging Face sum it up: the more control a person hands over to an agent, the more risks appear (Mitchell et al., 2025). An agent that only looks things up can get an answer wrong. One that acts can get a payment wrong, delete a customer database or send something on your behalf.

Questions to find out what you’re being offered

If someone pitches you an agent, these questions clear up a lot:

  • What tools does it have, and which systems can it access?
  • Who decides the next step, the model or a defined flow?
  • What can it do without asking for confirmation?
  • What happens when it gets something wrong, and how do you undo what it did?
  • Can you review afterward what it did and why?
  • When does the conversation go to a person?

The answers say more than the word “agent” on the slide.

At Monoceros we design assistants and agents always starting from the lowest level that solves the problem. We explain it in what changes when your app learns to talk, and it’s the starting point for our work in conversation and agent design.

References

Written by

Nieves Ábalos

Co-founder at Monoceros Labs

Co-founder of Monoceros Labs. A computer engineer researching dialogue systems since 2009, now doing a PhD in AI at the University of Granada.

Share this article:

Related posts

2018 and today: what's changed in building a chatbot

Building a chatbot in 2018 and building one today

From Alexa intents and flows to agents built on language models: what has changed in building a chatbot or voice assistant, and what hasn't.

Several agents for one call

How many agents does it take to answer a phone call?

Multi-agent systems: when splitting a task across several AI agents pays off, when it creates more problems than it solves, and what changes on a phone call.

Who checks the work?

In a newsroom with agents, who checks the work?

Newsrooms are automating more and more and moving toward AI agents. Why checking everything by hand doesn't scale, and why you need an expert in the loop.