With language models, adding a chat to an app seems like an afternoon’s work: an API call, a text box and you’re done. And it’s true you can have something that responds in very little time. The hard part starts afterward, when that conversation has to be reliable, useful and cheap to maintain.
In 2024 I wrote about this for Samsung Dev Spain (Ábalos, 2024). This post collects the main ideas and brings them up to date.
What changes for the people using the app
A conversational interface lets people ask for things in their own words instead of hunting for the right button. Done well, it simplifies the screen and makes the app more accessible. The conversation itself is also informative: someone who asks for “something gluten-free for dinner” is telling you what they need without you having to ask in a form.
The flip side is that expectations go up. If the app speaks naturally, people expect it to understand like a person, remember what was said earlier and know how to say no when something can’t be done. That’s why you still need to design the conversation, even if a model generates it.
How it’s built
The classic way, with intents and flows
The system classifies each sentence into an intent (“check balance,” “change password”), extracts the important details and follows a flow designed in advance. It’s predictable and easy to audit, but it only understands what someone anticipated, and it gets very expensive to maintain as it grows.
With a language model
The model interprets what the person says and generates the response. It understands many more ways of saying the same thing and responds naturally, but it can make up facts, skip steps or answer things it shouldn’t.
The hybrid way
The language model interprets and writes, while the business logic stays in code or in flows that always run the same way. Many platforms have gone in this direction. Rasa, for example, describes it this way in its CALM approach: the model interprets what the user wants and the logic decides what happens next (Rasa). For most apps that touch data or money, it’s the most sensible option.
And now, agents
Since 2024 a fourth option has joined them: agents, which besides conversing decide which tools to use to complete a task. Standards like the Model Context Protocol, which Anthropic introduced in November 2024, make it easier to connect an assistant to a company’s systems (Anthropic, 2024). They offer more autonomy, and that’s why they need even more limits.
The challenges that remain
Keeping context
A useful conversation remembers what’s been said: if you ask “and tomorrow’s?”, the system has to know you were talking about a booking. Language models handle context well within a conversation, but they have a memory limit, and sending them the whole history every time costs money. Between sessions, someone has to decide what’s stored, where and for how long.
Made-up answers
A language model answers confidently even when it doesn’t know. For questions about your products or your policies, the usual technique is retrieval-augmented generation, or RAG: before answering, the system searches your documents and passes the relevant passages to the model (Lewis et al., 2020). It greatly reduces the problem, but doesn’t eliminate it. You have to check that the answer says what the documents say.
When the conversation goes off track
People don’t follow the script: they change topic, correct what they said or ask for things the app doesn’t do. Classic systems responded with “I didn’t understand.” Language models make it possible to recover the conversation more naturally, as long as the design anticipates what to do in each case, including handing over to a person.
Testing what you can’t list
In an app with buttons, you can test every path. In a conversation, the paths are endless and the model never answers the same way twice. You have to test with many realistic conversations and measure, rather than reviewing them one by one. It’s a discipline of its own, and we explain it in evaluating an agent before your customers use it.
Cost, latency and security
Every response from a large model has a cost and takes time, which is more noticeable in voice. Sometimes a classic system handles the same thing just as well, faster and cheaper. And there are new risks: prompt injection, where someone writes text so the model ignores its rules, tops OWASP’s list of risks for language model applications (OWASP, 2025).
Saying it’s an AI
Since August 2, 2026, the EU AI Act requires systems that converse with people to tell them they’re talking to an AI, unless it’s obvious (Regulation (EU) 2024/1689, Article 50). It’s been good design practice for much longer: people adjust how they speak and how much they trust depending on who they know they’re talking to.
Where to start
Start with a specific, well-defined use case where you know what a good answer looks like. Design the conversation before choosing the model. Decide which parts need to be deterministic. And from day one, build a set of test conversations to measure every change against.
If that’s where you are, it may help to first read what to ask yourself before building an assistant. And if you’d like us to think it through with you, that’s what we do in conversation and agent design.
References