
Building a chatbot in 2018 and building one today
From Alexa intents and flows to agents built on language models: what has changed in building a chatbot or voice assistant, and what hasn't.
Published on · AI agents
Fast models, reasoning models and the new family of models that only make decisions, like Jev: what each brings to a voice agent, where latency rules, and how to combine them.


Daniel Kahneman popularized the idea that we think in two ways. System 1 is fast, automatic and intuitive: it recognizes a face, understands a sentence, dodges a ball. System 2 is slow and deliberate: it calculates, compares options, double-checks (Kahneman, 2011). We use the first almost all the time and the second only when needed, because it’s expensive.
Something similar has happened with language models. Today, models that answer instantly coexist with models that take their time to reason before answering. And in recent months a third type has appeared that doesn’t even write: it only decides. In a voice agent, choosing well among them is one of the decisions that most affects the experience.
Classic language models generate the answer directly, word by word. They’re fast and, for most conversational tasks, good enough.
In September 2024, OpenAI introduced o1, a model trained to generate a long chain of internal reasoning before answering (OpenAI, 2024). In January 2025, DeepSeek published R1, which showed that this capability could be achieved with reinforcement learning and put it within anyone’s reach with an open-weights model (DeepSeek-AI, 2025). Hybrid models soon followed: Claude 3.7 Sonnet, in February 2025, let you choose between an immediate answer or extended reasoning, and set how many tokens it could spend thinking (Anthropic, 2025). Today almost every provider offers something similar.
Reasoning models are clearly better at math, coding, planning and problems with many conditions. The price is time and cost: every reasoning token has to be generated, and paid for.
A late-2024 study looked at how these models behave with trivial problems. Its title says it all: Do NOT Think That Much for 2+3=? It found they spent a lot of resources on simple problems without improving the result (Chen et al., 2024). Overthinking wastes time and money, and sometimes leads the model to second-guess something that was right the first time.
In a text chat, waiting a few seconds for a well-thought-out answer may be acceptable. In a voice conversation, it isn’t. As we explain in is it my turn to speak?, the silences between turns in a conversation between two people last tenths of a second. An agent that takes five seconds to answer “what time do you open tomorrow?” seems broken, however good the answer.
That’s why, in voice, the model talking to the person has to be fast. Most of what happens in a customer service conversation (greeting, understanding what someone wants, asking for a detail, confirming, answering a frequent question) doesn’t need long reasoning.
Much of what an agent does isn’t talking but making small decisions. What does this person want? Are they asking to speak to someone? Do we have all the details yet? Are they getting upset? To answer that, a language model generates text that then has to be interpreted, and sometimes it answers something that wasn’t among the options.
In September 2026, TypeSafe AI introduced Jev, the first of what it calls System One models, a direct nod to Kahneman (TypeSafe AI, 2026). Jev doesn’t generate text. It receives the situation (the conversation, the customer’s data) and a closed question, and returns an answer of a predefined type: an option from a list, the probability that something is true or a score on a scale. Always with a calibrated probability, meaning one that genuinely reflects how sure it is.
Its value lies in three things:
They’re not for everything. They don’t write answers or explain their reasoning, they only work with text and they handle a limited number of options. They’re one more component, not a replacement.
It’s not just one company. There are already open alternatives like Laya, under a free license, which makes this kind of decision in more than a hundred languages with small models that fit on a single GPU (Laya). There are even dedicated benchmarks: JevBench pits more than a hundred systems against each other on accuracy, calibration, speed and cost, and many of the best are practically tied (JevBench). And OpenAI recently announced Decisions API, its own take on this kind of fast, structured decision-making. Everything suggests this category will be at the heart of many agents.
The most interesting solution is not to choose: combine both. In 2024, Konstantina Christakopoulou and colleagues proposed an architecture directly inspired by Kahneman, with two components. A fast “talker” that keeps up the conversation with the person, and a slow “reasoner” that plans, uses tools and makes the complex decisions in the background (Christakopoulou et al., 2024).
In a voice agent, that looks something like this: the fast model answers right away, asks for missing details and says it’s checking something (“one moment, let me check”), while a slower model reviews the conditions of a return or calculates a rate. When the reasoner finishes, the talker communicates the result. The person doesn’t hear long silences, and difficult decisions aren’t taken lightly. It’s similar to the hybrid architecture we describe in three ways to build a voice agent.
Decision models fit in as a third component. While the talker speaks, a model like Jev instantly resolves the small decisions: which flow to route to, whether a detail is missing, whether the person wants to talk to someone. When it’s confident, the agent moves on. When it isn’t, the decision goes to the reasoner or to a person.
Fast models:
Decision models, like Jev:
Reasoning models:
Measure each task separately. For each one, compare on your test conversations how often a fast model gets it right, a reasoning model and, if the task is a closed decision, a decision model; how long each takes and how much each costs. Public benchmark results point the way, but they don’t replace testing with your own data, especially in Spanish. Often the fast model is almost as accurate, and the difference doesn’t justify the wait. Other times the difference on a specific task is large, and that’s where it’s worth paying for the time, ideally in the background.
It’s a field that changes every few months, and what needs a reasoning model today will be handled by a fast one tomorrow. That’s why we prefer to design agents where swapping the model behind each component is easy, and to measure before deciding. If you want to explore it for your case, it’s the kind of work we do in applied research and prototypes.
Share this article:


From Alexa intents and flows to agents built on language models: what has changed in building a chatbot or voice assistant, and what hasn't.


Multi-agent systems: when splitting a task across several AI agents pays off, when it creates more problems than it solves, and what changes on a phone call.


Newsrooms are automating more and more and moving toward AI agents. Why checking everything by hand doesn't scale, and why you need an expert in the loop.