Agent, chatbot, or workflow? A 90-second decision
Three questions settle it: is the path fixed, is the conversation bounded, does it need tools and judgement. With what each shape costs to run and to keep.
Three questions, in this order. Does the same input always need the same steps? Build a workflow. Is every topic it covers and every action it takes on a list you could write down today? Build a chatbot. Does it have to call tools and work out the order for itself? That one, and only that one, is an agent.
An AI agent is software that is given a goal and a set of tools, and decides for itself which tools to call and in what order, until the goal is met or it gives up.
Does the same input always need the same steps?
yes -> workflow
no -> Can you list every topic it may cover and every action it may take?
yes -> chatbot
no -> Does it have to call tools and choose the order itself?
yes -> agent
no -> the requirement isn't written yet. Write it first
The other two shapes decide nothing, and that is the whole distinction. A workflow's path was chosen by whoever drew it. A chatbot picks words rather than actions. An agent picks actions, and everything expensive about agents follows from that: you cannot enumerate what it will do, so you cannot test it the way you test the other two, so you constrain it in code instead.
Is the path the same every time?
If you can draw the steps on a whiteboard and the drawing does not change depending on what the input says, you want a workflow. A trigger, a fixed sequence, a few conditions, a definition of done. An invoice arrives, fields come out of it, they get matched against a purchase order, the result posts to the ledger and a human gets a message when the numbers disagree.
A workflow can contain a language model without becoming an agent, which is where most of the confusion lives. That extraction step is a model call. So is classifying an inbound email, or drafting the first version of a reply. In each case your code owns the order of the steps and the model fills one of them in. Most of what is being sold as an agent this year is exactly that: good design, misleading label. Owning the sequence is worth more than it sounds, because the acceptance criteria stay a table of inputs and expected outputs that runs in CI forever. Hand the sequence to a model and you lose the table. Pay that price deliberately or not at all.
Is every topic on a list you could write down?
A chatbot is a conversation with a fence around it. It answers from a corpus you control, collects fields you specified, and hands off to a person at the edge of what it was built for. If the topics and the actions fit on one side of a sheet of paper, this is your shape, and it costs a small fraction of the alternative. The risk profile is the better argument, though. A bounded system fails by refusing politely, not by doing something.
Fences need enforcing, and that is where teams get caught. On the public assistant we run for a client marketplace, questions about tax treatment sit outside what it is competent to answer, and that rule is not a line in a prompt. A detector on the server sets a flag on the conversation and rewrites the reply into a handover to a human. The flag is sticky, so the assistant cannot wander back into the subject three turns later. The detector matches the tax abbreviation in capitals only, so an ordinary word spelled the same way in lower case cannot trip it. Somebody had to notice that, in production, after it happened.
Does it have to choose its own next step?
Reach for an agent when the number of possible paths is too large to enumerate, and not before. A question like "why has this order not been invoiced" can start in three different systems, and which one you look at second depends entirely on what the first returned. Nobody draws that on a whiteboard. An agent with a goal and a set of tools can work it out.
What you pay for is not mainly the model. You pay for a bounded loop, so a bad turn cannot run all night. You pay for checks written in code on the way out, because a prompt rule lowers a failure rate and code sets it to zero. You pay for an eval suite, since your regression test has quietly become a distribution rather than an assertion. And you pay for a record of what happened, because the first question after any incident is what the thing did and in what order.
Agents also fail quietly, and that property should decide the argument. We wrote up the six task types where AI coding agents still come back to a human, and the pattern holds outside coding: reach fails loudly, judgement fails silently, in output that reads perfectly well. If a wrong action would go unnoticed for a week, that is an argument for a workflow, or for a person standing at the point of no return.
Which shape fits which situation
| Situation | Right shape | Why |
|---|---|---|
| A form arrives and always needs the same five steps | Workflow | Nothing has to be decided, and a decision you did not need is a defect you did not need |
| Pulling fields out of documents that all look different | Workflow, model on one step | The judgement is per field, not per path. Your code still owns the sequence |
| Answering questions from your own documents | Chatbot | The corpus is bounded, so the answers can be |
| Qualifying an enquiry and booking a meeting | Chatbot with one action | A single write, on rails, does not need to be an agent |
| Working out why something failed across three systems | Agent | The next query depends on what the last one returned |
| Anything that moves money or cannot be undone | Workflow, or an agent with a human approval gate | The gate is the product decision. The rest is implementation |
| A process nobody has written down yet | None of them | Automating an undocumented process gives you a fast, confident version of whatever people were doing wrong |
What each shape costs to run and to keep
| Shape | Cost to run | Cost to keep |
|---|---|---|
| Workflow | Near zero per execution, or one model call on a single step | Cheap per change, expensive in aggregate. Every new case is another branch, and branches accumulate |
| Chatbot | A model call or two per turn, plus retrieval | Mostly content and boundaries. The fence needs moving as people ask things you did not anticipate |
| Agent | Several model calls per turn, growing with the conversation | Guards, evals and telemetry. The model is the cheap part |
Running cost is the number people put in a spreadsheet, and it is the one you can engineer down. On the assistant we run, a small fast classifier picks the flow and pulls out the entities before the expensive agent is invoked at all, and the price per token between those two roles differs by roughly tenfold. History is capped. Button clicks, a bare postcode and one-word answers inside a flow resolve with no model call at all, in under a second.
Maintenance is the number that surprises people. Every rule you decide the model must never break becomes code with its own tests, and it stays in the repository as long as the feature does. The second year of an agent costs more than the first.
The same product asks the question more than once
We have built and rebuilt one customer-facing assistant on a client marketplace three times in about eighteen months, and the shape answer changed each time.
The first generation was a fixed pipeline. A model read the visitor's message and pulled out the entities, a hand-written decision engine routed by task type, an executor ran the tools, a composer filled in a template, and a second model call rewrote that template into something that sounded like a person. It worked, and it was the wrong shape by the end: every new conversational path meant new template keys and another branch in the decision engine, and the rewriting step roughly doubled the model calls per turn.
The second generation was the agentic one. A cheap classifier picks the flow, then hands to one of four agents, each with its own steps and its own tools, inside a bounded tool-use loop. New paths became prompt and tool changes rather than new branches. The third changed no behaviour at all: it made the model layer swappable, so any provider's model can drive the same agents and the same guards. It also dropped two vendor SDKs the service had been carrying, leaving five runtime dependencies in total.
The part worth stealing is that the outward contract never moved. Same request shape, same response shape, same message topic, same downstream consumer. Two full rewrites of the hardest component in the system, and each time the blast radius was one service.
Two lessons, and the second is less comfortable. The shape question is not answered once, it gets re-asked as the product matures, and what makes re-answering affordable is deciding the boundary before deciding the shape. But we would not have started with the agent either. The pipeline taught us which conversational paths existed. Building the agent first would have meant guessing them, and paying agent prices to find out.
Why the word "agent" on a vendor page means very little
Gartner has a name for this. In a press release dated 25 June 2025 the firm coined "agent washing" for rebranding assistants, robotic process automation and chatbots as agentic without substantial agentic capability, and estimated that only around 130 of the thousands of vendors describing themselves that way were doing something real. The same release predicted that over 40% of agentic AI projects would be cancelled by the end of 2027 on cost, unclear value or inadequate risk controls, a forecast published alongside a poll of webinar attendees that gets quoted everywhere as though it were history.
Buy on behaviour instead of on the noun. What can it call, what can it write to, what happens when one of those calls fails, and can you see afterwards what it did. If the answers are one API, a chat window, an error message and no, you are buying a chatbot, which may be the right purchase at a much better price than the one on the slide.
The same discipline applies inside your own company, which is where an AI use policy sized for a small company earns its keep: the part that matters is the short list of actions requiring a human signature whatever shape the software is. And if you are measuring yourself against everyone else's announcements, the adoption numbers deserve a sceptical read.
Common questions
What is the difference between an AI agent and a chatbot?
A chatbot chooses words. An agent chooses actions. A chatbot answers within a bounded set of topics and hands off at the edge of them. An agent is given a goal and a set of tools, then decides which to call, in what order, and when the job is done.
If a workflow calls a language model, is it an agent?
No. The question is who owns the order of the steps. If your code owns it and the model only fills one in, such as extracting fields or classifying a message, you have a workflow with a model in it. Usually the better design, and far easier to test.
Which shape is cheapest to run?
The workflow, by a distance, because most of its steps cost nothing per execution. A chatbot costs a model call or two per turn. An agent costs several, and the number grows as the conversation does, which is why the biggest savings come from gates that skip the model entirely.
Can we start with a chatbot and move to an agent later?
Yes, and that is usually the right order. Pin the outward contract first: request shape, response shape, whatever sits downstream. We have replaced the inside of one assistant twice in eighteen months without the surrounding systems changing, because that boundary was settled early.
How do we tell whether a vendor's agent is really an agent?
Ask what it can call, what it can write to, what happens when a call fails, and whether you can reconstruct afterwards what it did. Teams who have built one answer quickly. Teams who have wrapped a chat window around a search index answer a different question.
Tell us what you're deciding.
Tell us where you think AI might fit and what data you already hold. On a 30-minute call we'll tell you what we would do first, even if the answer turns out not to involve AI.
More on AI consultancyWhat happens after you get in touch
We reply within one working day
By email, to arrange a time for the call.
A free 30-minute call
With a senior engineer, not a salesperson, about what you are building or deciding.
The fee in writing first
If a first step is worth taking, you get its exact fee in writing before any work starts.