Agentic AI: From Chatbots to Systems That Act
What separates an agent from a chat model, the loop at the heart of every agent, and why reliability is the hard part.
Debanjan Saha
· 3 min read
On this page
A chatbot answers. An agent pursues a goal. The difference is not a bigger model but a different shape of program: instead of one prompt in and one response out, the model runs inside a loop where it can choose actions, observe their results and decide what to do next.
The core loop
Almost every agent, however elaborate, reduces to the same cycle:
- Observe the current state: the task, the conversation, and the results of earlier actions.
- Decide what to do next, which may be calling a tool or giving a final answer.
- Act by executing the chosen tool.
- Repeat until the goal is met or a limit is hit.
This pattern was popularised by the ReAct paper, which interleaves reasoning steps with actions. In code it is surprisingly small:
def run_agent(goal, tools, llm, max_steps=10):
messages = [{"role": "user", "content": goal}]
for _ in range(max_steps):
reply = llm(messages, tools=tools)
if reply.tool_call is None:
return reply.text # the model decided it is done
result = tools[reply.tool_call.name](**reply.tool_call.args)
messages.append(reply) # keep the model's own request
messages.append({"role": "tool", "content": str(result)})
raise RuntimeError("Step budget exhausted")
Notice the max_steps guard. It is not decoration; it is the first line of defence against an agent that loops forever.
What makes something an agent
Four ingredients turn a language model into an agent:
- Tools. Functions the model can call: search, a code interpreter, a database query, an API. Tools are the agent’s hands.
- Memory. Short-term memory is the context window. Long-term memory is anything stored outside it and retrieved when needed.
- Planning. Breaking a goal into steps, and revising the plan when a step fails.
- Autonomy. How much the system decides for itself versus asking a human.
Standards such as the Model Context Protocol aim to make tools reusable, so one tool server can be plugged into many different agents and applications.
Why reliability is hard
Each step in an agent run can go wrong, and errors compound. If every step is 95% reliable, a ten-step task succeeds only about 60% of the time, since 0.95 to the tenth power is roughly 0.6. Common failure modes include:
- Choosing the wrong tool, or passing malformed arguments.
- Misreading a tool’s output and building on the mistake.
- Losing track of the goal as the context fills up.
- Taking an irreversible action on a wrong assumption.
Designing agents that behave
Several practices matter more than model choice:
Constrain the action space
Give the agent a small number of well-described tools with strict input schemas. Validate arguments before executing. Separate read-only tools from those with side effects.
Keep a human in the loop where it counts
Require approval for actions that are expensive, public or hard to undo, such as sending messages, spending money or deleting data.
Make runs observable
Log every step: the model’s request, the tool call, the result. When something goes wrong, you need to replay the trajectory to see where it diverged.
Evaluate trajectories, not just answers
An agent can reach a correct answer by a wasteful or dangerous route. Build test suites of realistic tasks and score cost, step count and safety alongside correctness.
Where this is heading
Agents are most useful today on tasks that are well specified, verifiable and tolerant of retries: coding with a test suite to check against, research with citations to verify, data cleaning with clear rules. As models improve at planning and at recovering from their own mistakes, the boundary of what can be delegated moves outward. The engineering around the model, meaning tools, guardrails and evaluation, is what determines how far you can trust it to go.
Discussion
Comments (Giscus) will appear here. Set
PUBLIC_GISCUS_REPO,PUBLIC_GISCUS_REPO_ID,PUBLIC_GISCUS_CATEGORYandPUBLIC_GISCUS_CATEGORY_IDto enable them.