Skip to content
Writing
Esc
Type to search all articles.
    All writing

    Agentic AI: From Chatbots to Systems That Act

    What separates an agent from a chat model, the loop at the heart of every agent, and why reliability is the hard part.

    DS

    Debanjan Saha

    · 3 min read

    – claps
    On this page

    A chatbot answers. An agent pursues a goal. The difference is not a bigger model but a different shape of program: instead of one prompt in and one response out, the model runs inside a loop where it can choose actions, observe their results and decide what to do next.

    The core loop

    Almost every agent, however elaborate, reduces to the same cycle:

    1. Observe the current state: the task, the conversation, and the results of earlier actions.
    2. Decide what to do next, which may be calling a tool or giving a final answer.
    3. Act by executing the chosen tool.
    4. Repeat until the goal is met or a limit is hit.

    This pattern was popularised by the ReAct paper, which interleaves reasoning steps with actions. In code it is surprisingly small:

    def run_agent(goal, tools, llm, max_steps=10):
        messages = [{"role": "user", "content": goal}]
    
        for _ in range(max_steps):
            reply = llm(messages, tools=tools)
    
            if reply.tool_call is None:
                return reply.text                      # the model decided it is done
    
            result = tools[reply.tool_call.name](**reply.tool_call.args)
            messages.append(reply)                      # keep the model's own request
            messages.append({"role": "tool", "content": str(result)})
    
        raise RuntimeError("Step budget exhausted")

    Notice the max_steps guard. It is not decoration; it is the first line of defence against an agent that loops forever.

    What makes something an agent

    Four ingredients turn a language model into an agent:

    • Tools. Functions the model can call: search, a code interpreter, a database query, an API. Tools are the agent’s hands.
    • Memory. Short-term memory is the context window. Long-term memory is anything stored outside it and retrieved when needed.
    • Planning. Breaking a goal into steps, and revising the plan when a step fails.
    • Autonomy. How much the system decides for itself versus asking a human.

    Standards such as the Model Context Protocol aim to make tools reusable, so one tool server can be plugged into many different agents and applications.

    Why reliability is hard

    Each step in an agent run can go wrong, and errors compound. If every step is 95% reliable, a ten-step task succeeds only about 60% of the time, since 0.95 to the tenth power is roughly 0.6. Common failure modes include:

    • Choosing the wrong tool, or passing malformed arguments.
    • Misreading a tool’s output and building on the mistake.
    • Losing track of the goal as the context fills up.
    • Taking an irreversible action on a wrong assumption.

    Designing agents that behave

    Several practices matter more than model choice:

    Constrain the action space

    Give the agent a small number of well-described tools with strict input schemas. Validate arguments before executing. Separate read-only tools from those with side effects.

    Keep a human in the loop where it counts

    Require approval for actions that are expensive, public or hard to undo, such as sending messages, spending money or deleting data.

    Make runs observable

    Log every step: the model’s request, the tool call, the result. When something goes wrong, you need to replay the trajectory to see where it diverged.

    Evaluate trajectories, not just answers

    An agent can reach a correct answer by a wasteful or dangerous route. Build test suites of realistic tasks and score cost, step count and safety alongside correctness.

    Where this is heading

    Agents are most useful today on tasks that are well specified, verifiable and tolerant of retries: coding with a test suite to check against, research with citations to verify, data cleaning with clear rules. As models improve at planning and at recovering from their own mistakes, the boundary of what can be delegated moves outward. The engineering around the model, meaning tools, guardrails and evaluation, is what determines how far you can trust it to go.

    Enjoyed this?

    A clap helps others find it.

    – claps

    Discussion

    Comments (Giscus) will appear here. Set PUBLIC_GISCUS_REPO, PUBLIC_GISCUS_REPO_ID, PUBLIC_GISCUS_CATEGORY and PUBLIC_GISCUS_CATEGORY_ID to enable them.

    Keep reading