AI agents: the learning route

AI agent planning: goals, observations and stopping rules

AI agent planning is the process of selecting steps and revising them as observations become available. A useful plan connects each step to a desired outcome, its prerequisites and a check. Planning proposes how to proceed; execution controls still determine which actions are permitted and when the system must stop.

Real team working together around laptops

A long list of steps can look impressive while hiding a missing input. This guide helps you make a plan testable. We will design a route to an introductory lesson and revise it when a catalogue lookup reveals an unmet prerequisite.

Key ideas

  • Specify the outcome before listing actions.
  • Treat the plan as revisable when evidence changes.
  • Tie each completed step to an observation or check.
  • Budget exhaustion is a stopped run, rather than completed work.

Define a result that can be checked

'Learn AI' is too broad to tell a system when to finish. 'Find an introductory lesson in the approved catalogue and return its prerequisite' is inspectable. Record the input source, permitted operations and acceptable output. If the learner's goal is unclear, define a clarification step instead of searching every subject. The result should tell a person what the system found and which information it still lacks.

Attach outcomes to steps

A planning step such as 'search' names an action, but not why it is needed. Write 'retrieve a lesson record for the chosen subject' and state the expected evidence. The next step can check the prerequisite only after that record exists. This makes dependencies visible. A plan with explicit prerequisites can pause on a missing record instead of continuing to an unsupported recommendation just because there are more items left on the list.

[2]

Revise from what actually happened

The loop needs to compare observations with the plan's assumptions. If the returned lesson is advanced, search for the required introduction or report that the approved catalogue lacks it. If the lookup fails, keep the failure distinct from an empty catalogue. ReAct is a research example of combining decision-making with actions and observations. For this classroom exercise, the important habit is simple: use new evidence to change the next step and preserve the reason for the change.

[1]

Design stopping conditions

Define success, inability to complete, awaiting review and a run limit as different outcomes. In the Python exercise, a step budget prevents endless proposals. A real application may also need limits on elapsed time, calls and usage. Set those limits in the controller rather than relying on the model to remember them. When a limit is reached, return what is known and what remains unresolved, without claiming that the original task has been completed.

[3]

Evaluate the plan and the final result

Check both whether the steps were allowed and whether they produced the requested outcome. A perfectly formatted plan can still recommend a lesson outside the catalogue. Save the goal, source identifiers, completed steps and final check in a compact record. If several agents share the work, use that record to distinguish a completed subtask from a pending handoff. Review failed examples to decide whether the goal, available tools or plan transitions need improvement.

In everyday language

A plan is a route on a map. Observations tell you whether a road is actually open. A good planner changes the route when the evidence changes, while the controller keeps you inside the allowed area and tells you when the trip must stop.

Try it yourself

Plan a catalogue recommendation in four steps. Add a missing-record branch and an advanced-lesson branch. Give every step an expected observation and define a three-call budget for the teaching exercise.

Expected result

A plan with explicit dependencies, honest unresolved outcomes and a stop when the configured call budget is reached.

Check your answer: If a tool call succeeds, has the whole task succeeded?

Only if its result satisfies the original acceptance check. A successful call can still return an irrelevant or incomplete record.

Questions

Should an agent follow its first plan exactly?

A plan should guide work while remaining revisable. New observations can reveal a missing prerequisite or an incorrect assumption. Preserve the goal and constraints, explain the change and check the revised path. Blindly following the first plan can turn an early error into later unsupported steps.

How do I stop repeated actions?

Define a step budget and track completed actions in the controller. Decide how to handle an identical request that produces no new information. Your policy may stop, ask for clarification or use an alternative permitted lookup. Test the rule rather than relying only on a prompt asking the model to avoid repetition.

Make every step explainable through its outcome and evidence. A useful plan can revise, pause and finish without pretending a limit is success.

Sources and further reading

  1. Yao et al. — ReAct: Synergizing Reasoning and Acting in Language Models ↗Sources checked:
  2. Anthropic — Building effective agents ↗Sources checked:
  3. LangChain — LangGraph overview ↗Sources checked: