Agentic AI Stop Conditions: A Practical Guide to Time, Cost, and Action Limits
An agent that can keep trying is not necessarily an agent that should. This engineering guide turns an open-ended task into an explicit run-budget policy: the conditions that let an agent finish, pause for review, retry a bounded operation, or stop before it creates a bigger problem.

Stop Conditions Are the Missing Half of Agent Design
An agentic system needs more than an objective and a tool list. It needs a reliable answer to a less glamorous question: when is this run no longer allowed to continue? A stop condition is a machine-checkable rule that changes the run state when a limit, a risk signal, or a required decision is reached. Depending on the policy, the next state may be completed, paused for human review, retried under tighter bounds, or ended with a clear failure record.
This is narrower than a general discussion of agents. The source pillar, AI Agents vs. Agentic AI: What Is the Difference and When Does It Matter?, explains the broader distinction and why systems that take actions deserve different design thinking. This supporting guide focuses on one operational control: a run budget. It is the boundary around one attempted job, not a claim that an agent is safe in every context.
A run budget makes several scarce things explicit: elapsed time, model and tool spending, the number and class of actions, the quality of evidence, and the degree of uncertainty that remains. It also makes escalation explicit. Instead of silently looping, an agent can say, in structured form, “I found conflicting account history,” “this refund changes a customer balance,” or “the available evidence does not meet the threshold.” That is useful information, not a failure to hide.
The word “stop” can sound absolute, but a good policy has more than one terminal behavior. A harmless, transient tool failure might permit one idempotent retry. A deadline may return the best supported partial result. A high-impact action should pause before it is sent. A policy violation should terminate immediately. The important distinction is that the response follows a predeclared rule rather than an improvisation made after the system has already crossed a boundary.
Think of the Run Budget as a Contract
A run budget is a small contract between product, engineering, operations, and the person affected by the outcome. It answers what this run is allowed to consume, change, infer, and defer. It should be created from the task type and context, not copied unchanged across every agent. Summarizing an internal document, drafting a customer reply, issuing a refund, and modifying production configuration have very different failure costs.
Keep the policy separate from the prompt. Prompt instructions can describe an objective and desired behavior, but they are not a dependable enforcement layer. The process that dispatches tools should hold the authoritative counters, permissions, approval gates, and state transitions. This separation also gives reviewers a concise place to audit the rule, instead of reconstructing it from natural-language instructions and model output.
A policy should distinguish a hard limit from a soft threshold. A hard limit always blocks continuation: no payment action after a permission check fails, no more tool calls after the maximum has been reached, and no execution after the deadline. A soft threshold asks for a different mode, such as shrinking the scope, requesting missing information, or asking a reviewer to approve a proposed action. Soft thresholds are valuable because they preserve useful work without converting every ambiguous case into an unsafe action.
One more distinction matters: a completed run is not necessarily a correct run. “The API returned success” and “the customer received the intended outcome” are different statements. Completion criteria should describe the evidence needed to declare success. Stop conditions describe when the run cannot responsibly seek more evidence or take another action. Both belong in the same policy.
Eight Stop Conditions Engineers Should Design Deliberately
1. Deadline or elapsed-time limit
Time limits prevent an agent from becoming a background process with an unclear owner. Measure wall-clock duration from a defined start point and include waiting time if the user experience depends on it. For asynchronous work, a deadline may mean “pause and create a resumable case,” not “discard everything.” Preserve the task input, the last verified state, and a concise reason so a later worker does not repeat unsafe actions.
Use a deadline to manage service expectations, resource contention, and stuck orchestration. Do not treat it as proof that a result is adequate. If the deadline arrives before the evidence threshold, return a partial answer labelled as incomplete or route the case for review. A timer should be checked by the orchestrator before each new step, not only at the beginning or end.
2. Step and tool-call limit
A step limit caps the number of model decisions, tool invocations, or state transitions in one run. It is a direct defense against accidental loops: search, summarize, search again, and repeat without gaining new evidence. Count the meaningful units separately. One model turn may choose several tools; one tool request may fan out into multiple remote operations. If these distinctions matter for cost or risk, the policy should expose them rather than hiding them in one vague “steps” counter.
When the limit is close, an agent can switch to a finalization mode: summarize verified facts, identify missing inputs, and avoid opening fresh lines of investigation. That behavior should be deterministic in the orchestrator. A model should not receive an unlimited final chance to “just check one more thing.”
3. Cost limit
Cost limits cover model usage and tools that charge per request, record, search, transaction, or compute unit. Set the budget using the pricing information available to the service owner, then reserve headroom for the final response and safe cleanup. The policy should track actual or conservative estimated cost as calls complete. Where exact cost arrives later, a conservative preflight estimate can still block an action whose worst plausible spend exceeds the run allowance.
Do not promise a universal number. A reasonable ceiling depends on the workflow, provider contracts, expected value, and who bears the cost. The implementation goal is simpler: an agent must not decide to spend beyond a configured budget merely because the next call might improve its answer.
4. Action-risk limit
Actions differ by reversibility, scope, sensitivity, and harm if wrong. Reading a public help article is unlike changing a shipping address; drafting a refund is unlike submitting it. Assign actions to risk classes and bind each class to a disposition. Low-risk, read-only actions may proceed within budget. Medium-risk actions may require a confirmation or a second validation. High-risk, external, destructive, financial, or sensitive-data actions should generally pause for an authorized reviewer unless a carefully governed exception exists.
Risk is not a score that makes accountability disappear. A classifier can help route work, but deterministic properties are stronger: does the action write external state, affect money, disclose data, widen permissions, or become difficult to undo? The NIST AI Risk Management Framework frames risk management as a governance activity, and its Generative AI Profile is a useful reference when teams map generative-AI risks to controls. Use such frameworks to structure decisions; still implement the specific gate in your system.
5. Uncertainty or insufficient-evidence limit
Agents frequently face missing facts rather than obvious errors. An uncertainty limit tells the system when it has too little verified support to make a recommendation or execute an action. Avoid treating a model’s confidence wording as enough. Build evidence requirements around the task: an account change may require a verified identifier and a current authorization record; a research response may require source agreement and traceable citations; an extraction task may require all mandatory fields to pass validation.
When the threshold is not met, the correct response is often to ask a targeted question, present alternatives, or escalate. This avoids a common anti-pattern: continuing tool calls simply because the agent has not found an answer it likes. The policy should name the missing evidence and record whether the gap is retrievable.
6. Conflicting-evidence limit
Conflicts deserve their own condition. A system can have plenty of evidence and still lack a safe basis for action because two authoritative records disagree. Define source precedence where it is legitimate to do so, such as a current system of record over a cached display. If the conflict cannot be resolved by that rule, pause. Do not tell the model to choose the most plausible record. The run log should include the conflicting values, their sources, timestamps where available, and the decision not to act.
7. State-change or concurrency limit
An action may become unsafe because the world changed after the agent planned it. Before a write, re-read the relevant version, balance, entitlement, or approval state. If it no longer matches the value used in planning, stop and re-evaluate. This is a familiar concurrency control principle expressed in agent workflows. It prevents an agent from applying an otherwise valid action to stale state.
Make idempotency part of this policy. A retry after an ambiguous network failure should use a stable idempotency key or an explicit status lookup, not blindly submit the action again. The question is not only “may we retry?” but “can we prove the prior attempt did not already change the state?”
8. Tool error and retry limit
Tool errors need a taxonomy. A timeout, rate limit, validation error, authorization failure, and remote service outage have different meanings. Retry only errors that are plausibly transient and only where the operation is safe to repeat. Set a small bounded retry count, use backoff appropriate to the dependency, and stop early when the error class rules out success. An authorization error is not made safer by repeated requests; a validation error usually means the input or schema needs attention.
A Run-Budget Field Template You Can Adapt
The following is a downloadable-style specification table: copy it into a design document, configuration record, or change request. The values are intentionally blank or qualitative. Filling them is a product and operational decision, not an invitation to use arbitrary defaults.
Version this template with the workflow. A policy revision should be reviewable like any other behavior change. It is especially important to retain the reason for a limit, because a future team may otherwise raise it to remove friction without understanding the risk it controlled.
Worked Scenario: A Customer Refund Request
Consider an agent that helps a support representative resolve a refund request. The agent can read the order, payment, shipment, and support systems. It can draft a proposed resolution. It may submit a refund only if the organization has explicitly authorized that action class and the policy conditions are satisfied. This example is illustrative; it is not a universal refund policy.
The run begins when a ticket with an order identifier arrives. The orchestrator creates a trace ID and attaches a short deadline, a bounded number of read operations, a cost allowance, and an action policy. It permits read-only lookups and drafting. It classifies a submitted refund as a high-impact external write because it changes financial state and may be hard to reverse.
- Validate the request. The system verifies that the ticket’s account is authorized to discuss the order and that the order identifier has the expected format. A malformed or unverified identifier is an insufficient-evidence stop: ask for clarification rather than search broadly for a likely match.
- Collect minimal evidence. It retrieves the current order status, payment state, return status when relevant, and any existing refund record. Each tool result is associated with the trace. If the system finds a prior refund with the same idempotency reference, it stops as complete or routes the ambiguity to a reviewer instead of creating another refund.
- Check for conflict. Suppose the support record says “approved,” while the payment system shows a disputed charge already under review. Those records point to different processes. The conflict rule pauses the run with a concise explanation and proposed next step. It does not ask the model to decide which department is right.
- Prepare, do not execute. If evidence is consistent, the agent drafts the refund amount and rationale. It re-reads the payment state immediately before any write. A changed payment version or new dispute status triggers a stale-state stop.
- Apply the risk gate. If the refund requires human approval, the agent sends an escalation payload: ticket summary, verified order facts, proposed amount, applicable policy reference, and a preview of the action. The agent has reached a successful pause, not an unfinished loop.
- Submit with safe retry handling. If the action is approved and delegated, the refund request uses an idempotency key. A timeout after submission is not a reason to submit again. The system first queries the payment provider for the idempotency key or transaction status. Only a policy-approved, provably uncommitted request may be retried.
Notice how this policy creates better customer communication. The agent can say that a request is awaiting review because account records conflict, rather than inventing a denial or making a financial change based on an incomplete picture. The agent also produces a clean audit trail for the representative who takes over.
Evaluate Stop Behavior, Not Just Final Answers
A test suite for an agent should include cases where the correct outcome is not a completed task. Test the state transition and the evidence recorded, not only the natural-language response. The Anthropic guide to building effective agents emphasizes choosing the simplest workable pattern and making systems understandable; bounded stop behavior helps make that pattern observable. Where an SDK offers safety and guardrail hooks, such as the OpenAI Agents SDK guardrails documentation, treat the hook as an integration point—not as a replacement for workflow-specific policy.
| Test case | Expected disposition | Assertion |
|---|---|---|
| Deadline expires before evidence is complete | Pause or partial result | No new tool call occurs after the deadline; trace names the missing evidence. |
| Agent reaches final allowed tool call | Finalize or escalate | Counter is exact; no hidden retry starts a new attempt. |
| Cost estimate for next action exceeds budget | Stop before dispatch | Tool is not called; remaining budget and reason are logged. |
| High-impact write lacks approval | Pause for review | Preview is preserved; external state is unchanged. |
| Two authoritative records disagree | Escalate | Both values and sources appear in the escalation payload. |
| State version changes after planning | Re-evaluate or pause | Original planned write is not submitted against the new version. |
| Read-only dependency returns a transient timeout | Bounded retry | Retry count and delay comply with policy; final failure is explicit. |
| Write request times out after dispatch | Check status before retry | Idempotency lookup happens before any second write. |
Add adversarial tests too. Feed the agent a prompt that asks it to ignore a deadline, an instruction hidden in retrieved content, or a request to split an action into smaller calls to bypass a per-action threshold. The enforcement layer should hold. Test malformed tool outputs, missing trace fields, duplicated webhooks, and reviewers who reject an escalation. A safe policy is only real when these unglamorous paths are exercised.
Operating metrics that reveal policy quality
Use metrics to diagnose policy behavior, not to reward maximum autonomy. Useful measures include: the distribution of stop reasons; the share of runs that end in a clean completion, safe pause, retry exhaustion, or policy denial; tool calls and elapsed duration per task type; budget headroom at completion; the rate of escalations later approved, modified, or rejected; duplicate-action prevention events; stale-state blocks; and error classes by dependency. Review representative traces alongside aggregates. A low escalation rate could mean a well-designed workflow, or it could mean the gate is not being reached.
Common Mistakes That Defeat a Good-Sounding Policy
Another mistake is turning every stop into a generic error. A user and an operator need different, useful messages: “I need an account identifier,” “this request is awaiting approval,” “the payment record changed while I was preparing the action,” and “the service was unavailable after a bounded retry.” Do not expose confidential system details, but do expose the category of next step when it is safe to do so.
A Practical Build Sequence
1Inventory actions and states. Draw the workflow’s external reads, writes, approvals, retries, and terminal states. Identify which writes are irreversible or high-impact.
2Name the evidence for success. For each action, list the verified records and freshness checks required before dispatch. Decide what happens when records conflict.
3Define budget fields and dispositions. Put deadline, step, cost, retry, and action-class rules in a versioned policy object. Include an owner and escalation destination.
4Enforce at the orchestration boundary. Check policy before tool dispatch and after every result. Keep tool credentials scoped so a bypass is difficult. Record counters independently from model text.
5Build idempotency and resumption deliberately. Give write actions stable identifiers and make paused work resumable only after revalidating time-sensitive facts.
6Test stop paths first. Simulate exhaustion, conflict, stale state, malformed results, and timeout ambiguity. Verify that no external action occurred when a gate should block it.
7Review traces and refine. Adjust a rule only after examining real task outcomes, reviewer feedback, and the failure mode the rule is intended to control.
The pseudocode is illustrative, not production-ready. Real systems need authentication, authorization, concurrency control, durable state, structured error handling, observability, and workflow-specific validation.
Frequently Asked Questions
What is the difference between a stop condition and a guardrail?
A guardrail is a broad term for a control that constrains behavior. A stop condition is a specific, observable rule that transitions an individual run to complete, pause, retry, escalate, or fail. A run budget often combines several stop conditions with action permissions and logging.
Should every agent have a cost limit?
Every production-like workflow should have a deliberate consumption policy, even if the practical limit is enforced at a service or account level. The right mechanism depends on whether the agent uses metered models, paid tools, or scarce shared capacity. The core principle is that spending is bounded by policy rather than an unending loop.
When should an agent stop instead of asking a human?
Stop immediately for a hard policy violation, a forbidden action, or an unrecoverable error. Pause or escalate when a human can resolve missing authorization, conflicting evidence, or a high-impact decision. The disposition should be defined before the run starts.
Can a confidence score decide whether an agent may act?
It can be one signal, but it should not stand alone for consequential actions. Prefer explicit evidence requirements, validated fields, source precedence, freshness checks, and approval gates. Model confidence is not the same as verified authority or correctness.
How many retries are safe?
There is no universal count. Retry only transient, eligible errors; use bounded attempts; and require idempotency or a status check for writes. Authorization and validation failures usually need correction or escalation, not repeated calls.
What should an escalation contain?
Include the task summary, trace ID, verified evidence, missing or conflicting facts, proposed action or decision, relevant policy rule, and the exact approval requested. Keep it concise enough for a reviewer to act without redoing the investigation.
Conclusion: Make Stopping a First-Class Capability
Reliable agentic systems do not prove their value by taking every action they can imagine. They prove it by recognizing when a bounded run has enough evidence to finish, when a transient failure merits a safe retry, and when authority or certainty is missing. A run-budget policy turns those judgments into inspectable engineering behavior.
Start with one workflow that has real external consequences. Define its time, cost, step, action, evidence, state, and retry boundaries; test the stop paths; then review the traces with the people accountable for the outcome. For the wider architectural context behind these controls, return to AI Agents vs. Agentic AI: What Is the Difference and When Does It Matter?.
