Context Engineering for AI Agents: A Practical Guide to Context Windows, Memory, and Retrieval
AI CORE · Practical systems guide

Context Engineering for AI Agents: A Practical Guide to Context Windows, Memory, and Retrieval

Reliable agents are not created by a clever prompt alone. They are designed by deciding what information enters each model call, what stays out, what is retrieved on demand, and what is preserved after the task. This guide turns context engineering for AI agents into an operating discipline.

Illustration of an engineer assembling context layers for an AI agent

Context Engineering for AI Agents: The Quick Answer

Context engineering for AI agents is the practice of deliberately selecting, structuring, updating, and removing the information an agent receives while it works. That information can include system instructions, the user’s request, prior conversation turns, files, retrieved passages, tool results, intermediate plans, and summaries of earlier work. The goal is not to stuff the largest possible amount of text into a context window. The goal is to give the model the smallest trustworthy set of information that lets it make the next good decision.

Think of an agent as a capable collaborator with a desk. A context window is the active space on that desk, not the entire office archive. The agent needs a clear assignment, relevant evidence, tools, and enough short-term state to continue. It may also need access to a filing system for documents and durable notes. If every document, transcript, log, and prior thought is dumped onto the desk, the useful signal becomes harder to locate and the agent can follow stale or conflicting material. If too little is present, it guesses, repeats work, or asks tools the wrong question.

This distinction matters because an agent is usually a loop, not a single response. It reads a state, chooses an action, observes a result, and repeats. Context is therefore the agent’s control surface. Each turn changes what the model can see and consequently changes the action it is likely to choose. A well-designed context layer makes the desired action obvious; a weak one forces the model to infer missing policy, project facts, or task boundaries.

Bottom line: use the context window for current instructions and evidence, retrieval for external knowledge, long-term memory for durable facts, and compaction for preserving only the useful state of a long-running task. Treat every item as a decision: does it improve the next action enough to earn its place?

This guide focuses on practical design rather than vendor-specific magic. The exact APIs differ, but the underlying work is stable: define instruction hierarchy, model the task state, retrieve evidence with provenance, bound tool outputs, write memories carefully, compact when needed, and evaluate the whole loop. For adjacent foundations, see our guide to how AI agents work with tools and memory.

A Better Mental Model: Context Is a Runtime Assembly, Not a Prompt

The word “prompt” is useful for a direct question to a chatbot, but it is too small for production agents. An agent’s input is assembled at runtime from multiple sources with different reliability, lifetimes, and permissions. A customer-support agent, for example, might receive non-negotiable policy, a customer’s current request, an authenticated account record, a retrieved product policy, a shipping tool response, and a short summary of the conversation. Those pieces should not be treated as one undifferentiated block of prose.

Instead, treat context as a typed bundle. Every element should answer four questions: where did it come from, who or what is allowed to change it, how long should it remain relevant, and what decision is it intended to support? A policy instruction belongs in a protected layer. A tool result should carry a timestamp and source identity. A retrieved passage should retain its document reference. A memory should explain whether it is a user preference, a task fact, or a hypothesis that still needs checking.

This model prevents a common mistake: confusing persistence with relevance. A fact can be important enough to store but not relevant enough to inject into every turn. A user may prefer concise answers, which can be a durable preference. Their question about last week’s billing issue may be a short-lived task fact. A current account balance may be authoritative for one turn but stale a few minutes later. Each has a different context lifecycle.

It also clarifies why a larger context window is not a complete architecture. A larger window can hold more material, and that can remove some immediate constraints. It does not decide which source is authoritative, solve contradictions, ensure a retrieved document is current, stop untrusted text from influencing instructions, or keep a long tool transcript coherent. Context engineering is the work of making those decisions explicit.

Anthropic’s effective context engineering for AI agents frames the discipline around curating what a model sees at each step. The practical implication is simple: a system should optimize the agent’s next decision, not merely maximize the amount of available information. That is an engineering problem involving data contracts, task design, evaluations, and observability.

The Six Context Layers Every Agent Team Should Distinguish

Names vary between products, but separating these layers avoids many design errors. They can be sent together to a model, yet they should remain distinct in your implementation and evaluation. When an agent performs poorly, this separation gives you something concrete to inspect: was the instruction wrong, the working state incomplete, the tool result noisy, the memory stale, the retrieval weak, or the compaction misleading?

1. System instructionsStable rules, role boundaries, safety constraints, output contracts, and tool-use policy. These define how the agent should operate.
2. Working contextThe active task, current plan, recent relevant turns, selected files, and temporary state needed for the next action.
3. Tool outputsStructured observations produced by APIs, code execution, databases, browsers, or other tools. They are evidence, not instructions.
4. Long-term memoryDurable user preferences, stable project facts, verified decisions, and useful lessons that may matter in future sessions.
5. RetrievalOn-demand passages, records, and artifacts selected from a larger corpus because they are relevant to the present query.
6. CompactionA controlled summary or state transition that preserves decisions, open questions, and evidence while retiring low-value history.

System instructions: define the operating contract

System instructions answer questions the user should not have to repeat: what the agent is for, what it must not do, when it needs confirmation, which tools it may use, how it should report uncertainty, and what output format it should produce. These instructions should be clear, testable, and as stable as possible. They are not a place to paste every organizational wiki page. Long generic policies can obscure the few rules that actually matter at decision time.

Good instructions establish precedence and boundaries. For example: “Use the account tool for current account facts; do not infer a balance from conversation text.” Or: “Before performing an external write, summarize the intended change and request confirmation.” These rules guide tool behavior without pretending that a prose instruction can replace permissions in the tool layer. Enforce critical authorization in code as well as language.

Working context: preserve the task’s local state

Working context is the agent’s scratchpad for the current job. It includes the explicit goal, constraints, latest user turn, selected evidence, a compact plan, and the results needed to choose the next action. It should be current and scoped. In a coding agent, it may include the failing test, relevant files, repository conventions, and a diff summary. In a research agent, it may include the research question, inclusion criteria, sources already reviewed, and unresolved claims.

Working context should not become an accidental transcript archive. Retaining a recent exchange can help continuity, but retaining every prior exchange often creates ambiguity. Use a state object where possible: goal, completed steps, pending steps, assumptions, and blockers. The model may still consume it as text or structured input, but the application owns its meaning and can update it deterministically.

Tool outputs: preserve evidence, reduce noise

Tool outputs have special status because agents act on them. A database response, API payload, compiler error, browser page, or search result can be more authoritative than conversational recollection. Keep outputs structured when possible, with identifiers, timestamps, status, and clear fields. The agent should be able to tell a successful response from an error and a fresh result from a cached one.

Do not blindly paste massive raw outputs into the next model call. A long log can contain one useful error line; a search response can contain snippets that are not the primary source; a web page can include untrusted instructions. Filter, truncate, summarize, or render a safe schema before returning it to the model. The agent needs evidence sufficient to decide, not an uncontrolled document dump.

Long-term memory: store durable, attributable knowledge

Long-term memory survives beyond a single task. It can make an agent feel consistent, but it is also a source of subtle risk. Save memories that are stable, useful, attributable, and appropriate to retain: a user’s stated preference, a verified project convention, a confirmed decision, or a recurring workflow choice. Do not automatically convert every message into durable truth. A user may speculate, change their mind, or describe a temporary condition.

Memory needs metadata: source, time, confidence, scope, and a way to revise or delete it. Consider whether a fact belongs to an individual, a project, or an organization. Keep sensitive data out unless the product has a justified retention and access model. Personalization and user control must be designed together, not bolted on afterward; product teams should make memory reviewable, correctable, and removable. For a deeper conceptual comparison, read AI agent memory explained: how context becomes continuity.

Retrieval: find current evidence when it is needed

Retrieval is not memory. Memory is a selected, durable representation of what the system has learned or been told. Retrieval searches a potentially large, changing corpus for material relevant to a current request. A support agent can retrieve the current policy; a code agent can retrieve the relevant module or API documentation; a research agent can retrieve source passages. The retrieved material should remain linked to its origin so that the system and user can inspect it.

Retrieval-augmented generation, popularized in the Google RAG paper, is powerful because it lets a model ground a response in external information without permanently training that information into the model. But retrieval is a pipeline: query formulation, candidate selection, ranking, filtering, context assembly, and answer generation. A failure in any stage can look like model hallucination. Our guide to RAG hallucinations and source grounding explains why retrieved text alone is not a guarantee of truth.

Compaction: make a long task continue without carrying everything

Compaction is the deliberate replacement of sprawling history with a compact state that preserves what matters. A good compaction includes the objective, decisions made, evidence references, completed actions, open questions, constraints, and the next recommended step. It should not claim that an uncertain point is resolved merely because the original nuance no longer fits.

Compaction differs from casual summarization. Casual summarization asks, “What was said?” Operational compaction asks, “What must a future agent see to safely continue this work?” That means retaining identifiers, links, decisions, and uncertainty. It also means discarding greetings, duplicate tool traces, failed paths that no longer matter, and verbose observations that have already been distilled into verified state.

A Practical Context Architecture for an AI Agent

A practical architecture does not need to be complicated on day one. Start with an explicit request boundary and add layers only when the task needs them. The key is to make every transition observable. You should be able to reconstruct which instructions, memories, retrieved documents, and tool outputs shaped an action. That trace is vital for debugging, evaluation, compliance, and user trust.

Step 1: Normalize the incoming task

Convert the user request into a task object. Extract the goal, desired deliverable, constraints, identity or authorization context, and any required confirmation. Do not silently fill gaps that would change a meaningful outcome. A clear task object reduces the temptation to carry an entire conversation into each turn. It also creates a stable place to record what the system believes it is doing.

Step 2: Apply protected instructions and policy

Build the instruction layer from a controlled source, not from a document the user can edit. Include role, output requirements, tool policy, escalation rules, and explicit guidance on untrusted content. Keep the language concrete. “Cite the source field when making a factual claim from retrieval” is easier to evaluate than “be accurate.” When policy is complex, encode hard checks in the application rather than relying exclusively on the model to interpret prose.

Step 3: Load only relevant durable memory

Search memory using the task and retrieve a small set of candidates. Validate their scope before injecting them. A project preference should not override a user’s current request; a memory from another customer should never cross a tenancy boundary; a stale decision should not outrank fresh authoritative data. Add an explicit label such as “saved preference” or “previously confirmed project decision” so the model can reason about provenance.

Step 4: Retrieve external evidence with source identity

Generate a retrieval query from the task, then fetch and rank candidates. Prefer authoritative material for claims that require authority: official documentation, signed records, primary research, controlled internal policies. Attach source title, URL or identifier, version or date when available, and a short excerpt. Let the model see enough context to use the material while preserving a path back to the original. LangChain’s overview of context engineering for agents is helpful here: retrieval is one component of a broader process of selecting and managing information.

Step 5: Plan, act, observe, and update state

The agent should form a bounded plan before it starts a series of tool calls. The plan can be brief, but it makes evaluation easier: were the chosen tools suitable, was the sequence reasonable, and did the agent stop when evidence was sufficient? After each tool call, normalize the result, update task state, and decide whether the next action still advances the goal. Do not let raw output accumulate without limit.

OpenAI’s Agents guide and documentation emphasizes combining models with tools and orchestration. In practice, orchestration is where context engineering becomes visible. The application decides what the model sees next, which tools are available, how results are returned, and when a human must decide.

Step 6: Compact and hand off

When the task becomes long, compact it at an intentional checkpoint. Save a structured handoff that records the task goal, completed actions, current evidence, pending questions, and any conditions that need user confirmation. If a human or another agent later resumes, it should not have to replay an entire transcript to understand the status. A compact handoff improves continuity and makes the system more resilient to retries or interrupted sessions.

Architecture rule: keep authority separate from convenience. A retrieved policy can inform the model; an authorization service should decide whether an action is allowed. A memory can personalize tone; it should not silently create permission. A tool result can provide evidence; it should not become a new instruction just because it appears in the same context.
Illustrated context engineering pipeline showing instructions, task state, retrieval, tools, memory, and human approval
A context architecture should make sources, authority, and handoffs explicit.

Decision Table: What Belongs Where?

The following table is a starting point for design reviews. It is intentionally qualitative. Exact thresholds depend on the task, model, regulatory environment, and tool behavior. The useful question is not “Can this fit?” but “What is the safest and clearest lifecycle for this information?”

Information typeBest homeInject into every turn?Key control
Role, safety constraints, output contractSystem instructionsUsually yes, in concise formVersion control and tests
Current user goal and constraintsWorking contextYes while task is activeExplicit task object
Current account status or inventoryTool output or retrievalOnly when neededFreshness and source identity
User’s stable formatting preferenceLong-term memoryWhen relevantConsent, editability, scope
Product documentation or policiesRetrieval corpusRetrieve on demandVersioning and citations
Long task transcriptCompacted state plus trace storeNo; include a compact handoffPreserve decisions and uncertainty
Raw logs and large payloadsArtifact store; selected excerpts in contextNoFiltering and bounded views
Authorization and entitlementsApplication or policy engineExpose outcome, not sole enforcementServer-side enforcement

Use a context budget even when the model window is large

A context budget is a deliberate allocation of attention, not merely a token limit. Reserve space for stable instructions, the current task, essential evidence, and the expected response or tool call. If a tool result grows, decide what must be kept before adding more. This forces prioritization early and prevents a late-stage failure where the model has seen too much history but not enough room for a useful action.

Budgets should be content-aware. A legal or clinical workflow may need exact citations and quoted evidence. A coding workflow may need a diff, failing test, and a focused interface contract. A support workflow may need the latest policy and account record. There is no universal “best context length.” There is only a defensible selection for the next decision.

Interactive Context Design Check

Use this lightweight check before adding another source to an agent turn. It is not a model quality score. It is a prompt to make the information lifecycle explicit.

Choose an item and timeframe to see a suggested home.

Common Context Engineering Failures and How to Fix Them

Failure: treating the entire chat history as truth

Conversation history contains requests, guesses, corrections, old plans, and sometimes contradictory statements. When every message remains equally visible, the model may follow an outdated instruction or repeat an abandoned approach. Fix this by separating immutable system instructions, current task state, confirmed decisions, and archival history. Summarize the history only after validating what should survive.

Failure: retrieval without provenance or ranking discipline

Weak retrieval can return similar-sounding but irrelevant passages. Worse, it can surface an old version of a policy alongside the current one. Fix this with metadata filters, authority-aware ranking, freshness rules, document versioning, and source citations in the answer. When evidence conflicts, the agent should surface the conflict rather than quietly blend the claims. This matters especially when readers are trying to understand why AI models hallucinate: a fluent answer is not the same as a grounded answer.

Failure: unbounded tool output

Tools frequently return more data than a model needs. Search results, logs, HTML pages, database rows, and command output can all swamp the active task. Fix it at the tool boundary. Return a schema with the essential fields, cap list sizes, offer pagination or detail lookup, and include an explicit indicator when data was truncated. Store the full artifact elsewhere for audit and direct inspection.

Failure: memory that accumulates without governance

An agent that remembers everything will eventually remember something wrong, sensitive, obsolete, or out of scope. Fix this by defining memory write criteria. Require explicit user statements or verified outcomes; attach a source and timestamp; support review and deletion; and expire facts that naturally age. Treat memory as a small knowledge base with stewardship, not as an automatic transcript compression feature.

Failure: summaries that erase uncertainty

A terse summary can improve speed while quietly removing caveats. A future turn then acts as if a tentative inference were confirmed. Fix this by storing uncertainty directly: “hypothesis,” “needs verification,” “source conflict,” or “pending user decision.” Include a link or identifier to the evidence that supports each material conclusion. The best compaction is short but auditable.

Failure: putting security policy only in prompts

Instructions help the model follow expected behavior, but they are not a security boundary. Tool permissions, identity checks, tenant boundaries, and approval requirements need application-level enforcement. The model can propose an action; trusted code should decide whether the action is permitted. This protects against model errors and malicious or irrelevant content embedded in documents, pages, or tool outputs.

Failure: optimizing a single response instead of the full loop

An answer can look excellent in isolation yet be a poor agent action. It may choose an unnecessary tool, retrieve a broad corpus, fail to record state, or create a next step that cannot be verified. Evaluate trajectories: did the agent get to a correct result, use the right evidence, respect constraints, recover from errors, and leave behind a usable handoff? Agent quality is a property of the loop.

Split illustration comparing an overloaded AI agent context with a curated evidence-first context
More context is not automatically better context; selection and lifecycle matter.

How to Evaluate Context Engineering Before You Scale

Start with a small, representative set of tasks rather than a broad benchmark alone. Include ordinary successful requests, ambiguous requests, stale-document cases, conflicting-source cases, long-running tasks, tool failures, and requests that should require human confirmation. For each, record the context assembled, the action sequence, citations or tool evidence, final output, and human judgment. This makes failures diagnosable instead of mysterious.

Useful measures are behavioral and qualitative: whether the agent obeyed instruction precedence, retrieved an appropriate source, cited the source correctly, used a tool only when necessary, preserved user control, recognized uncertainty, and maintained a coherent state after compaction. Also inspect cost and latency as operational constraints, but do not optimize them by stripping away necessary evidence. A cheap answer that confidently acts on outdated policy is not a useful improvement.

Observability makes this practical. Trace the task version, instruction version, memory candidates considered, retrieval query and selected chunks, tool inputs and normalized outputs, compaction events, and final response. Redact sensitive data from traces according to your security model. Our article on AI agent observability and tracing explains why a trace is the bridge between a surprising behavior and a concrete engineering fix.

When you run experiments, change one variable at a time where possible. Compare a raw transcript against a structured task state. Compare top-k retrieval against authority-filtered retrieval. Compare a free-form summary against a schema that requires open questions. The point is not to declare one architecture universally best; it is to learn which context choices improve your specific tasks without damaging safety or maintainability.

Context Engineering Checklist for AI Agent Teams

Before launch

  • Define system instructions, tool policy, and output contracts in versioned sources.
  • Model the active task as explicit state rather than only chat history.
  • Identify authoritative sources for every high-stakes fact type.
  • Keep raw artifacts outside the prompt and return bounded tool views.
  • Define memory write, review, correction, deletion, and expiration rules.
  • Build compaction that retains decisions, evidence IDs, uncertainty, and next steps.
  • Test adversarial and conflicting-content cases, not only happy paths.

During operation

  • Log which context sources shaped each consequential action.
  • Check freshness and tenancy before injecting retrieved data or memories.
  • Ask for human confirmation before external or irreversible actions.
  • Make citations or evidence references visible when users need verification.
  • Monitor repeated tool loops and oversized payloads as design signals.
  • Review compacted state for unsupported certainty and missing blockers.
  • Use production failures to extend your evaluation set and refine contracts.

The checklist is deliberately conservative. Context engineering is where capability, product design, data governance, and safety meet. A small agent can often begin with a task state and one retrieval tool. A mature agent may need tenant-aware memory, document permissions, multi-step orchestration, and audit-grade traces. The principle remains the same: expose only the information that helps the next safe, grounded action, and preserve enough evidence to explain why the action happened.

Sources and Further Reading

This article is an architectural guide, not a substitute for security, privacy, or domain-specific compliance review. Verify product behavior and documentation before implementing a production workflow.

FAQ: Context Engineering for AI Agents

What is context engineering for AI agents?

Context engineering is the practice of selecting, structuring, updating, and removing the instructions, state, evidence, memories, retrieval results, and tool outputs an agent receives at each step. Its purpose is to help the model take the next grounded and appropriate action.

Is context engineering the same as prompt engineering?

No. Prompt engineering focuses mainly on phrasing an instruction. Context engineering includes prompts but also covers the runtime assembly of task state, tools, retrieved knowledge, long-term memory, compaction, authority, and lifecycle rules.

What is the difference between memory and retrieval?

Memory stores selected durable facts or preferences for future use. Retrieval searches a larger corpus for information relevant to the current task. A current policy is usually retrieved; a stable user preference may be remembered.

Should an agent include its entire chat history in every request?

Usually no. Keep the current task, recent relevant turns, and a structured summary of important prior decisions. Preserve the full transcript outside the active context for audit or later lookup when needed.

What should compaction preserve?

Preserve the objective, decisions, completed actions, evidence references, open questions, constraints, uncertainty, and next step. Remove duplicate chatter and raw details that can be recovered from artifacts.

How do tool outputs affect agent reliability?

Tool outputs are often the agent’s most important evidence. Normalize them, preserve source and freshness metadata, bound their size, and keep untrusted content from being treated as instructions.

Can a larger context window replace RAG?

No. A larger window can hold more text, but RAG still helps locate current and relevant evidence from a larger corpus. It also supports provenance, access control, ranking, and selective context assembly.

How should teams evaluate context engineering?

Evaluate complete task trajectories, including instruction following, retrieval quality, tool use, evidence attribution, compaction, recovery from errors, and human approval behavior. Keep traces so failures can be diagnosed and fixed.