Context Engineering for AI Agents: Less Context, Better Decisions

Why agents need curated working context, just-in-time retrieval, compaction and durable state instead of ever-growing prompts.

A curated set of context cards selected from a large archive

An agent does not become dependable simply because more tokens fit in its window. It improves when the right evidence, instructions and state arrive at the moment of decision.

Context is the agent’s working state

Context includes system instructions, conversation history, tool definitions, retrieved documents, tool results and temporary plans. Anthropic frames context engineering as curating this full token set, not merely polishing prompt wording. The objective is a small, high-signal state that supports the next action.

More material can introduce stale rules, duplicated evidence and distracting tool output. Capacity and relevance are different constraints. Teams should track what enters context, why it is there and how long it remains authoritative.

Retrieve information just in time

Store durable information outside the prompt and keep lightweight references such as file paths, document IDs or search queries. Let the agent fetch the specific segment it needs. This pattern reduces repeated payloads and makes provenance easier to inspect.

Retrieval needs boundaries: permission checks, version metadata, source ranking and limits on untrusted instructions. A tool result is data, not an automatic extension of the system prompt. Separate content to reason about from commands the agent is allowed to follow.

Compact without erasing commitments

Long-running work eventually needs compaction. A useful summary preserves goals, accepted decisions, completed work, unresolved blockers, artifact locations and verification evidence. It discards repeated chatter and bulky outputs that can be fetched again.

Make summaries structured and reviewable. If a commitment matters—such as ‘do not send without approval’—store it as explicit state rather than hoping a narrative summary retains it. Link the summary to source artifacts so a later session can recover detail.

Use subagents to isolate noisy work

A focused subagent can search a large corpus or test several approaches in a clean context, then return a compact result. Isolation keeps exploration debris out of the coordinator’s decision state. It also enables parallel work where tasks are genuinely independent.

The coordinator still needs acceptance criteria and verification. A confident summary is not proof that a file exists or an action succeeded. Require paths, source links, test output or read-back evidence appropriate to the task.

Observe context as a system metric

Log context size, retrieval sources, compaction events, tool payload size and the age of state used for decisions. Sample failures for irrelevant or missing context. Cost and latency matter, but so does whether the system acted on an obsolete fact.

A good context architecture makes failures legible. Engineers can see whether the agent lacked evidence, ignored a rule or chose poorly despite having both.

How to put the idea into practice

Begin with one bounded workflow related to context engineering for ai agents: less context, better decisions. Write a one-page baseline before changing the system: current completion time, quality checks, common failure categories, escalation path and the person responsible for the outcome. Select a representative sample rather than only the cleanest examples. Include ordinary cases, difficult edge cases and at least one case where the correct result is to stop or ask for more information.

Run the candidate beside the existing process before allowing it to replace that process. Review both successful and failed outputs, because a lower error rate can still conceal a new high-impact failure. Record the exact configuration used for each test and keep artifacts that let another reviewer reproduce the result. At the end of the pilot, decide whether to expand, revise or stop using thresholds agreed in advance—not a retrospective impression of the best demonstration.

Questions to ask a vendor or internal team

Ask what evidence supports the central claim, which system and data versions produced that evidence, and what conditions were excluded. Request results for the languages, input types and risk categories your deployment will encounter. Ask how changes are announced, how regressions are detected and how a customer can export logs needed for an incident review.

Also ask what the system does when confidence is low, a dependency fails or the request falls outside its supported scope. A dependable product should have a defined failure state, not merely a more polished answer. Ownership matters: identify who can pause the workflow, who approves exceptions and who informs affected users if a material error escapes into production.

A practical checklist

  • Define the user task and the failure that matters before choosing a model or tool.
  • Keep a small, versioned test set drawn from real work, including awkward and adversarial cases.
  • Record model, prompt, tools, retrieval settings, data version, latency and cost for every run.
  • Require human confirmation for irreversible, high-impact or externally visible actions.
  • Review failures by category, not just by one average score, and add regressions to the test set.

Related reading from Meydo Journal

Primary sources