All English guides
context engineeringAI agent failuresprompt engineering vs context engineeringRAG limitationsAI agent context window

Context engineering: 7 reasons your AI agent keeps failing

Rewriting the prompt again won't fix an agent that read a stale policy or lost track of its own state. Here are seven context failures to check before you touch the model.

LeanX··5 min read

Also available in Korean: Read the Korean original

Short answer: Context engineering means choosing the goals, knowledge, state, tools, permissions, and output format a model needs at the moment it acts. Prompt engineering polishes the instruction. Context engineering builds the whole environment the model judges from.

When an agent gets something wrong, a lot of teams reach for the prompt first. They write longer instructions, add "never make a mistake," and spell out the role in more detail. It helps for a moment, then breaks again on a different input.

That's because the real cause usually sits outside the instruction. The agent might have read a stale policy document. It may have lost state from an earlier step. Or a tool returned a result in a format it didn't expect. A better prompt alone rarely fixes any of that.

Prompt engineering vs. context engineering

DimensionMain questionExample
Prompt engineeringWhat do I instruct, and how?Role, task steps, output format, examples
Context engineeringWhat do I show or hide at the moment of judgment?Retrieved docs, memory, state, tool results, permissions, budget

A good prompt still matters. In real workflows, though, it's only one part of everything the model actually receives.

Seven ways context engineering breaks silently

1. The goal isn't cut into task-sized units

"Do customer support well" is too big a goal. It doesn't say what counts as done, how far the agent can act, or when to hand off to a person.

Cut it down. Try: "Classify the inquiry into one of six types, find two relevant policies, draft a reply, and hold refund, legal, or privacy cases for review." A goal this specific can actually be measured and operated.

2. Retrieved knowledge is wrong or stale

Adding RAG doesn't automatically make an agent factual. If duplicate documents, retired policies, untitled files, and documents with different permission levels all get retrieved together, the model picks whichever one sounds most plausible.

Every document needs an owner, a scope, a version, an effective date, and an expiry flag. Pass the title, last-edited date, and source link along with each retrieved result, so a person can check the basis too.

3. Too much context buries what matters

A bigger context window doesn't mean stuffing in every document is good design. Long meeting notes, unrelated conversation, and duplicate policies mixed together make the model lose track of priority.

Choose only what each step needs. Give the classification step its criteria, the reply step its relevant policy and customer history, and the approval step its risk signals and evidence.

4. State from earlier steps disappears

In long-running work, using chat history alone as memory makes it unclear what's actually been completed. The agent may repeat research it already did, or draft a reply to an email it already sent.

Structured state fixes this. Store a stage like new → researched → draft_ready → pending_review → approved → sent. Then define the transition conditions and the owner for each step.

5. Tool inputs and outputs are ambiguous

A tool with a clear input and output schema is safer than one described loosely as "look up the company in the CRM." Decide upfront which fields are required, such as company name, domain, and country. Also decide whether a no-match result returns an empty array or an error.

Returning a tool's result as one blob of natural language makes the next decision shaky. Return a structured ID, source, date, confidence score, and error code instead.

6. Permissions and approval rules aren't in context

A model knowing the answer and a model having the authority to act on it are two different things. External sends, payments, deletions, and sensitive lookups need their own policy layer.

Give the agent the current user's role, allowed tools, spending limits, forbidden targets, and approval conditions explicitly. Enforce the permission check again in system code. Don't leave it to the model's own judgment.

7. Failures never feed back into an eval set

If an error that happens in production only ever gets discussed in a chat thread, the same mistake repeats. Save the failing input, the context at the time, the model's output, the human's correction, and the error type.

Add that record to a weekly eval set, and it becomes clear whether the fix belongs in the prompt, retrieval, tools, or policy. Swapping the model comes after that, not before.

Why RAG alone isn't enough

RAG retrieves the knowledge a task needs. Context engineering is broader: it also covers current task state, user permissions, tool results, remaining budget, past failures, and output format.

If your internal AI keeps getting things wrong, check the list below before you touch your vector database's top-k setting.

  1. Are the completion condition and forbidden actions clear?
  2. Does the current step receive only the documents it needs?
  3. Can you verify each document's version and source?
  4. Is state stored as data, not just conversation?
  5. Are tool inputs and outputs validated against a schema?
  6. Are permissions and human approval enforced in code?
  7. Do failures accumulate into an eval set?

A one-page context design doc for small teams

You don't need a large platform to start. A single page per workflow works. Write down, for each task:

  • Task goal and completion condition
  • Input data and its trust level
  • Searchable knowledge and its owner
  • State values and transition conditions
  • Available tools and their permissions
  • Human approval conditions
  • Output schema and error codes
  • Eval set and operating metrics

This doc can reduce how often engineers, the business team, and leadership describe the same failure in different language.

Where to start

Pick one workflow that keeps failing, and fill in the one-page doc above before you edit another prompt. These references go deeper into the same ideas:

FAQ

What is context engineering?

It's the design of selecting the goals, knowledge, state, tools, permissions, and output format a model needs at the exact moment it takes action.

Sources

  1. Anthropic — Effective context engineering for AI agents
  2. Anthropic — Building effective agents
  3. Model Context Protocol — official documentation

Want help picking your first AI pilot?

Book a free call