All English guides
AI agent harnessagent execution environmentAI agent permissionslong-running AI agentsagent memory design

The AI agent harness: what to define before a longer prompt

Before you write a longer prompt, set up the agent's harness. Define what it can read, which tools it can use, and what it carries forward. Then set how far it can act and when it stops for a human.

LeanX··4 min read

Also available in Korean: Read the Korean original

Short answer: Before you write a longer prompt, define the agent's harness. It has five pieces: files to read, tools, working memory, an execution scope, and a rule for when to stop for a human. Together, these five pieces let the agent do the same job again in a way someone can review.

"Read this document and draft a proposal" sounds simple. But real work also decides where the material lives, whether it can leave the company, and who reviews the result. Without that context, a longer prompt alone might produce one good-looking result. You won't know what to trust or fix next time.

OpenAI's May 2026 Agents SDK Build Hour looks at long-running agents that work across documents, tools, and systems. It explains that they need files, tools, memory, an execution environment, and a controlled sandbox. The talk covers these components in general, not one team's performance numbers. The table below applies that idea to a repeated team task, as LeanX's own suggestion.

Look at the agent's workspace before its output

When you hand work to a person, you also tell them where the files live and which systems they can use. You explain what they're allowed to touch and how the work gets accepted. An agent needs the same. "Just do a good job" doesn't tell you which reference it used when the source changed, which tool failed, or when it should have stopped.

Your first design doesn't need a complex multi-agent setup. Take a weekly customer-interview summary as an example. The agent can read a folder of anonymized transcripts, draft into a fixed summary format, and let a reviewer edit before anything is final. The scope of that workspace is also the boundary of quality and accountability.

One table for the five harness pieces

PieceQuestion to answer firstExample first version
Input filesWhat can it read, and where's the current version?Read only 20 approved interviews plus the template
ToolsWhich systems can it read or change?Search documents and save drafts only
Working memoryWhat should carry over to the next step?Sources used, open questions, unresolved items
Execution permissionHow far can it act without a person?Up to drafting; a reviewer approves anything sent out
Stop & reviewWhen does it stop, and who checks?Pause on missing evidence, sensitive data, or a tool error

What matters in this table isn't a perfect answer. It's surfacing the blank spots. If a source or an owner is missing, that's a work rule a person still needs to set before the agent takes over. Running an AX (AI transformation) self-assessment to separate repeat tasks from approval points makes it easier to pick that first scope.

Split permissions into read, draft, and execute

Splitting an agent's permissions into just "can" and "can't" creates dangerous edge cases. Read is the permission to find material. Draft is the permission to produce an internal result. Execute is the permission to change something external, like a system, a customer, or a record. Most first pilots don't need more than read and draft.

A quote-drafting agent, for example, can read an approved price list and write a draft into an internal document. But sending the email to the customer, changing a price, or finalizing a contract still needs a person's approval. This split doesn't shrink what the agent can do. It protects how far you can roll things back if something goes wrong.

Working memory should be a handoff note, not a chat log

Long-running work doesn't need every conversation kept forever. It needs a handoff note the next run can actually check. That means the material version used, the criteria decided, what's still unconfirmed, why a human made a change, and what to ask next. This can reduce repeat failures.

Don't keep sensitive customer data or secrets longer than necessary. Set a retention period and who can access it, and consider storing a reference location or an anonymized case number instead of the raw text. Follow your security and privacy team's policy for how long any of this is kept.

Start with one task, one output, one reviewer

Aiming straight at "automate our whole sales process" makes it impossible to tell where something broke. Pick one task that repeats within a week, one output in a fixed format, and one person who reviews it. Look at 10 to 20 representative cases. Collect the reasons drafts got accepted, edited, or held back. That's where your next fix comes from.

Agent harness card

Task name / task owner:
Trigger condition and output:
Files to read and where the current version lives:
Allowed tools, and read / draft / execute permission:
Handoff note to leave for the next run:
Human reviewer and approval condition:
When to stop immediately, and how to roll back:
Reasons for accepting, editing, or holding a draft, and the next fix:

This card is a work agreement, not a technical spec. You learn more from the day an agent correctly flags something to hold than from the day it goes well. Once the execution environment is set, you finally have a fair basis for comparing models, prompts, and tools.

Where to start

Pick one task that repeats within a week and is easy to reverse. Fill in the harness card above for it. Keep permissions at read and draft, and name one reviewer. Then run 10 to 20 representative cases and note why each draft was accepted, edited, or held.

FAQ

What is an agent harness?

It's the work environment that lets an agent finish a job safely and repeatably. It bundles files, tools, memory, execution permissions, and stop-and-review criteria.

Should you build a long-running agent from the start?

No. Start with one task that repeats but is reversible, and confirm the output format and review rules before you expand its scope.

What permissions should you give an agent?

Separate read, draft, and external-execution permissions. High-impact actions, such as sending customer email, changing prices, or finalizing contracts, should need explicit human approval.

Sources

  1. OpenAI — Build Hour: Agents SDK

Want help picking your first AI pilot?

Book a free call