AI agent permissions: split read, draft, and execute
Give an AI agent one stage at a time. Let it read the files it needs and draft into a review space. Execute only after a specific approval. Here's how to split permissions by stage.
Also available in Korean: Read the Korean original
How far can you let an AI agent go on its own?
Split its permissions into three stages: read, draft, and execute. At first, the agent only reads the files it needs and saves drafts to a set review space. Some actions affect the outside world, such as sending, paying, or deleting. It's best to give those a separate execution step that confirms the target and the content first.
"Read this email and reply" bundles four separate jobs: opening the email, finding the policy, writing the reply, and actually sending it. To make AI automation something people can trust, both the user and the system need to see those separate steps and respect them.
What a recent study shows about the limits of input screening
On September 10, 2026, Check Point Research published the PuzzleMask study. It looks at conditions where a fast screening model misses a request hidden in plain prose. A downstream model with more reasoning and tool access can then read that hidden request, and it is the model that holds the tools. The researchers are careful to note this doesn't prove the downstream model breaks its own safety rules. One experiment shouldn't be generalized into "every AI service is vulnerable this way."
A GeekNews summary also covers the study. Turn it into a design question instead: when screening misses something once, what can the AI actually do next? Anthropic's May 2026 post on containment design covers this too, describing limits on execution environment and access scope, not just approval prompts.
Split permissions by task stage
| Stage | Allow | Needs separate confirmation |
|---|---|---|
| Read | Look up the assigned inbox or policy docs | Access another team's or customer's records |
| Draft | Save a reply or edit to a review space | Overwrite the original policy or customer record |
| Execute | Apply approved content to an approved target | Run with a changed recipient, amount, or target |
This table is LeanX's suggestion for teams, not a fixed rule. "Read-only" doesn't mean the agent can read everything. Scope lookups to the customer or time range the task needs, and check whether results can be forwarded to an outside address.
Make the approval prompt specific
"Approve this AI action?" tells the reviewer nothing. Show what will change before they approve, and label the button with the real action. The wording below is an example LeanX suggests, based on public sources.
Draft reply ready for review.
To: customer address confirmed by the rep
Content: 1 reply about a delivery date question
Attachments: none
Basis: current delivery policy document
Flag: customer's requested date is not confirmed yet
[Edit draft] [Send this to 1 recipient]
If the recipient or attachments change after approval, don't reuse the old approval. The server should check that the actual target and content still match what was approved. What matters isn't that a button was pressed. It's which action was approved.
Five checks for the team and engineering to run together
- Scope the target: define which folders it can read, which records it can edit, which addresses it can send to.
- Separate the tools: saving a draft and actually sending should use different functions with different permissions.
- Verify content right before execution: recheck recipient, amount, attachments, and required fields.
- Prevent duplicates: keep an execution ID and status so a retry after a timeout doesn't run the same action twice.
- Plan to stop and recover: decide who can cut the connection if something looks wrong, and how far you can roll it back.
A prompt that says "be careful" can't substitute for this. Enforce hard limits, like permissions and spending caps, in the tools and the server. Let the AI propose results inside those limits, not set them.
Checks you can run without customer data
Test with dummy documents and sending turned off, rather than running attack prompts against a live account. These are boundary checks, not an attempt to reproduce the attack technique itself.
- Does a document that references another team's files expand the agent's access scope?
- Does changing the test recipient after a draft is written trigger a recheck?
- Does sending the same execution request twice still produce only one result?
- If a required policy document is missing, does the agent hand off to a person instead of guessing?
- Can you tell who approved what from the execution log alone?
Where to start
Pick one workflow that already sends something external, such as an email, a payment, or a record update. Write its read / draft / execute boundaries using the table above, then run it through the five checks. Clear boundaries also mean fewer pop-ups for small, low-risk steps. Let approved reads and drafts flow without interruption, and save confirmation screens for changes the user is actually responsible for.
FAQ
Does writing precautions into the prompt count as permission control?
Instructions help, but they don't limit actual access. Constraints that must always hold, like target, amount, and send permission, need to be enforced in the tool and the server as well.
Sources
Want help picking your first AI pilot?
Book a free call