All English guides
AI agent pilotAX pilot checklistAI agent operationsAI agent permissionsAI agent ROI

The 2-week AX pilot checklist for rolling out an AI agent

AI agent rollouts fail without a task contract, staged permissions, human review, and logs. Run this 2-week pilot checklist before you scale up.

LeanX··8 min read

Also available in Korean: Read the Korean original

Short answer: Rolling out an AI agent isn't a headcount project. It's a project to design roles, permissions, review, and logs so you can safely hand off repetitive work. Pick one task and produce small outputs, again and again, for two weeks. Build an operating loop a person can actually check. That is more likely to succeed than building an autonomous agent on day one.

Most companies already have plenty of AI tools. Subscribing to one doesn't change how work gets done. Employees open a chatbot, ask the same question, and paste the answer into a document or chat. If someone still checks it from scratch at the end, AI just became one more task.

The starting point for AX (AI transformation) isn't "which model is smartest." It's deciding which task gets finished, with whose help, to what standard, and with what evidence left behind. This guide walks through an operating design a team can validate in two weeks when it first rolls out an AI agent.

An "agent" here isn't a robot that makes every decision on its own. Think of it as a software worker. It uses defined tools, finds the information it needs, produces a result, and hands it to the next person or system. What matters isn't how much autonomy it has. It's whether it knows where to stop, and stops there.

How is an AI agent different from a chatbot?

Mixing the two up blurs your expectations and your evaluation criteria. A chatbot is an interface that answers questions. An agent carries out multiple steps toward a goal. But doing multiple steps doesn't automatically make it a safe production system.

TypeMain roleKey question before adopting
ChatbotAnswers questions and explains informationHow do you check accuracy and freshness?
Automated workflowMoves data and sends alerts in a set orderWhere does it stop when there's an exception?
AI agentPicks tools and runs multiple steps to produce a resultDoes it stay inside its allowed tools and permissions?
AX operating systemRuns work, people, AI, systems, and measurement as one loopWho approves, how do you undo failure, what gets improved?

Teams don't always need the most autonomous agent. A task with a clear, repeated sequence is more stable as a fixed workflow. A task where the source material or the root cause changes every time is where a narrow, limited agent can help.

Five reasons agent pilots fail

  1. The task is too big. "Automate sales" isn't something you can execute. "Classify web inquiries within 5 minutes and draft a reply" has a start and an end.
  2. There's no definition of done. Replace "seems well written" with required fields, source links, banned phrases, and confirmation that it reached the right person.
  3. Permissions are too broad. Giving write, delete, or send access to a task that only needed read access turns a small error into a big one.
  4. Human review comes too late. Checking everything at the very end lets bad data spread through multiple systems first. Put an approval point in before risk builds up.
  5. Success is measured by automation rate alone. Handling 80% automatically isn't a win if a person needs longer to fix the other 20%.

None of these five go away by swapping the model. You have to break the task down again, define the input/output contract, and define what happens on failure.

How do you pick the first pilot task?

A good first task isn't an impressive one. It's frequent, has reasonably fixed inputs, and produces a result a person can review quickly. If a task meets four or more of the following, it's a pilot candidate:

  • You've handled 20 or more in the same format in the last month.
  • The input data's location and format are roughly fixed.
  • You can turn the required output fields into a checklist.
  • A person can catch a mistake before it's sent, deleted, or finalized.
  • You can start measuring processing time and rework time now.
  • If it fails, you can fall back to the manual process.

Avoid handing a first pilot decisions that are hard to reverse: performance reviews, finalizing refunds, making legal commitments to a customer, deleting operational data. Being technically possible is different from being something your organization can run safely.

How do you write a task contract for an agent?

A short task contract reproduces better than a long prompt. It also tells the agent what not to do. Six lines are usually enough to start:

Goal: Classify the last 24 hours of web inquiries by type and draft a reply.
Input: Inquiry text, customer tier, recent contact history, current policy docs.
Allowed tools: CRM lookup, policy document search, draft document creation.
Forbidden: Confirming refunds, promising discounts, downloading customer data, auto-send.
Done when: Category, source link, draft, confidence level, and assigned reviewer all exist.
On failure: If there's no source or confidence is low, route to human review and log why.

Keep this contract in your team docs and repo rules, not just in the prompt. That way the standard survives a change of person or tool.

Raise permissions step by step, starting from read-only

Agent permissions are a design question, not a trust question. Instead of connecting every system at once, raise access in stages:

StageAllowed scopePass criteria
0. ObserveRead sample data; show results on screen onlyYou find wrong inferences and missing patterns.
1. DraftDraft documents or tickets; no outbound sendingHuman edit time and missing-source rate are within range.
2. Internal handoffRoute to a reviewer's queue and internal alertsRouting errors and duplicate handoffs are under control.
3. Limited executionOnly reversible internal updatesAudit log and rollback steps are confirmed.
4. Conditional external actionSend or change only when conditions passHuman approval and exception handling are running in practice.

The bar for raising a stage isn't "the results look good." It's that failure stays contained, and you can reconstruct who did what, and when.

How to run a 2-week AX pilot

Days 1–2: fix the task and the baseline

Anonymize 20–30 representative recent cases as your baseline. Record time per case, time a person spent fixing it, how often something was missed or reworked, and any external tool cost. Skip this step and all you'll have at the end is "it feels faster."

Days 3–4: test in draft mode

Let the agent read, classify, and draft — nothing further. Every result carries its source and a confidence level. Pull the wrong cases from your baseline and use them to fix the task contract and your test cases.

Days 5–7: connect to an internal queue

Send drafts to a reviewer's internal queue, and log approvals, edits, and rejection reasons in a structured way. From here, how fast the person can judge the result matters more than the model's raw output.

Days 8–10: extend to a limited slice of real work

Connect one low-risk category to real work. Prefer recommend-and-confirm over auto-send. Deliberately feed in exceptions, duplicates, blank inputs, and permission errors to test the failure path.

Days 11–14: decide whether to continue

Compare processing time, edit time, approval rate, rework rate, and cost per case against your baseline. If the numbers look good but customer impact or a security issue showed up, stop. If you continue, reinforce exception handling for the same task before moving to the next one.

What should you measure?

MetricQuestionHow to read it
Time to first resultHow fast does a draft or category come back?Look at the change in wait time, not raw speed.
Human edit timeHow many minutes does a person need to finish the result?The real savings, including review cost.
Approval rateWhat share of first results pass the rules and get approved?Checks both the task contract and input quality.
Escalation rateWhat share gets handed to a person, and why?Tells you if it's a risk signal or a working safeguard.
Total cost per caseWhat do model, API, tool, and review costs add up to?Net benefit, not automation rate.
Recovery timeHow long to get back to the normal flow after an error?Turns operational risk into a cost figure.

Don't set a single goal for the pilot. Pair "cut processing time 30%" with safety goals too — "zero external send errors," "unsourced drafts under 5%."

Where should people stay in the loop?

If a person has to re-read every result from scratch, it isn't automation. But removing people entirely isn't AX either. People should focus on:

  • Decisions that affect customers, revenue, or legal liability
  • Exceptions with no supporting source in policy documents
  • Work touching personal data, permissions, or a security event
  • A new case type that falls outside existing rules
  • Whether to fold an automation result into the next rule update

Good AX doesn't have people generate the AI's output for it. It has AI handle repetitive drafting and organizing, and people spend their time on exceptions and improving the standard.

Security and privacy checklist

  • Have you listed what data the agent can and can't read?
  • Are you using a least-privilege service account instead of an operator's own login?
  • Do you mask or minimize personal data before it reaches the model?
  • Do you log tool calls, inputs, outputs, approver, and timestamp?
  • Is there a stop condition so failures don't retry forever?
  • Does a person approve any external send, delete, payment, or permission change?
  • Can you switch back to manual processing immediately if something breaks?

Where to start: the operating contract comes before the agent

An AI agent isn't an employee that automatically does the job well. It's a software worker that produces a result within a defined goal and toolset, then hands judgment to a person. So the real work of adoption isn't picking a model. It's tying scope, input, permissions, done-criteria, approval points, logs, and metrics into one contract.

Don't promise a big automation from day one. Pick one repeated task this week and write the six-line contract above for it. Validate it in draft mode for two weeks, and log every case a person had to fix and why. See where people actually spend their time, then widen permissions gradually. That process, not the agent itself, is what makes AX stick inside an organization.

FAQ

What task should you start an AI agent rollout with?

Pick something frequent, with fairly clear inputs and a definition of done, that a person can review quickly. Classifying web inquiries, summarizing internal research, and drafting reports are typical examples.

Is it safe to send an AI agent's output straight to a customer?

Not at first. Validate it in draft mode and internal-handoff mode before that. Keep a human approval step for anything hard to reverse, such as external sends, refunds, deletions, or permission changes.

How long does an AI agent pilot take?

If you scope it to one task, two weeks is enough to validate a baseline, drafts, approvals, failure handling, and cost and quality metrics. That's separate from how long a full organization-wide rollout takes.

Sources

  1. Anthropic: Building effective agents
  2. Claude Code documentation
  3. OpenAI Codex documentation
  4. GitHub Actions documentation
  5. NIST AI Risk Management Framework

Want help picking your first AI pilot?

Book a free call