The 2-week AX pilot checklist for rolling out an AI agent
AI agent rollouts fail without a task contract, staged permissions, human review, and logs. Run this 2-week pilot checklist before you scale up.
Also available in Korean: Read the Korean original
Most companies already have plenty of AI tools. Subscribing to one doesn't change how work gets done. Employees open a chatbot, ask the same question, and paste the answer into a document or chat. If someone still checks it from scratch at the end, AI just became one more task.
The starting point for AX (AI transformation) isn't "which model is smartest." It's deciding which task gets finished, with whose help, to what standard, and with what evidence left behind. This guide walks through an operating design a team can validate in two weeks when it first rolls out an AI agent.
An "agent" here isn't a robot that makes every decision on its own. Think of it as a software worker. It uses defined tools, finds the information it needs, produces a result, and hands it to the next person or system. What matters isn't how much autonomy it has. It's whether it knows where to stop, and stops there.
How is an AI agent different from a chatbot?
Mixing the two up blurs your expectations and your evaluation criteria. A chatbot is an interface that answers questions. An agent carries out multiple steps toward a goal. But doing multiple steps doesn't automatically make it a safe production system.
| Type | Main role | Key question before adopting |
|---|---|---|
| Chatbot | Answers questions and explains information | How do you check accuracy and freshness? |
| Automated workflow | Moves data and sends alerts in a set order | Where does it stop when there's an exception? |
| AI agent | Picks tools and runs multiple steps to produce a result | Does it stay inside its allowed tools and permissions? |
| AX operating system | Runs work, people, AI, systems, and measurement as one loop | Who approves, how do you undo failure, what gets improved? |
Teams don't always need the most autonomous agent. A task with a clear, repeated sequence is more stable as a fixed workflow. A task where the source material or the root cause changes every time is where a narrow, limited agent can help.
Five reasons agent pilots fail
- The task is too big. "Automate sales" isn't something you can execute. "Classify web inquiries within 5 minutes and draft a reply" has a start and an end.
- There's no definition of done. Replace "seems well written" with required fields, source links, banned phrases, and confirmation that it reached the right person.
- Permissions are too broad. Giving write, delete, or send access to a task that only needed read access turns a small error into a big one.
- Human review comes too late. Checking everything at the very end lets bad data spread through multiple systems first. Put an approval point in before risk builds up.
- Success is measured by automation rate alone. Handling 80% automatically isn't a win if a person needs longer to fix the other 20%.
None of these five go away by swapping the model. You have to break the task down again, define the input/output contract, and define what happens on failure.
How do you pick the first pilot task?
A good first task isn't an impressive one. It's frequent, has reasonably fixed inputs, and produces a result a person can review quickly. If a task meets four or more of the following, it's a pilot candidate:
- You've handled 20 or more in the same format in the last month.
- The input data's location and format are roughly fixed.
- You can turn the required output fields into a checklist.
- A person can catch a mistake before it's sent, deleted, or finalized.
- You can start measuring processing time and rework time now.
- If it fails, you can fall back to the manual process.
Avoid handing a first pilot decisions that are hard to reverse: performance reviews, finalizing refunds, making legal commitments to a customer, deleting operational data. Being technically possible is different from being something your organization can run safely.
How do you write a task contract for an agent?
A short task contract reproduces better than a long prompt. It also tells the agent what not to do. Six lines are usually enough to start:
Goal: Classify the last 24 hours of web inquiries by type and draft a reply.
Input: Inquiry text, customer tier, recent contact history, current policy docs.
Allowed tools: CRM lookup, policy document search, draft document creation.
Forbidden: Confirming refunds, promising discounts, downloading customer data, auto-send.
Done when: Category, source link, draft, confidence level, and assigned reviewer all exist.
On failure: If there's no source or confidence is low, route to human review and log why.
Keep this contract in your team docs and repo rules, not just in the prompt. That way the standard survives a change of person or tool.
Raise permissions step by step, starting from read-only
Agent permissions are a design question, not a trust question. Instead of connecting every system at once, raise access in stages:
| Stage | Allowed scope | Pass criteria |
|---|---|---|
| 0. Observe | Read sample data; show results on screen only | You find wrong inferences and missing patterns. |
| 1. Draft | Draft documents or tickets; no outbound sending | Human edit time and missing-source rate are within range. |
| 2. Internal handoff | Route to a reviewer's queue and internal alerts | Routing errors and duplicate handoffs are under control. |
| 3. Limited execution | Only reversible internal updates | Audit log and rollback steps are confirmed. |
| 4. Conditional external action | Send or change only when conditions pass | Human approval and exception handling are running in practice. |
The bar for raising a stage isn't "the results look good." It's that failure stays contained, and you can reconstruct who did what, and when.
How to run a 2-week AX pilot
Days 1–2: fix the task and the baseline
Anonymize 20–30 representative recent cases as your baseline. Record time per case, time a person spent fixing it, how often something was missed or reworked, and any external tool cost. Skip this step and all you'll have at the end is "it feels faster."
Days 3–4: test in draft mode
Let the agent read, classify, and draft — nothing further. Every result carries its source and a confidence level. Pull the wrong cases from your baseline and use them to fix the task contract and your test cases.
Days 5–7: connect to an internal queue
Send drafts to a reviewer's internal queue, and log approvals, edits, and rejection reasons in a structured way. From here, how fast the person can judge the result matters more than the model's raw output.
Days 8–10: extend to a limited slice of real work
Connect one low-risk category to real work. Prefer recommend-and-confirm over auto-send. Deliberately feed in exceptions, duplicates, blank inputs, and permission errors to test the failure path.
Days 11–14: decide whether to continue
Compare processing time, edit time, approval rate, rework rate, and cost per case against your baseline. If the numbers look good but customer impact or a security issue showed up, stop. If you continue, reinforce exception handling for the same task before moving to the next one.
What should you measure?
| Metric | Question | How to read it |
|---|---|---|
| Time to first result | How fast does a draft or category come back? | Look at the change in wait time, not raw speed. |
| Human edit time | How many minutes does a person need to finish the result? | The real savings, including review cost. |
| Approval rate | What share of first results pass the rules and get approved? | Checks both the task contract and input quality. |
| Escalation rate | What share gets handed to a person, and why? | Tells you if it's a risk signal or a working safeguard. |
| Total cost per case | What do model, API, tool, and review costs add up to? | Net benefit, not automation rate. |
| Recovery time | How long to get back to the normal flow after an error? | Turns operational risk into a cost figure. |
Don't set a single goal for the pilot. Pair "cut processing time 30%" with safety goals too — "zero external send errors," "unsourced drafts under 5%."
Where should people stay in the loop?
If a person has to re-read every result from scratch, it isn't automation. But removing people entirely isn't AX either. People should focus on:
- Decisions that affect customers, revenue, or legal liability
- Exceptions with no supporting source in policy documents
- Work touching personal data, permissions, or a security event
- A new case type that falls outside existing rules
- Whether to fold an automation result into the next rule update
Good AX doesn't have people generate the AI's output for it. It has AI handle repetitive drafting and organizing, and people spend their time on exceptions and improving the standard.
Security and privacy checklist
- Have you listed what data the agent can and can't read?
- Are you using a least-privilege service account instead of an operator's own login?
- Do you mask or minimize personal data before it reaches the model?
- Do you log tool calls, inputs, outputs, approver, and timestamp?
- Is there a stop condition so failures don't retry forever?
- Does a person approve any external send, delete, payment, or permission change?
- Can you switch back to manual processing immediately if something breaks?
Where to start: the operating contract comes before the agent
An AI agent isn't an employee that automatically does the job well. It's a software worker that produces a result within a defined goal and toolset, then hands judgment to a person. So the real work of adoption isn't picking a model. It's tying scope, input, permissions, done-criteria, approval points, logs, and metrics into one contract.
Don't promise a big automation from day one. Pick one repeated task this week and write the six-line contract above for it. Validate it in draft mode for two weeks, and log every case a person had to fix and why. See where people actually spend their time, then widen permissions gradually. That process, not the agent itself, is what makes AX stick inside an organization.
FAQ
What task should you start an AI agent rollout with?
Pick something frequent, with fairly clear inputs and a definition of done, that a person can review quickly. Classifying web inquiries, summarizing internal research, and drafting reports are typical examples.
Is it safe to send an AI agent's output straight to a customer?
Not at first. Validate it in draft mode and internal-handoff mode before that. Keep a human approval step for anything hard to reverse, such as external sends, refunds, deletions, or permission changes.
How long does an AI agent pilot take?
If you scope it to one task, two weeks is enough to validate a baseline, drafts, approvals, failure handling, and cost and quality metrics. That's separate from how long a full organization-wide rollout takes.
Sources
Want help picking your first AI pilot?
Book a free call