← Blog

Article | LUMATARRA

A practical 30-day AI agent pilot plan

Use this 30-day AI agent pilot plan to define scope, run shadow mode, keep human approval, measure one result, and make a clear go/no-go decision.

A controlled AI agent pilot moves one business workflow through scope, shadow mode, human-approved execution, and an evidence-based decision.

A useful AI agent pilot does not begin with more features. It begins with a stop condition and a person who owns the final decision.

That may sound cautious, but it is what makes a 30-day pilot move quickly. The team knows which workflow is in scope, which sources the agent may use, what it may prepare, what it must never do without approval, how exceptions move to a person, and what evidence will support a continue, change, or stop decision.

This AI agent pilot plan is designed for one workflow that has already passed a basic readiness check. If the trigger, approved sources, owners, or outcome are still unclear, use the 10-point AI agent readiness checklist first. A pilot should test a defined operating job, not discover what the job was supposed to be.

What a 30-day AI agent pilot must prove

An AI agent proof of concept can show that a model produces an interesting answer. A business pilot must prove something more useful: that one controlled workflow can improve one operating result without creating unacceptable risk or shifting hidden work to the people reviewing it.

By day 30, the team should be able to answer five questions with evidence:

  1. Did the agent use only the approved sources and follow the defined scope?
  2. Did its proposed or completed routine work meet the agreed quality bar?
  3. Did human approval and exception handling work under real operating conditions?
  4. Did the target measure improve without creating a larger review or cleanup burden?
  5. Should the business continue, change, or stop this version of the workflow?

The pilot is not successful because the agent ran. It is successful when the business can make a better decision about the next version.

Complete the pilot charter before day one

Write the charter in plain operating language. A leader, process owner, reviewer, and builder should all be able to read it and describe the same pilot.

AI agent pilot charter worksheet

  1. Workflow or job: Name one repeated business job, not a department or broad transformation goal.
  2. Trigger: State the observable event or schedule that starts the work.
  3. Approved sources: List the systems, documents, datasets, or locations the agent may use and identify the source of record.
  4. Routine output: Define the draft, classification, summary, routing decision, review packet, or other bounded output the agent should prepare.
  5. Prohibited actions: State what the agent must not decide, send, approve, change, or assume.
  6. Human approval point: Name exactly what a person reviews before any consequential action occurs.
  7. Exception owner and path: Name who receives missing, disputed, sensitive, or out-of-scope cases and how the handoff occurs.
  8. Accountable process owner: Name the person responsible for the workflow and its business result.
  9. Baseline metric: Record the current value of one operating measure before the agent enters the workflow.
  10. Pilot target: Define the improvement or learning threshold that would support continuing.
  11. Stop or rollback condition: Define the quality, safety, data, or workload boundary that pauses the pilot.
  12. Go/no-go decision owner: Name the person who will decide continue, change, or stop at the end of the pilot.

Do not leave ownership in a group label such as “the team.” One person may hold several roles in an owner-led company, but each responsibility must still be explicit.

Choose one baseline and one target

A pilot becomes difficult to evaluate when every possible benefit is treated as a primary goal. Choose one operating measure that reflects the job.

Useful measures include:

  • minutes required to prepare a recurring operating brief;
  • elapsed time from intake to correct routing;
  • percentage of cases arriving with required information;
  • review time per document or case;
  • number of follow-ups left without an owner;
  • percentage of agent outputs accepted without material correction; or
  • number and type of exceptions identified before a handoff.

Record the baseline using the current process before the pilot changes it. Define the target in the same unit. If the baseline is 45 minutes of preparation per weekly brief, do not switch to “number of outputs generated” during the pilot. Activity is not the result.

Also measure approval workload as a guardrail. A faster draft is not an improvement if reviewers spend more time correcting it than the original work required.

Set the boundary before connecting tools

The approved source list is both a data boundary and an operating instruction. For each input, state whether it is the source of record, supporting context, or excluded material.

The pilot should not fill a missing fact with an unsupported assumption. It should flag the gap, preserve the available evidence, and route the case to the exception owner.

Without explicit human approval, the pilot must not make:

  • customer commitments or promises;
  • financial transactions, approvals, or changes;
  • legal, policy, privacy, or security decisions;
  • production-system mutations;
  • sensitive external communications; or
  • conclusions that depend on missing or disputed information.

These are practical pilot controls, not legal advice. The point is to keep authority with the people who already own the decision while the agent handles bounded preparation and coordination work.

LUMATARRA’s approach to AI strategy and readiness starts with this operating design: define the job, sources, controls, owners, and measurable result before choosing how much technology to connect.

The 30-day AI agent pilot plan

The sequence below assumes the workflow happens often enough to produce representative cases during the month. If it runs only once per quarter, extend the observation window rather than pretending a small sample proves reliability.

Week 1: Narrow the job and establish control

Goal: Freeze the pilot charter, capture the baseline, and make the boundaries testable.

  • Observe the current workflow and write its trigger, standard path, handoffs, and definition of done.
  • Confirm approved sources and identify the source of record for each required fact.
  • Capture the baseline metric using the current human-run process.
  • Name the process owner, approval owner, exception owner, and go/no-go decision owner.
  • Map risks and approval points, including the actions the agent is prohibited from taking.
  • Collect representative cases, including ordinary work, missing information, disputed facts, and known exceptions.

Exit check: Every item in the pilot charter has an answer, the baseline is recorded, and unresolved scope or source questions have an owner.

A week-one build that races ahead of this work usually creates rework. The fastest path is to make the operating contract clear enough that the agent can be tested against it.

Week 2: Run in shadow mode

Goal: Compare agent proposals with normal work without allowing autonomous commitments or production changes.

  • Let the agent read only the approved inputs and prepare the defined routine output.
  • Keep the current human process as the operating source of truth.
  • Compare the proposed output with the human result using agreed review criteria.
  • Record corrections, unsupported assumptions, missing citations, routing errors, and exceptions.
  • Measure preparation and review time separately so hidden workload remains visible.
  • Refine instructions or source handling only when a change remains inside the approved charter.

Exit check: The team can describe the common failure modes, the agent reliably stops at prohibited actions, and reviewers agree that limited human-in-the-loop execution is safe to test.

Shadow mode is not a demo. It is a controlled comparison that shows whether the agent can follow the job before it participates in the job.

Week 3: Execute with a human in the loop

Goal: Use the agent in the live workflow while a named person approves consequential outputs and exceptions follow the defined path.

  • Allow only the routine action or handoff authorized by the charter.
  • Require the approval owner to review the defined evidence before any consequential action.
  • Route missing, sensitive, disputed, or out-of-scope cases to the named exception owner.
  • Log the source evidence, proposed output, reviewer decision, material edits, and final disposition.
  • Read back the target operating measure and the approval workload at an agreed cadence.
  • Pause immediately if a stop condition is crossed; do not tune around a safety boundary while the pilot continues.

Exit check: The approval path works in practice, exceptions reach a person, and the team has enough live evidence to compare the pilot with the baseline.

“Human in the loop” must name a decision, not merely a person who could intervene. The reviewer needs to know what evidence to inspect, what they may change, and what the workflow does after approval or rejection.

Week 4: Compare evidence and decide

Goal: Evaluate the operating result, failure modes, and workload shift, then make an explicit decision.

  • Compare the pilot result with the original baseline and target using the same unit.
  • Review recurring errors, exceptions, unsupported assumptions, and near misses.
  • Measure how much work moved to reviewers, exception owners, or downstream cleanup.
  • Confirm that prohibited actions remained blocked and approval evidence is complete.
  • Decide whether the next version should continue unchanged, change scope or controls, or stop.
  • Document the decision, owner, required changes, and the next measurement period.

Exit check: The go/no-go owner records a continue, change, or stop decision and the evidence supporting it.

Use explicit stop and rollback conditions

A stop condition protects the business from normal pilot pressure. When a deadline is close, teams are tempted to explain away a bad case and keep moving. A predefined boundary turns that moment into a decision rather than a debate.

Examples include:

  • the agent uses or exposes an unapproved source;
  • a prohibited action is attempted;
  • a material unsupported assumption reaches the approval step;
  • exceptions are not reaching the named owner;
  • the error rate exceeds the chartered threshold;
  • approval or correction time exceeds the original workload; or
  • the workflow cannot return cleanly to the prior operating path.

The rollback should be simple: pause the agent, return to the known human-run process, preserve the evidence, and let the process owner decide what must change before testing resumes.

Make the day-30 decision in operating language

The final review should not end with “the AI looked promising.” Choose one of three decisions.

Continue

Continue when the target result improved, controls worked, exceptions were manageable, and review effort remained proportionate. The next step may be a longer measurement period or a carefully bounded increase in volume. It is not automatic permission to remove approval.

Change

Change when the job is valuable but the evidence shows a specific design problem: the source boundary is too broad, the routine output is poorly defined, an exception category is common, or review workload is too high. Update the charter, preserve the baseline, and test the revised version against a new decision date.

Stop

Stop when the workflow is not stable enough, the required data is not trustworthy, the controls do not hold, the value does not justify the burden, or the pilot crosses a stop condition. Stopping is a useful result when it prevents a weak proof of concept from becoming unmanaged production work.

From a pilot to a practical AI operating system

A strong pilot creates an operating pattern the business can reuse: one job, approved context, explicit controls, named owners, measurable evidence, and a decision loop. That is more durable than a collection of disconnected AI tools.

For examples of the role-based work this pattern can support, review AI agents for business operations. The platform comes after the operating contract. Whether the implementation uses Microsoft 365, Teams, Copilot Studio, Power Platform, Fabric, Power BI, Azure, or another approved component, the pilot still needs the same boundaries and evidence.

The best first pilot does not try to prove that AI can do everything. It proves that one workflow can become faster, clearer, or more reliable while people keep authority over consequential decisions.

Book an AI Operating Review

FAQ

How long should an AI agent pilot run?

Thirty days is often enough to observe one repeatable workflow, compare a baseline with a controlled result, capture exceptions, and decide whether to continue, change, or stop. The right duration still depends on how often the workflow occurs and whether the team can collect enough representative cases.

What should an AI agent pilot measure?

Measure one operating result tied to the workflow, such as preparation time, routing speed, review time, missing-information rate, or unassigned follow-up. Record the baseline before the pilot and compare the same measure during controlled execution.

What is shadow mode in an AI agent pilot?

In shadow mode, the agent observes inputs and prepares a proposed output without making commitments or changing production records. A human compares the proposal with the normal process, records errors and exceptions, and decides when the pilot is ready for limited execution.

Where should human approval sit in an AI agent pilot?

Human approval should sit immediately before any action that creates a customer commitment, financial action, legal or security decision, production change, sensitive communication, or other consequence the agent is not authorized to make.

When should a company stop an AI agent pilot?

Stop or roll back when the pilot crosses a predefined safety, quality, workload, or data boundary; when exceptions cannot be handled reliably; or when the measured result does not justify the added review burden. The stop condition and decision owner should be named before the pilot starts.