PRACTICAL GUIDE
How to run an AI Employee pilot and decide whether it deserves to scale
A practical guide to turning an AI automation idea into a measurable, controlled pilot that supports a business decision.
· IA Empleado
AI pilots often fail for two opposite reasons: they remain demos that are too small to prove real value, or they try to cover so many processes that it becomes impossible to know what works. An AI Employee pilot should sit between these extremes. It needs to operate on a real process, with enough data and integrations to be representative, while keeping clear boundaries so risk can be controlled and outcomes can be attributed. This guide proposes a practical sequence for selecting the case, measuring the starting point, designing the intervention, operating with oversight and deciding whether the system should scale.
01
1. Choose a process with enough volume
You need enough cases to observe patterns, errors and exceptions. If a process occurs twice a month, the pilot may take too long to generate evidence. Prioritise frequent activities where the team can provide a representative sample within a reasonable period.
Volume does not necessarily mean thousands of cases. What matters is enough repetition to compare behaviour and distinguish coincidence from consistent improvement. Define a minimum number of observed cases before closing evaluation.
02
2. Check that the process is stable enough
A process undergoing constant redesign is a poor candidate. If rules, owners or systems keep changing, any result will be mixed with those changes. A pilot works best when there is a relatively stable way of doing the work today.
This does not mean the process must be perfect. It may contain manual tasks, friction and exceptions, precisely because you want to improve it. What matters is understanding the current workflow and being able to describe what changes during the test.
03
3. Write a concrete value hypothesis
State what you expect to improve and why. For example: if the AI Employee classifies requests and drafts responses using CRM context, the team can reduce first-response time without increasing corrections. The hypothesis connects technology to a measurable consequence.
Also state what you are not trying to prove. An email pilot may evaluate classification and drafting, but not necessarily sales, annual satisfaction or global staffing savings. Limiting the hypothesis keeps evaluation honest.
04
4. Measure the current process before intervening
Across a reasonable sample, record volume, time, errors, rework, escalations, manual steps and approximate cost. Where possible, measure by case type because simple and complex cases often behave very differently.
Document how the baseline was obtained and its limitations. You do not need accounting-level precision to make a good decision, but you do need a consistent enough reference to compare the pilot with previous work.
05
5. Describe the exact scope
Specify inputs, outputs, tools, users and included case types. Also list exclusions: unsupported languages, amounts above a threshold, special customers, sensitive documents or systems not yet connected.
A written scope prevents success from depending on different expectations across teams. It also lets you analyse an error in context: a case outside the pilot should not count as a system failure, but it may become a candidate for a later phase.
06
6. Select the minimum necessary data
Identify what information the AI Employee actually needs to complete the process. Avoid connecting complete databases for convenience. The smaller the data set, the easier it becomes to control permissions, quality, privacy and behaviour.
Before launch review duplicates, missing fields, inconsistent formats and conflicting sources. Data problems quickly become apparent AI problems. Correcting them or designing exception rules improves interpretation of results.
07
7. Design integrations with least privilege
Each connector should expose only the operations required by the pilot. If the agent needs to read orders and create tasks, it does not need broad permission to delete customers or modify accounting. Narrow scopes limit error impact and simplify auditing.
Separate read and write access where possible. During early iterations you can enable real queries while keeping writes behind human approval. This validates context and integration before granting more autonomy.
08
8. Decide what the person does and what the AI does
Map the division of work: the agent may read, summarise, classify, retrieve data, propose and execute certain actions; the person may approve, resolve exceptions or make higher-impact decisions. This boundary should be clear before the pilot.
Avoid using people as validators of every detail when it adds no value. If everything is reviewed line by line, the pilot measures an assistance tool rather than automation. Design review in proportion to risk.
09
9. Set business, quality and safety metrics
Do not evaluate speed alone. A workflow can become faster while producing more errors. Combine time, volume, resolution, rework, human intervention, cost, critical errors and policy compliance. The combination should reflect the outcome that matters to the business.
Define thresholds or target ranges before launch. For example, reduce time without exceeding a certain correction rate and without critical incidents. This prevents success from being declared based only on the metric that happened to improve most.
10
10. Define a pilot group and observation period
Start with a small but representative group of users or cases. Avoid selecting only simple examples because they create an overly optimistic picture. Include enough variability to encounter common exceptions without exposing the whole business.
Instead of setting only an end date, define a minimum volume as well. If one week has low activity, you may not gather enough evidence. Closure should happen when a useful sample has been reached and relevant conditions have been observed.
11
11. Instrument the pilot from day one
Assign case identifiers and record stages, tools, errors, times, approvals, corrections and outcomes. Without observability it becomes difficult to explain why a metric changes or investigate discrepant cases.
Keep versions of models, prompts, rules and integrations. When you adjust the system during the pilot, you can separate results by version and avoid mixing old and new behaviour into one average.
12
12. Design pause and rollback criteria
Specify which situations require stopping writes, returning to human approval, isolating an integration or reverting to the previous version. Examples include critical errors, sustained correction increases, inconsistent data or unexpected behaviour after a release.
Test these mechanisms before launch. Knowing that a safe mode theoretically exists is not enough if nobody knows how to activate it or if it takes too long. The ability to step back increases confidence for responsible experimentation.
13
13. Review failed cases with a common taxonomy
Classify each problem by likely source: data, model, instructions, integration, rule, permission, interface or process. Add severity and whether it was detected automatically or by a person. This taxonomy turns anecdotes into patterns.
Not every failure justifies changing the model. If most issues come from an incomplete data source or a slow API, improvement should target that component. The pilot helps discover what really limits value.
14
14. Calculate total operating cost
Include model usage, APIs, infrastructure, licences, storage and human review time. A pilot may look cheap in AI usage while becoming expensive if it generates many corrections or requires heavy oversight.
Compare cost per case and, where possible, cost per correctly resolved case. The second metric penalises rework and gives a more realistic view of efficiency than a technical consumption figure.
15
15. Run a final review using predefined criteria
Present the baseline, outcomes, case distribution, quality, cost, incidents, feedback and exceptions. Explain which improvements were made during the pilot and separate results by version where relevant.
Then compare each criterion with the agreed target. Avoid general conclusions such as it worked well. A useful decision identifies what is ready, what needs another iteration and what should not scale yet.
16
16. Scale with a controlled expansion plan
If the pilot meets its criteria, gradually expand volume, users, integrations or autonomy. Change one meaningful dimension at a time and observe the effect. This sequence makes regressions easier to attribute and keeps control.
Keep everything learned as an operating standard: permissions, observability, oversight, tests, rollback, metrics and owners. The goal of a successful pilot is not simply to finish the test, but to turn what worked into repeatable, sustainable operations.
TAKEAWAYS
Key ideas
Choose a frequent, stable and measurable process.
Define the hypothesis, scope and baseline before building.
Start with minimum data and integrations and limited autonomy.
Set success, pause and rollback criteria before launch.
Measure quality, business impact, cost and human intervention during execution.
Scale only when evidence meets the agreed criteria.
GO DEEPER
AI Employee pilot: validate value, safety and operational fit before scaling.
A good pilot does not try to automate the whole company. It selects one concrete process, defines a baseline, limits scope, connects only the systems that are necessary and establishes success criteria before starting. This makes it possible to prove savings, quality and control with real data, detect friction early and decide with evidence whether to expand, adjust or stop the initiative.
APPLY IT