PRACTICAL GUIDE
How to design approval workflows for AI Employees without slowing automation
An effective approval workflow protects important decisions without turning every automated task into another manual queue. This guide explains what to review, what context to show and how to evolve autonomy.
· IA Empleado
Human oversight fails when it is added at the end as a generic button. An approver receives a proposal without context, opens several applications to verify it and eventually accepts almost everything out of fatigue. The problem is not approval itself; it is the design. A good workflow separates preparation, authorisation and execution, triggers review only through clear rules, presents enough evidence, expires when business state changes and leaves traceability for learning. The goal is not to maximise or minimise approvals, but to assign people and automation the work each can perform more reliably.
01
1. List the real actions in the process
Start with concrete verbs: read customer, classify email, propose reply, create task, change status, issue refund, modify bank details or send a document. A broad label such as manage customer hides risk differences that are essential for control design.
Include actions that happen outside the main interface. An agent may send email, call an API, write to CRM or generate a file for another system. The inventory must cover real effects, not only steps visible to the user.
02
2. Build an impact and reversibility matrix
For every action assess financial, contractual, reputational, privacy, security and customer-experience impact. Then add reversibility: correcting an internal label is usually easy, while transferring money or sending confidential information to the wrong recipient may be difficult or impossible to undo.
You do not need a universal formula. A simple low, medium and high scale can work when criteria are defined. What matters is that two teams classify similar cases consistently and that the matrix can be revised as processes or risks change.
03
3. Assign an operating mode to each level
Low-risk actions can execute automatically with logging and limits. Medium-risk actions can start in proposal mode and require approval. High-impact actions may need dual validation, segregation of duties or remain outside the agent's scope.
Avoid using model confidence as the only boundary. A high-confidence output can still be a bad operation when it affects critical data. Confidence is an additional signal; action type and business rules remain decisive.
04
4. Define deterministic approval triggers
Turn the matrix into executable rules: amount above a threshold, discount outside range, external recipient, sensitive document category, irreversible change, blocked customer, contractual exception or contradictory sources. Every rule should be testable with concrete cases.
Keep authority rules outside the prompt. The model can explain why it believes a situation meets a condition, but a policy layer should check values, permissions and state before deciding whether the action requires review.
05
5. Design the approval request as a product
A good approval request summarises the case in seconds: action, rationale, impact, source data, proposed changes, triggered rules and links to the source system. Show essentials first and allow drill-down when the approver needs more evidence.
Avoid dumping entire conversations or documents by default. Extract the fragments supporting the proposal and respect data minimisation. Less noise reduces decision time and lowers the risk of exposing information that was not needed for approval.
06
6. Expose uncertainty and missing data
The approver needs to know which parts are verified and which are inferred. Mark unconfirmed fields, unavailable sources, multiple matches and contradictions. An apparently confident proposal that hides these signals creates artificial trust.
When a required value is missing, consider blocking approval until it is completed. A person should not become a mechanism for bypassing basic validation. Human review adds judgment, while deterministic rules continue protecting process integrity.
07
7. Bind approval to specific parameters
An authorisation should describe exactly what may execute: tool, entity, fields, values, amount and version of relevant state. Avoid generic authorisations such as do it or proceed that can be reinterpreted when context changes.
Generate an approval identifier and bind it to the operation. If important parameters change afterwards, invalidate the authorisation and request a new one. This prevents a legitimate human decision from backing an action different from the one reviewed.
08
8. Revalidate immediately before execution
Price, stock, balance, customer status, permissions or availability may change between proposal and approval. Before writing, query again the data affecting operation validity. If a relevant difference exists, stop the flow.
Revalidation also protects against races between automations. Two processes may try to modify the same record. Use versions, expected states or concurrency controls where supported so later changes are not overwritten.
09
9. Add expiry to decisions
Define a validity window according to the process. A commercial reply may remain valid for hours while an inventory-dependent operation may be valid only for minutes. Expiry should be visible to the approver and checked automatically before execution.
When it expires, do not reuse the previous yes. Rebuild the proposal with current data and show what changed. This reduces frustration and helps the person understand why an apparently similar action needs review again.
10
10. Design identity, permissions and segregation of duties
Record who approves and which role they held at that time. Some decisions may require a department owner, finance or security. Membership in a chat or shared inbox should not automatically equal authority for every operation.
For higher-impact actions apply separation between the person preparing and the person authorising, or even dual approval where internal policy justifies it. The objective is not universal bureaucracy, but explicit authority for each type of effect.
11
11. Define what happens on silence or rejection
A workflow needs clear states: pending, approved, rejected, expired, cancelled and escalated. If nobody responds, apply the SLA and predefined safe fallback. Never turn lack of response into consent for a sensitive action.
A rejection should close the operation or return it to preparation with instructions. Avoid endless loops where the agent presents the same proposal. Capture a structured reason and use it to correct data, rules or approach before requesting review again.
12
12. Design escalation without bypassing controls
Escalation means finding another authorised approver, increasing priority or routing to an operational channel. It does not mean elevating agent privileges because the case is urgent. Urgency changes timing treatment, not action authority.
Define alternate owners and coverage schedules for critical processes. If nobody authorised is available, the safe option may be to pause, inform the customer or execute a reversible alternative. Document these fallbacks before they are needed.
13
13. Audit the outcome, not just the approval click
Traceability should continue after yes. Record whether execution succeeded, what the destination system returned, whether retries occurred and whether the effect matched the proposal. An approval does not prove the operation completed correctly.
Use an end-to-end correlation identifier joining proposal, approval, execution and outcome. This makes errors investigable without relying on screenshots or manual reconstruction across email, CRM and technical logs.
14
14. Use sampling to reduce approvals without losing control
When an action proves stable, replace some pre-review with post-control through samples. Select random cases and add targeted samples for exceptions, model changes, new connector versions or higher-risk segments.
Define which outcome forces a temporary return to individual approval. Sampling works only when there is a response mechanism. Without rollback thresholds, reviewing samples becomes passive observation with no operational effect.
15
15. Measure oversight quality and cost
Track approval volume, average time, rejection rate, corrections, escalations, expiries and human minutes consumed. Add outcome metrics such as prevented errors, later incidents and team satisfaction. Without that combination you can optimise speed at the expense of safety.
Segment by action and risk. A global average can hide that one tool performs very well while another creates friction. Metrics should support decisions about adjusting thresholds, improving context or keeping mandatory review.
16
16. Evolve autonomy with entry and exit criteria
Before moving an action from proposal to automatic execution, define the observation period, minimum volume, required quality, absence of critical errors and rollback coverage. Promotion should be deliberate and versioned with the policy.
Also define rollback criteria: rising errors, data changes, a new integration, internal policy changes or incidents. Being able to reduce autonomy quickly is as important as expanding it. A mature system does not assume yesterday's decision remains valid forever.
TAKEAWAYS
Key ideas
Useful approval is designed by action and risk, not as a generic step for everything.
Authority rules should be deterministic, testable and external to the prompt.
The approver needs context, evidence, uncertainty and expected effect in one view.
Authorisation should bind to specific parameters, expire and be revalidated before execution.
Rejections, corrections and sampling are signals for improving the system and evolving autonomy.
The ability to reduce autonomy quickly is an essential part of operational control.
GO DEEPER
Human-in-the-loop AI automation: scale without giving up critical decisions.
An AI Employee does not have to choose between being useful and being controlled. Good design automates reading, classification, preparation and low-risk actions while reserving financial, legal, reputational or hard-to-reverse decisions for a person. Well-designed human oversight is not a brake: it is an operating layer that allows autonomy to expand with evidence.
APPLY IT