All posts

Research / Operations

When Should an AI Assistant Read, Recommend, or Act?

Set AI assistant permissions by action and consequence. Compare reading, recommendations, approved execution, and bounded automation with practical examples.

By Waypoint ExponentialPublished Revised
A teal cube waits on a cream track before a terracotta arch and brass barrier, illustrating an AI assistant's boundary before taking action

Give an AI assistant permission for a specific operation on a specific set of records. It can read an order without changing it, recommend a correction without approving it, or execute an authorised update. Choose the boundary by the consequence of an error, the evidence available, and the team's ability to recover. A convincing answer doesn't establish permission to act.

Separate the action levels

A practitioner discussion about production agents asks which operations need approval and how to avoid requests waiting indefinitely. Those questions belong in the workflow design. The discussion identifies reader concerns; its anecdotes don't prove that a particular approval product or threshold works for your business.

Use the following action levels as a proposed planning framework. The final two levels distinguish a person approving each operation from an owner authorising a narrow class of operations in advance. Your existing financial, security, and business approval rules still apply.

A proposed action boundary for each operation
LevelAssistant's jobExample boundary
ReadIt retrieves authorised information.It finds the delivery status for the current customer's order.
RecommendIt prepares a proposal for a person to decide.It drafts a correction with the source record and reason.
Execute after approvalIt performs the operation a qualified person approved.It applies the exact reviewed update if its conditions still hold.
Act within an agreed policyIt performs eligible operations without individual review.It adds an internal routing tag within a tested, limited workflow.

Assign a level to each operation, rather than the whole assistant. A support assistant can read tickets and draft replies while a staff member approves sending. Permission to add an internal tag needn't include permission to close the ticket, issue a refund, or export the customer's history.

Reading also needs a boundary. Decide which records and fields the assistant can retrieve, who can receive the answer, and which service processes the data. Read access can disclose information even when the assistant doesn't change a database. Keep a customer's lookup within that customer's authorised records.

Classify the actual consequence

Start by asking what happens after a wrong operation. An internal draft can wait for correction. An incorrect routing tag can hide urgent work. A message can disclose information as soon as the recipient sees it. Reversing a database value won't necessarily reverse those downstream effects.

Check scale as well as severity. A minor change to a single record may fit a narrow policy; the same change across thousands of records needs a different review. Consider the people affected, external commitments, and time available to detect a problem before another process acts on it.

OWASP's Excessive Agency guidance identifies unnecessary tools, broad permissions, and excessive autonomy as sources of damaging actions. It recommends limiting capabilities and enforcing authorisation in downstream systems. For a reading task, expose the required lookup and use an appropriately scoped identity; don't attach unrelated write or delete operations.

Record whether the action changes a business commitment. Updating a delivery preference before dispatch differs from promising a delivery date after stock allocation. Both may look like a small CRM edit, but the second affects other teams and the customer. Have the process owner classify them separately.

Don't let the model's own confidence decide its authority. Confidence can inform an evaluation if you have tested its relationship to actual errors, but it doesn't establish a user's permissions or a transaction's eligibility. Define mandatory review cases through the approved process and enforce them independently.

Follow an order through different boundaries

Consider an illustrative distributor handling an emailed order for replacement parts. The assistant reads the message and retrieves the customer's current order records. The lookup allows that account's records and the fields needed for order entry. It doesn't include unrelated accounts or payment credentials.

The assistant extracts a draft order and flags a packaging ambiguity: the message asks for ten boxes, while the catalogue stores individual units. It recommends a conversion using the relevant catalogue entry. Staff inspect the original request, the conversion, and the proposed quantity before accepting the draft.

Keep creating the order separate from accepting the extraction. A staff member may agree that the message asks for ten boxes without approving stock allocation, price, or credit terms. The assistant should show which business decision the approval covers and leave the other decisions with their existing owners.

After the customer and catalogue checks, a qualified operator can approve creating the exact order. If the source record changes while the proposal waits, refresh the affected proposal and request a new decision. An approved draft doesn't authorise a different quantity or a newly inferred customer account.

A small internal classification step may fit bounded automation. For example, the owner may allow a routing tag on eligible drafts while reserving order creation and customer communication for review. Treat this as a proposed local policy, with an explicit exception route; it isn't a claim that tagging is always low risk.

For a subscription business, apply the same separation to refunds. The assistant can gather the transaction and relevant policy, then recommend a response. The authorised staff member makes the financial decision under the company's controls. An automated recommendation doesn't establish permission to move money.

Give reviewers a decision they can inspect

Show the proposed change beside the evidence the reviewer needs. For an order, show the customer account, source request, current values, proposed values, and affected business step. If the assistant changes a draft after rejection, make that change visible rather than presenting the revised proposal as an unchanged request.

Use a reviewer who has authority for the operation. A salesperson can review a customer message without having authority to approve a credit exception. A process owner can accept a workflow's design without holding every transaction approval. Name the appropriate reviewer and their backup for each decision.

OWASP's AI Agent Security Cheat Sheet recommends validating consequential actions outside the agent. Its approval controls bind the actor and exact operation to a time-limited record. If the target or parameters change, the executor needs a fresh approval. An agent-supplied confirmation flag doesn't prove that a person reviewed the action.

Set a response path for waiting proposals. Record who receives the request, when it expires, and where an unresolved case goes. Expiry should leave the case in an explicit state, such as awaiting manual handling. Silence from a reviewer shouldn't become approval.

Protect review capacity. If staff skim proposals because the queue is too large, reduce the pilot's volume or simplify the decision presented. Ask reviewers to explain a disputed case using the source evidence. That exercise can reveal whether the screen supports a real judgement or encourages a quick click.

Enforce the approved action at execution

Use a small operation with a clear contract. An order tool might accept the approved draft reference and return the resulting order reference. Define which fields it can change and the conditions the application checks. A general instruction to update customer records leaves the actual boundary unclear.

Anthropic's engineering guide to agent tools recommends tools with a distinct purpose, explicit parameters, and evaluations grounded in real work. That supports testing a narrow order operation against realistic cases before offering a broad collection of system commands. The guidance doesn't prove that a particular tool is safe for your workflow.

Give the operator a visible distinction between a proposal, an accepted decision, and a completed operation. For this illustrative order flow, the completion record should identify the order the ERP actually created. If creation fails, keep the approved proposal available for controlled resolution without showing a successful order to staff.

Define what happens when the rule service, identity check, or required evidence is unavailable. For the proposed workflow, suspend the affected write and route the case to an operator. Don't let an integration failure silently expand the assistant's permissions. Keep an agreed manual route available during the pilot.

Handle retries and uncertain outcomes

A timeout can leave the caller unsure whether an operation happened. For example, the ERP may create an order before the client loses its connection. Treat the case as unresolved until you can establish the result; asking the model to generate another order request can create a duplicate.

Stripe's idempotency documentation gives a concrete provider example: a repeated request with the same key returns the saved result, and the service checks that the parameters match. Stripe may remove keys after at least 24 hours. These are Stripe's semantics, so inspect your own destination's guarantees and retention period rather than assuming every API behaves the same way.

For the order example, store a stable operation reference before the first attempt and reconcile it with the destination's record. Use the destination's duplicate-prevention mechanism where it supports the operation. Keep a changed order proposal separate from retrying the original request, and preserve the receiving system's record of completion.

Test recovery beyond restoring an old value. If a routing update already triggered warehouse work, the owner needs a way to find that work and coordinate correction. Record affected cases and downstream contacts. A technical rollback and business recovery can require different people.

Earn broader permission with operating evidence

Start with the smallest permission that meets the use case. Run proposals alongside ordinary work and compare them with staff decisions before allowing an eligible class of writes. Choose representative cases and difficult exceptions; a demonstration on tidy records won't establish the proposed boundary.

For this order pilot, test a missing account, ambiguous packaging, changed stock status, duplicate submission, and reviewer absence. Inspect the actual record and the action history for each case. Include an input that tries to instruct the assistant to bypass the approved process; the case should preserve the established action boundary.

Measure wrong actions, missed exceptions, review time, unresolved outcomes, and recovery effort. Define acceptance criteria before reviewing the results. Keep severe failures visible separately from routine corrections so a high average success rate doesn't conceal an unacceptable failure.

Expand permission through a named release decision. The owner should specify the newly eligible cases and the evidence supporting them. Recheck the boundary when the workflow gains another customer, destination, or business action. Our pilot evidence guide covers the broader decision to scale.

Make the boundary part of the operating agreement

Keep an action register alongside the workflow documentation. For each operation, state its permitted records, level, eligibility conditions, approver or policy owner, recovery route, and review date. The person covering an absent colleague should be able to identify what they may approve without guessing.

An SME can start with a short register for its first workflow. A corporate platform team can provide the execution controls while the business unit owns eligibility and operating decisions. A PE operating team can share the method across portfolio companies while each company approves its own business boundaries.

For venture diligence, ask to see an actual proposal, its approval or policy decision, and the resulting destination record. Ask what happens during an uncertain outcome and who can suspend writes. Those examples give evidence about the product's operating design and the manual work it still needs.

Use the deployment-team responsibility map to assign the work and the post-launch ownership guide to maintain it. Before expanding authority, have the receiving operator demonstrate an exception and a recovery case. Approve the boundary your team can actually operate.

Put the work into practice

Operations consulting and business process automation

We help company leaders improve the work between people, data, and systems. Our consulting and engineering team maps the process, fixes the handoffs, and builds the automation your staff will use.