All posts

Research / Organisation design

Put AI Quality Review Where the Decision Is Made

Design AI quality review with accountable decision owners, staffed approval queues, clear escalation rules, and records that connect approval to actual action.

By Waypoint ExponentialPublished Revised
Teal cubes approach a brass-framed cubic checkpoint on a cream platform, with a terracotta cube in a side bay, representing review before an AI action

Put AI quality review beside the business decision, with the evidence and authority the reviewer needs to act. Name the person who owns the outcome, staff the approval queue, and connect each approval to the exact action the system takes. This gives an operations team a practical way to control work without turning every generated output into another manager's inbox task.

Define the decision and its accountable owner

“A human checks the AI” leaves several operating questions unanswered. Which person checks it? What are they checking against? Can they reject the recommendation, change the proposed action, and stop further work? A reviewer who can only click approve carries responsibility without enough control.

Start with a specific decision. An accounts payable assistant can extract invoice fields, propose a cost code, and prepare a payment instruction. Each step has a different consequence. The finance team may accept automatic extraction while requiring an authorised person to release a payment. Write down the boundary in terms of the business action.

Name the operating owner who sets the acceptance rule and answers for the process result. Name the reviewer who handles an individual case, the engineering owner who maintains the execution controls, and the assurance owner who checks whether the review process works. In a small company, a person may hold several roles; make the duties explicit and arrange a separate check where the stakes warrant it.

The NIST AI Risk Management Framework 1.0 calls for documented responsibilities, trained personnel, and defined roles for human oversight. Its Manage function also addresses responsibility for overriding or deactivating systems. NIST currently says it is updating the framework. These principles support the operating design here; the design isn't a certification checklist.

Keep the decision owner close to the domain. A procurement lead understands supplier substitutions and delivery commitments; a central AI team understands model and platform behaviour. Give both a route to intervene, while the procurement lead owns the business acceptance criteria. A model improvement doesn't by itself settle whether a substitution meets a customer's contract.

For a VC investor assessing a startup, ask who owns decisions after the demonstration team leaves a customer. For a private equity operating partner, check whether each portfolio company has a receiving owner and paid review capacity. A shared platform team can support the controls, but local management still needs authority over the affected operation.

Place review before the relevant commitment

Map where the workflow makes a commitment: sending a message, changing a record, placing an order, or granting access. Put the required approval immediately before that action, with enough time for a person to understand the case. Reviewing a summary after the system sends the order only supports detection and recovery.

Separate preparation from execution. Let the assistant gather evidence and prepare a proposed change inside its allowed scope. Use the application and its permissions to enforce whether it can execute. A prompt telling the model to ask permission doesn't enforce a payment limit or prevent a tool from changing a customer record.

Choose controls according to the consequences, reversibility, and evidence you have about the task. A suggested internal tag may tolerate correction after the event. A supplier bank-detail change needs the existing verification process before any subsequent payment. A reviewer can't compensate for missing identity checks by reading an AI explanation carefully.

Examples of where to put review in an AI workflow
Proposed actionDecision ownerReview point
Apply an internal tag.The queue lead owns the tagging rule.Check a defined sample and exceptions.
Send a customer quote.The commercial lead owns terms.Review before sending when approval rules apply.
Change supplier bank details.The finance owner controls verification.Complete independent verification before use.
Grant system access.The access owner approves the scope.Check identity and permissions before granting.

These examples describe possible control placements, rather than universal approval policies. Your company must set the actual rules for its systems and obligations. Keep existing approvals in the same workflow where possible so staff don't approve the same commitment again in a detached AI dashboard.

If you automate a bounded class of decisions, define what belongs in that class and test the routing. Check cases near the boundary, missing fields, unfamiliar counterparties, and conflicting source records. A gate that routes the wrong case to automatic execution defeats an otherwise careful review design.

Give reviewers evidence they can check

A review screen should show the proposed action, its scope, and the evidence that supports it. For a quote, show the customer, line items, prices, delivery promise, and terms that change. Let the reviewer open the approved pricing source and customer request without searching another system for each field.

Distinguish the original source from the AI's summary. A fluent explanation can repeat the same mistake as the proposed output. Show the underlying record or document, its effective date, and any disagreement between sources. If the evidence doesn't resolve the decision, let the reviewer request clarification or take the case out of the automated workflow.

Microsoft Research's human-AI interaction guidelines, published in 2019, recommend contextually relevant information, efficient correction and dismissal, and controls over system behaviour. Apply those principles to review: make rejection and editing straightforward, and explain what happens after each choice. The guidelines don't establish that a particular approval screen prevents errors.

Don't ask the reviewer to infer reliability from a model's self-reported confidence. If you use a score to route cases, evaluate whether it predicts errors on your own task and population. Missing evidence, a prohibited action, or an out-of-scope request can require escalation regardless of that score.

Train reviewers on the actual failure patterns. Give them practice cases with wrong source matches, plausible but unsupported claims, and changes hidden inside otherwise correct output. Evaluate whether they detect and explain the problem. Put this exercise in a controlled environment so a missed test case can't affect a real customer.

Keep independent assurance outside the daily approval pressure. An assessor can review a sample of accepted and rejected cases against the original evidence and business outcome. Approval rates alone don't tell you whether staff caught mistakes; a high rate may reflect good outputs, poor scrutiny, or a queue that rewards speed.

Staff the queue and measure its capacity

Approval creates an operating queue. Give it a roster, a backup, and a service expectation that matches the commitment. “The manager will check it” fails when that manager has other work or takes leave. Show who currently holds the case and when another reviewer takes over.

A June 2026 discussion about an agent approval queue describes work waiting when its usual reviewer went away without coverage. This is a reader's reported experience, not evidence of typical failure rates. It raises a concrete design question: who receives the pending work when the named person isn't available?

Calculate demand before expanding production. Consider this illustrative example, with assumed volumes rather than an industry benchmark. A workflow handles 600 cases each day and sends 20% for review. That produces 120 reviews. At four minutes each, the queue needs eight hours of direct review time per day.

Two reviewers with three hours each allocated to this work provide six hours of capacity. They can finish 90 cases at the assumed review time, leaving 30 extra pending cases each day. The gap exists before meetings, difficult escalations, or rework. Nominal headcount doesn't tell you how many hours staff can actually spend in the queue.

Measure the distribution of review time and arrivals. An average can hide a lunchtime surge or a group of cases that takes much longer. Track the oldest pending item, missed deadlines, and cases waiting on specialist input. Leave capacity for variation; a queue planned at its full theoretical capacity has little room for a complicated case.

When demand exceeds capacity, reduce the incoming scope, assign more trained cover, or improve the process that creates avoidable exceptions. Don't silently turn overdue items into automatic approvals. Include review and downstream repair in the project's operating cost; our guide to converting time saved into business benefits explains why freed capacity needs its own evidence.

Set escalation and stop rules

Give each escalation a receiving role and a next action. A reviewer who finds contradictory delivery dates can send the case to the planner who owns the source schedule. An attempted unauthorised action needs the engineering or security responder who can contain it. A generic “needs review” label can conceal both problems in the same queue.

Write thresholds in terms staff can observe. Examples include an unverified bank-detail change, a quote outside an approved margin rule, a missing required source, or a case approaching its contractual cutoff. Use your own policy values and evidence. A universal model-confidence threshold doesn't describe these business conditions.

Distinguish rejecting a case from stopping a workflow. A malformed document may need manual handling while other cases continue. Repeated use of the wrong customer record may justify suspending that affected task. Name who can pause execution and who approves restarting after engineering fixes and tests the cause.

Choose the timeout behaviour explicitly. For a consequential action awaiting approval, a missed deadline can hold the action and alert the responsible person. For time-critical work, provide an agreed manual route. Make sure staff can recover the original inputs and finish the job without relying on the failing assistant.

Test a pause in the application, including work already queued. Turning off new requests won't necessarily stop pending execution. Check what happens to approved but unsent actions, retries, and work another system already received. Record what staff must reconcile before restarting to avoid duplicate commitments.

Bind the approval to the action and its record

An approval must apply to the action the person inspected. If the system changes the amount, recipient, permissions, or message after approval, it needs a fresh decision. Engineers should preserve the reviewed payload and check it at execution, with controls that reject an expired or mismatched approval.

A September 2026 preprint on “Loopjacking” examines cases where a human reviews an action but the system uses that approval for a materially different operation. The author reports findings in a selected set of products and versions, and explicitly says they don't estimate prevalence across the market. Treat it as a specific failure mode to test in your implementation.

Keep a linked record of the proposed action, evidence references, reviewer identity, decision time, and final execution result. Include the workflow configuration and model version identifiers where they affect reconstruction. An approval log without the actual outcome can't tell you whether the system sent the approved quote or failed before sending it.

Use a stable case identifier across the review queue and receiving system. Record whether the action succeeded, failed, or has an unresolved outcome. If a network timeout leaves delivery uncertain, reconcile with the receiving system before retrying. Review quality includes avoiding duplicate actions after an ambiguous response.

Preserve enough evidence to investigate without copying every sensitive document into another log. Use controlled references or snapshots where appropriate, restrict access, and agree retention with the relevant data owner. Check that reviewers can still retrieve the evidence they need during the retention period.

For an illustrative quote workflow, the record can link the original request, approved price-list version, proposed quote revision, reviewer approval, and the message actually sent. If a customer later queries a term, the commercial owner can inspect that chain rather than reconstructing it from a chat transcript.

Assign ownership for fixing recurring failures

A reviewer can repair today's case while the process produces the same error tomorrow. Give recurring failures an owner who can change the source record, business rule, integration, or model workflow. Keep that work visible alongside the queue so case handling doesn't consume the entire improvement budget.

Use a short reason code and a concrete example when a reviewer rejects or edits an output. Separate “wrong source document” from “source document has the wrong value”. Engineering owns retrieval or mapping faults; the relevant business owner owns its source content. Assign disagreements about policy to the person authorised to decide the rule.

Have the process owner review repeated corrections on a scheduled cadence. Ask what caused them and how the proposed repair will change the operating result. A prompt edit may address wording while leaving an outdated price feed untouched. Verify the changed source or control before removing a review requirement.

Test fixes against examples of the failure and a broader held-out set of cases. Reviewer edits provide evidence for investigation; they don't automatically become training data or proof that the next version improves. Track whether the change reduces that error without introducing another failure elsewhere.

Tell reviewers when a rule or source changes and what they now need to check. Keep a way to challenge an accepted result after the case closes. Some mistakes emerge only when a customer responds or another team tries to fulfil the commitment. Connect that downstream signal back to the original case and its owner.

Test the whole review process before expanding

Begin with a bounded workflow and a staffed pilot. Measure accepted outcomes, review time, queue delay, and downstream rework against the previous process. Include the cases the assistant cannot handle. A fast path for selected easy cases doesn't establish that the whole operation costs less or finishes sooner.

Run controlled exercises for reviewer absence, missing evidence, deadline expiry, payload changes after approval, and an uncertain execution result. Confirm who receives the alert and what they do next. A documented escalation rule needs a working receiving team and application behaviour.

Use assurance sampling to examine both reviewed and automatically executed cases. Define how you select the sample and which populations it covers. Zero observed errors in a small sample can't prove that rare failures won't occur. Keep severe incident counts and unresolved cases visible beside any aggregate acceptance rate.

Expand only when the receiving owner can explain the current limits and support the larger workload. If volume doubles, check whether review demand also doubles and whether specialists can absorb more exceptions. Changes to the model, source system, or permitted actions can alter the result even when the interface looks the same.

At the investment or management review, ask for a recent case traced from input to final outcome, the current queue ageing, and an example of a recurring error the team fixed. These records help a VC assess delivery discipline and a PE operating team compare control maturity across companies without claiming that different workflows have identical risks.

For your first implementation, choose a decision your team already understands and attach review to its existing operating system. Assign an owner and backup, make the evidence inspectable, and test a rejection through to its final outcome. Broader autonomy becomes a decision you can support with observed results and available operating capacity.

Put the work into practice

AI implementation and delivery

We help SMEs and scale-ups put AI into a specific business workflow. We define the problem, prepare the data, build the software, and help your team operate it in production.