Research / Change management
Make Staff Corrections Part of the AI Improvement Loop
Capture employee corrections inside normal work, review their meaning, and turn them into tested AI fixes without adding an unpaid feedback workload.

Staff corrections improve an AI workflow when someone captures the evidence, checks what the correction means, and owns a tested fix. Editing a generated answer doesn't automatically teach the system. Build the feedback loop into normal work, fund the time it takes, and show employees what changed after they reported a problem.
Start with the work employees already do
A service coordinator checks a drafted customer reply, changes its delivery date, and sends it. The customer receives a correct answer. The system may still generate the same wrong date tomorrow. Unless the correction reaches someone who can investigate and change the workflow, the coordinator keeps paying the repair cost.
That repeated repair can disappear from an AI project's dashboard. Drafting time falls, while checking, rewriting, and explaining mistakes occupy another part of the day. If you ask staff to log every repair in a separate spreadsheet, you also create a new administrative task. Include those minutes in the operating model before claiming a saving.
A worker's account on Reddit describes automation leaving them with additional checking and preparation, and managers asking for feedback after introducing changes. It's an individual account, not evidence of how often this happens. It does identify a question your own rollout needs to answer: who absorbs the work that the headline saving leaves out?
Observe a complete case with the people handling it. Record where they notice a problem, what evidence they consult, and how they correct the outcome. Ask what prevents a report during a busy shift. A feedback design that needs staff to reconstruct the case later loses context and makes participation harder.
Agree the purpose before collecting data. For example: reduce recurring delivery-date errors in drafted service replies without increasing the team's total handling time. That gives employees a reason to contribute and gives the project a testable outcome. Collecting corrections because they might become training data is too vague to guide a working process.
Capture corrections inside the workflow
Microsoft's HAX Toolkit guideline on granular feedback recommends letting users provide feedback during regular interaction and on individual outputs. Apply that principle at the place where staff already edit or reject the draft. A whole-answer thumbs-down gives less diagnostic context than identifying the date field they changed.
Where the workflow permits it, record the original field and its corrected value automatically when the employee saves. Keep the relevant case reference, system version, and event time. Offer a short optional reason, such as wrong source, missing context, or wording preference. Let staff flag uncertainty rather than forcing a confident diagnosis.
Capture only what the investigation needs under your agreed access and retention rules. A restricted reference to the source record may be enough; a second copy of the full customer file may add exposure without helping. Keep source access controls intact, and define who can view the correction record. Employees shouldn't need to paste private customer details into a general feedback channel.
Keep the urgent service path short. Staff need to complete the customer task even when they can't explain a correction immediately. Make optional feedback easy to skip, and give the receiving team enough context to investigate without a mandatory essay. Required safety or approval steps still apply; optional product feedback mustn't masquerade as another approval gate.
Explain the capture plainly in the interface and staff briefing. Say whether the system records edits, who reviews them, and whether any approved examples can enter tests or training. Don't label the button “Teach the AI” unless the product actually implements that behaviour and you can explain its scope. A saved correction can enter a review queue without changing any model weights.
Keep feedback connected to the existing case or issue system. Staff can follow a reference to see its status without maintaining a parallel log. For an SME, a controlled issue queue may be sufficient. A large organisation may need integration across several tools, but the person correcting the answer should still have a single place to act.
Check what each correction means
A correction is evidence of a disagreement. The receiving team still needs to determine whether the employee fixed a factual error, applied a local rule, expressed a preference, or made another mistake. Treat the edited value as a candidate label until a qualified owner checks it against the applicable evidence.
NIST's AI Risk Management Framework, GOVERN 5.2, calls for mechanisms that regularly incorporate adjudicated feedback into system design and implementation. For this workflow, adjudication means resolving what the correct outcome is before using the case to change the system. The framework doesn't make every staff edit a reliable training example.
Consider two branches that quote different delivery times. Both employees may follow their branch's approved rule. A central team that treats the most frequent correction as universally correct can break the other branch's work. Keep the applicable location, policy version, and date with the example, then ask the business owner whether the rule should differ.
| Staff changed | Check first | Possible action |
|---|---|---|
| A factual value. | Verify the source and date. | Repair source access or data. |
| A policy decision. | Confirm the local rule. | Clarify policy or routing. |
| The wording. | Separate style from meaning. | Adjust approved instructions. |
| An intended action. | Check authority and impact. | Review action controls. |
| An uncertain answer. | Ask a domain owner. | Keep it out of trusted labels. |
Give the reviewer the original input and output, the correction, the relevant source reference, and the system version where permitted. Record the failure category and the reason for the agreed answer. If policy owners disagree, show that unresolved status. Don't let an engineer settle a commercial policy dispute by selecting the easiest answer to implement.
Reuse needs a separate decision. A record that staff can access for customer service isn't automatically suitable for a shared test set or a vendor's training process. Check the intended use, permissions, and data handling before transferring it. Use a sanitised or constructed case when it preserves the failure accurately; verify that removing details hasn't changed what the test measures.
Fund the receiving team and protect candour
Name a triage owner who checks new reports and groups duplicate failures. Name a business owner who resolves the correct outcome, and an engineering owner who can implement changes. Set a review cadence and reserve capacity. Several roles can belong to the same person in a small team, but the responsibilities still need an owner.
Budget the staff time that creates the evidence. In an illustrative team, 20 employees each spending three minutes a day on extra feedback across 20 working days contribute 1,200 minutes, or 20 hours per month. A further two hours of triage each week adds eight hours in a four-week month. Those assumed 28 hours exclude engineering and domain review.
The arithmetic doesn't tell you whether feedback is worthwhile. It tells you what to include in the decision. Compare the burden with observed recurring repair and downstream mistakes. If the receiving team can't act on the incoming queue, simplify collection or narrow the pilot while the owner resolves capacity. A larger archive won't compensate for an unattended queue.
Managers need to adjust workload expectations. Ask for feedback during allocated working time, and recognise diagnosis as part of the project. Don't assume a skilled employee can maintain examples between customer calls while their existing target stays fixed. Put responsibility for maintaining the test set with the project team rather than an informal volunteer.
Explain how the organisation will use employee identifiers. Correction counts reflect case mix, reporting habits, and system failures, so they don't establish individual performance. Keep process improvement reporting separate from staff appraisal and restrict identifiable access to people who need it. If managers propose another use, explain and review it before collecting data for that purpose.
Let employees challenge a triage decision and report through an alternative channel when the normal route feels uncomfortable. Include new staff, different shifts, and the people who avoid the AI tool. A queue dominated by confident early adopters can miss the failures that caused others to stop using it.
Follow a delivery-date correction through the loop
Take an illustrative service desk that drafts replies from an order system and a delivery-policy library. An employee changes a promised Friday delivery to Tuesday after checking the order's confirmed date. They finish the customer reply, and the workflow records the field change with controlled references to the draft and order.
The employee selects “wrong date” and adds no further explanation during a busy shift. The triage owner reviews the case later and finds similar reports. They group the cases under a shared issue while keeping their individual context. That gives engineering a pattern to investigate without asking each employee to file a second ticket.
The business owner confirms that the reply must use the confirmed order date, and that missing confirmation requires escalation. Engineering finds that the workflow retrieves an outdated generic delivery paragraph before consulting the order. The immediate fix changes source priority and the missing-date rule. No model retraining is necessary for this particular diagnosis.
The reviewer creates approved test cases for a confirmed date, a missing date, and a conflict between the order and generic policy. These cases have different purposes: checking the authoritative source, checking escalation, and checking precedence. Engineering also runs the existing suite to check that the change doesn't disrupt other replies.
After a reviewed release, the owner tells staff which source rule changed and what to do when a date is missing. They check comparable live cases and the staff's repair time. If the wrong dates continue, they reopen the issue with the new evidence. Closing a ticket after editing a prompt doesn't prove that the working process improved.
Connect feedback to tests and a release
Make the status visible: received, awaiting evidence, decision agreed, fix planned, testing, or released. If the team rejects a proposed change, give a reason and identify the applicable rule. Staff need to know whether the report lacked evidence, reflected an allowed local variation, or reached a queue that the project hasn't funded.
After adjudication, connect the issue to an approved example and its expected outcome. Record the version of the rule that supports the answer. A test can then check that a later change preserves the behaviour. Keep examples with unresolved labels out of claims about accuracy, while tracking their unresolved status in the issue queue.
Anthropic's January 2026 guide to agent evaluations describes teams moving from manual feedback to evaluations as systems scale. It also explains that automated evaluations work alongside production monitoring and human review. Apply that combination here: a test result supports a release decision, and live evidence checks whether the fix helps in actual work.
Keep a separate set of cases for judging generalisation, with access that prevents developers from repeatedly tuning against it. Passing the exact examples that drove a fix proves a narrower point than improving unseen cases. Report which set informed development, which set assessed the release, and which live cases support the result.
Approve the change through the normal release process and record what actually changed: a source document, retrieval logic, instructions, or action control. If the team proposes training or fine-tuning, assess that as its own change with approved data and evaluation. Staff feedback can inform several kinds of improvement; the receiving owner chooses the appropriate mechanism.
Send a short, specific update to the contributing team. Include the affected workflow, the released version or date, and any limit that staff still need to handle. Link back to the issue. A release note such as “Confirmed order dates now take priority; missing dates still need review” lets employees change their working expectations.
Measure the burden and the business outcome
Measure time spent capturing feedback and investigating it alongside ordinary review and repair. Include engineering, domain decisions, and recurring maintenance in the cost view. Compare cases with similar complexity and volume. A drop in drafting time can coexist with an increase in total effort, and you need the complete workflow to see it.
Track repeat failures after each relevant release. For the service-desk example, count incorrect delivery promises against the number of eligible replies, and check downstream customer contacts about dates. State the observation period and how reviewers found errors. A handful of staff reports can't establish a reliable failure rate without coverage of the underlying cases.
Report queue age and the share of adjudicated issues that reached a tested change. A slow queue can show missing decision capacity; repeated requests for more evidence can show a poor capture design. Use these measures to improve the process. Don't reward raw report volume, which can encourage duplicate tickets and conceal whether anything changed.
Check for missing voices and changing behaviour. Fewer corrections may mean better output, lower usage, less reporting time, or staff losing confidence that anyone will act. Review a sample of accepted outputs and interview people who stopped contributing. Compare the volume and mix of eligible cases before interpreting a falling correction count.
For venture investors, request a small traceable sample from employee report to adjudication, release, and observed outcome. Ask how much time the startup's customers still spend fixing the workflow. For PE operating teams, compare the receiving capacity and local rules across portfolio companies before sharing labels. A common software platform doesn't make every company policy identical.
For a corporate programme, show which central team owns recurring failures and which local team approves rules. Keep the improvement budget visible next to the claimed operating saving. If the same issue persists across business units, review the shared source or component before asking every unit to produce more examples.
Run a bounded pilot with visible decisions
Start with a workflow where staff already make observable corrections and a business owner can resolve the answer. Pick a bounded group and agree what successful improvement looks like. Avoid collecting across the whole organisation before you know whether the receiving team can investigate the records or employees can contribute within their workload.
Before starting, show staff a real capture example, the permitted data, and the status they can see afterward. Reserve their contribution time and the receiving team's review time. Confirm an escalation route for serious failures that can't wait for a scheduled triage meeting. The feedback queue supports improvement; incident handling still needs its own response.
Choose a review date that fits the case volume and severity. At that review, inspect a few completed loops and unresolved cases with the staff who contributed. Ask whether the evidence was sufficient, whether the decision made sense, and whether the fix changed their work. Use those observations to adjust the capture design before expanding.
Continue when the pilot produces verified fixes and the full cost compares favourably with the outcomes you value. Narrow or stop it when staff spend time labelling cases that nobody can adjudicate or release. Give the team the evidence behind that decision and keep a working manual route for the customer task.
Your first practical step is to choose a recurring correction and name its receiving owner. Give that owner time to follow it through a decision, a test, and a live check. When employees can see the result of their contribution, you have a concrete process to repeat and a basis for deciding how much improvement work to fund.
