Guide

AI Adoption at Work: Choose the Task Before the Tool

Published 19 min read

Your company buys an AI assistant and gives your team access. Six weeks later, the usage dashboard shows that about a third of the team opens it in a typical week. Someone in leadership asks the obvious question: why are people resisting?

That question assumes the tool fits the work and the people are the problem. Sometimes that is true. Sometimes the tool may save no time on a particular task, produce errors that are slow to find, or lack access to the data the task requires.

So the useful questions about AI adoption come in a different order. Which tasks in this workflow suit the tool? Which people can use it well on those tasks? And what work outcome should improve if it is working? This guide answers them with a task-selection matrix and a pilot that measures finished work.

In short: pilot AI only on tasks where a good output saves real time, a wrong one is cheap to catch, and the people using it can check its work. Then judge the pilot on finished work and quality, not on how often people log in.

Start with tasks, not the tool

“Use the AI assistant” can mean drafting a reply, summarizing a document, checking a calculation, or translating a message. Each is a different task with different benefits and risks.

This is the organizational version of an argument I make about products. In Behavior Market Fit, I argue that the behavior you choose largely decides whether a product has a chance, and that you should compare candidate behaviors before you build anything. The behavior change guide puts the sequence as problem, behavior, solution, product. For AI at work, the sequence is outcome, task, person, tool.

The running example. This case is hypothetical; the team, people and details are invented for teaching. Maya leads a 14-person customer support team at a company that sells scheduling software to dental clinics. Her company has approved an AI assistant built into the help desk, connected to ticket text and the public help-center articles. She has five candidate tasks:

  • A. Drafting replies to routine “how do I” tickets. These are most of the team’s volume.
  • B. Summarizing long ticket histories before an escalation to engineering.
  • C. Calculating prorated refunds when a clinic changes plans mid-cycle.
  • D. Writing the weekly trends report for the product team from ticket tags.
  • E. Translating replies for Spanish-speaking clinics. Two agents are bilingual.

Maya’s outcome is resolving clinics’ problems correctly and quickly. AI use is a possible means, not the outcome.

Why usage numbers mislead

A large randomized study shows why a usage dashboard cannot settle the question. Eleanor Wiske Dillon and colleagues randomly gave Microsoft 365 Copilot to some workers and not others across 66 large firms and 7,137 knowledge workers, from September 2023 to October 2024. Three of the four authors are at Microsoft Research, so this is a study of Microsoft’s own product.

Use peaked early, at about 55% of treated workers, and then fell. After about 12 weeks, weekly use settled at just under 40%, and 20% of treated workers did not use it at all in months four to six.

The work effects were narrow. Workers given access spent about 1.4 fewer hours a week in email, a 12% reduction; among those who used the tool, the authors estimate about two hours, or 17%. But treated and control workers replied to the same number of email threads, attended the same number of meetings, and completed the same number of documents. The authors could not observe the quality of anyone’s work; their data came from activity logs.

The time savings showed up in email, which each person manages alone, and not in meetings or document writing. The authors suggest that changing shared work requires colleagues to agree on new norms, which an individual license does not create. If that explanation is right, tasks one person controls will change first.

The same tool helps some tasks and hurts others

Fabrizio Dell’Acqua and colleagues call this a jagged technological frontier: the same tool can improve one task and worsen another of similar apparent difficulty. Their field experiment with 758 Boston Consulting Group consultants, published in Organization Science in 2026, randomly assigned people to work with or without GPT-4 in 2023.

On 18 product-development tasks the researchers had confirmed were within the tool’s capabilities, consultants with AI completed 12.2% more tasks and worked about 25% faster. Their human-graded quality scores were roughly 30% to 34% above the control group’s average. On a business case designed to sit outside the frontier, the result reversed. Control participants reached the correct recommendation about 84.5% of the time; the two AI groups did so 60% and 70.6% of the time, an average drop of 19 percentage points.

AI-assisted answers were also rated more coherent and persuasive whether or not they were correct. A wrong answer that reads well is harder to catch. The study used one firm, elite consultants, one outside-the-frontier task, and a 2023 model. The frontier’s location moves as tools change; its jaggedness is the lesson.

Two other studies add the question of who benefits. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered rollout of a GPT-based assistant that suggests responses to customer support agents at a software company. Their data covered 5,172 agents, 1,636 of whom had access to the tool.

Issues resolved per hour rose 15% on average. Less experienced and lower-skilled agents improved in both speed and quality. The most experienced and highest-skilled agents saw small gains in speed and small declines in quality.

Shakked Noy and Whitney Zhang gave ChatGPT to half of 453 college-educated professionals doing occupation-specific writing tasks. Average time fell 40% and graded quality rose 18%. Their earlier working paper notes the main limit: the tasks were short, self-contained and lacked context-specific knowledge.

For a team lead, their follow-up matters most. People given ChatGPT in the experiment were twice as likely to report using it in their real jobs two weeks later, and 1.6 times as likely after two months. The working paper’s preliminary follow-up gives detail: among people new to ChatGPT, 26% of the exposed group were using it at work, against 9% of the control group. Yet people not using it at work mostly said it lacked the context-specific knowledge their writing required.

So one study points to two causes of non-use. Unfamiliarity can be fixed by a chance to try the tool; poor fit cannot. A usage dashboard counts both as the same zero.

Build the task-selection matrix

Rate each task on six properties as Favorable, Mixed or Unfavorable. Do not add the ratings into a single score; several work as gates.

  1. Usefulness. How much time or quality could a good output gain? Favorable: a frequent task where a good output saves substantial time or improves the result. Unfavorable: a rare or quick task.
  2. Error consequence. What happens if a wrong output is used? Favorable: errors are cheap and caught before anyone relies on them. Unfavorable: an error could reach a customer’s money, health, legal record, or a decision that is hard to reverse.
  3. Verification cost. For someone competent to judge the output, how long does checking take? Favorable: much less time than doing the task. Unfavorable: nearly as long as doing it.
  4. Required skill. Can the people who would use the tool on this task judge its output and ask for a better one? Favorable: they can. Unfavorable: they cannot. If the answer differs between groups, rate each group separately.
  5. Permissions and data access. Is the tool approved for the data the task involves, and can it reach that data? Favorable: yes on both. Unfavorable: either answer is no.
  6. Workflow effort. Does the tool sit where the work happens? Favorable: inside the existing system. Unfavorable: extra copying, switching between systems, or coordination with other people.

Then apply four decision rules, in order:

  1. Unfavorable permissions means blocked. Fix the access or drop the task. Do not count non-use on a blocked task as resistance.
  2. Someone using the tool must be able to check its output. If required skill is Unfavorable for a group, that group does not use the tool on this task.
  3. Unfavorable error consequence plus unfavorable verification cost means do not pilot now. That is where a fluent wrong answer does the most damage.
  4. Rank the remaining tasks by usefulness, then by the fewest Unfavorable ratings. Pilot one or two. Treat every Mixed rating as something the pilot must measure.

The UK Cabinet Office’s Mitigating Hidden AI Risks Toolkit (June 2025), drawn from its rollout of a government tool called Assist, calls this problem task-tool mismatch. It advises testing what a tool is good at, telling people which tasks it suits, and pointing them to other tools for the rest. That is practice guidance from one rollout, not a controlled test.

Maya’s matrix

These ratings are Maya’s judgments before any pilot, to be tested.

Task (and people, if ratings split) Usefulness Error consequence Verification cost Required skill Permissions Workflow effort Decision
A. Draft routine replies, agents past onboarding Favorable Mixed Favorable Favorable Favorable Favorable Pilot first (rule 4)
A. Draft routine replies, new hires Favorable Mixed Favorable Mixed Favorable Favorable Second round, with a daily review sample (rule 4)
B. Summarize ticket history Mixed Mixed Unfavorable Favorable Favorable Favorable Third in rank; quick timing test only (rule 4)
C. Calculate prorated refunds Mixed Unfavorable Unfavorable Mixed Unfavorable Unfavorable Do not pilot (rules 1 and 3)
D. Weekly trends report Mixed Mixed Mixed Mixed Unfavorable Unfavorable Blocked (rule 1)
E. Translate replies, bilingual agents Mixed Mixed Favorable Favorable Favorable Favorable Second in rank; pilot after A (rule 4)
E. Translate replies, all other agents Mixed Mixed Favorable Unfavorable Favorable Favorable Do not use (rule 2)

Task A is the strongest candidate. Experienced agents know the product, so checking a draft takes less time than writing one. A wrong instruction would annoy a clinic and probably reopen the ticket: a real cost, but a recoverable one. That is why error consequence is Mixed, and why the pilot must track reopened tickets.

Task B looks attractive until you ask how an agent would check the summary. To confirm that it did not drop an important detail, the agent has to read the history, which is the work the summary was supposed to save. It ranks third, so Maya only times ten escalations with and without a summary.

Task C fails two rules. A wrong refund costs money, checking it means redoing the calculation, and the tool cannot see billing data. The better fix may not involve AI at all: a proration calculator inside the billing system addresses the actual problem. Choosing the task first keeps that alternative visible.

Task D is blocked: nobody should be asked why they skip a tool that cannot read the report’s data.

Task E shows the person question directly. A bilingual reader can check a translated reply much faster than writing one, so verification cost is Favorable. But agents who cannot read Spanish could not check the output at all, so rule 2 excludes them.

Decide which people should use it on which tasks

The studies above suggest a risk worth planning for. Less experienced agents gained most in the support study. Wrong AI-assisted answers sounded more convincing in the consulting study. So the people who gain most may be the least able to catch mistakes.

That does not mean keeping new staff away from AI. It means matching the person, the task and the check. Maya’s new hires can use drafts for routine tickets, where they have the most to gain, with a senior agent reviewing a sample each day during their first months. Experienced agents can use drafts or ignore them.

The new hires join after the first pilot, not during it. That pilot needs a clean quality baseline from agents who can judge drafts unaided. Their Mixed skill rating is then measured in its own round (rule 4).

The check needs as much design as the task. The automation bias entry explains why “review carefully” is weak advice: a concrete check names an independent source of evidence and a trigger. For routine replies, the agent confirms that the help-center article behind the draft matches the clinic’s question and product version, then compares the draft with that article. The human-in-the-loop AI guide covers how to test whether a review step catches the errors that matter.

Run a pilot that measures finished work

A pilot should answer one question: does using the tool on this task, by these people, improve the outcome enough to justify its cost and risk? Usage helps interpret the answer; it cannot be the answer. If you reward usage, people will generate drafts they do not need, and the number will rise while the work stays the same. That is Goodhart’s law in a help desk.

Here is Maya’s pilot for task A.

  • Who and when. The 12 agents past onboarding. Pick six at random to start in week 1; the other six start in week 5 unless the pilot has been stopped. Weeks 1 to 4 compare people with and without the tool during the same period, which guards against seasonal changes in ticket volume.
  • Primary outcome: completed work. Routine tickets resolved per agent-hour, from the help desk.
  • Quality. The share of tickets reopened within seven days. Each week, a senior agent scores ten random replies per agent for correctness, completeness and tone, without being told the agent’s group. The reviewer may still guess from writing style; record that limit.
  • Serious errors. Any reply with instructions that could corrupt a clinic’s appointment data. Log every one, and stop the pilot for review after the first.
  • Time, including checking. Handle time per ticket, so verification counts as part of the cost.
  • Usage, as a diagnostic. Whether each draft was sent as is, edited, or discarded. Heavy discarding is information about fit.
  • Non-users. Ask agents who rarely use drafts to walk through the last ticket where they skipped one. The UK toolkit makes the same point: gather feedback from non-users, not only enthusiasts.
  • Decision rules, set before the pilot starts. Expand if resolved tickets per hour are at least 10% higher than in the comparison group during weeks 1 to 4, with a reopen rate no more than one percentage point higher. Revise if speed improves but quality falls. If quality improves without a speed change, decide whether that gain alone justifies the cost. Drop the task if neither moves.

These thresholds are Maya’s choices, not benchmarks; set yours from what a gain is worth and what an error costs. Twelve people over four weeks give a rough estimate, so decide in advance what you will do if the result is unclear. The research methods guide shows how to state the comparison and decision rule before you see the data.

One design trap is specific to AI. Brynjolfsson and colleagues found that agents kept some gains during periods when the tool was down, which suggests they learned from its suggestions. So “on” and “off” weeks for the same person can blur the comparison. A staggered start avoids that problem within a person. It does not stop spillover in a shared queue, where comparison agents may copy good AI-drafted replies, so note any shared templates that appear.

Task E waits for A’s result. If A goes well, the two bilingual agents try translation for four weeks, reading every translated reply before sending, while Maya tracks time per Spanish-language ticket and their corrections.

Read the result before you decide

What the pilot shows What it likely means Next step
Resolved per hour up, reopen rate and review scores steady The task fits and the check works Make the task standard practice; keep sampling quality
Resolved per hour up, reopen rate or review errors up People are trusting drafts they should be checking Delay the second group and redesign the check
Little change, many drafts discarded Poor fit for this task, at least with this tool Drop the task or test a narrower version
Gains for newer agents, none for experienced agents Fit differs by person Make the tool optional for experienced agents; keep review for newer ones
Low use concentrated among people who never tried it Access, skill, unclear expectations or fear Ask them; fix what they report; then measure again

Whatever the pattern, update the matrix: some judgments are now measurements.

Fear of AI, resistance, or poor fit?

Fear of AI at work is common. In a Pew Research Center survey of 5,273 employed US adults in October 2024, 52% said they were worried about the future impact of AI use in the workplace, and 36% said they felt hopeful. Fear is one explanation for low use among several, each needing a different response. Two agents from Maya’s week 1 group show the difference.

Dev, one of the most experienced agents, rarely uses drafts. The resistance story writes itself. But Dev resolves more tickets per hour than almost anyone, and when he uses a draft, he rewrites most of it. A rough timing of ten tickets shows his editing takes about as long as his writing.

That is row four of the result table, and it resembles what Brynjolfsson and colleagues found among highly skilled agents. Dev’s low use is a sensible response to a task where the tool adds little for him; pushing him would likely cost time without improving answers.

Lucia has never opened a draft. When Maya asks, Lucia mentions a rumor that the team will shrink once “the AI handles the easy tickets.” Here fear is a plausible explanation, and it needs a direct answer about roles, as the UK toolkit advises. If no cuts are planned, Maya can say so. If roles will change, she should say that plainly too, because reassurance she cannot keep will make the next rollout harder.

Either way, Lucia’s non-use says nothing yet about whether drafts fit her work. Sticking with the current way of working is not proof of status quo bias when real switching costs exist. Diffusion of innovations research stresses relative advantage and compatibility, which the matrix rates task by task. If use of a chosen task still lags for unclear reasons, the Behavioral State Model example shows how to test competing explanations.

When low use really is resistance, and when the matrix stops working

Sometimes the resistance story is right. Suppose a task rates well, a pilot shows faster work with steady quality, and access and training are in place. If some people still never try the tool, fear, habit or concerns about status deserve direct attention.

The change management models guide compares approaches to that kind of change and the evidence for each. In the behavior change guide, I argue that most change management initiatives specify desired outcomes without specifying the behaviors that make them up. AI change management has the same weakness when the goal is “embrace AI” rather than a named task done a named way.

The matrix has limits too. It assumes someone can check the output in reasonable time; a forecast that resolves in a year needs a different evaluation, such as comparing past predictions with what happened. It is built for tools people use directly, not fully automated systems. And the ratings expire: a task rated Unfavorable this quarter may deserve a fresh test next quarter.

Try it on a new case

A finance manager at a mid-sized nonprofit has an approved AI assistant and two candidate tasks. The first is drafting the narrative sections of grant budget reports. Her team wrote the budgets, and every grant report already goes through a finance review before submission. The second is assigning each credit card transaction to the right account code from its receipt; those transactions feed audited financial statements. Which should she pilot first, and what should she measure?

An answer. Pilot the budget narratives first. They recur and take time to write, a wrong sentence would be caught in the existing finance review, and the reviewers wrote the budgets, so checking is quick.

The transaction coding is higher volume, but errors flow into audited accounts and checking each code means rereading each receipt. Error consequence and verification cost both risk Unfavorable, and the permissions question for financial data comes first.

To learn about the second task safely, run the tool on last quarter’s transactions, already coded correctly, and compare. For the narratives, measure time to a submitted report and reviewer corrections, not how often the assistant was opened.

Choose the next task deliberately

Back to the dashboard that showed a third of the team using the tool. That number cannot tell you whether people are resisting, because it does not say which tasks they tried, whether the tool helped, or whether they could check its work. Low use on a well-chosen, tested task is the point where a conversation about fear or habit becomes useful.

If you would like outside help, I advise leaders on behavioral prediction and behavior change: what customers and employees are likely to do, and what that means for their business. Details are on my consulting page. It is my own service, so weigh the suggestion accordingly; everything in this guide works without it.

Your next step: list the tasks in one workflow, rate them with the worksheet below, and pick the one or two you will pilot first.

AI Task Selection and Pilot Worksheet

Use this worksheet with the guide AI Adoption at Work: Choose the Task Before the Tool. Part 1 helps you choose which tasks suit the tool, Part 2 matches people to tasks, Part 3 plans a pilot that measures finished work, and Part 4 helps you read the result. A filled, hypothetical example follows the blank version.


Part 1: Choose the tasks

Workflow: ______________________________________________

Outcome the work should improve (not “AI use”): ______________________________________________

Tool and what it is approved to access: ______________________________________________

Rate each task

Rate every task on all six properties as F (Favorable), M (Mixed) or U (Unfavorable). Do not add the ratings into a score. If a rating differs by person, write one row per group of people.

Property Question Favorable Unfavorable
1. Usefulness How much time or quality could a good output gain? A frequent task where a good output saves substantial time or improves the result A rare or quick task
2. Error consequence What happens if a wrong output is used? Errors are cheap and caught before anyone relies on them An error could reach a customer’s money, health, legal record, or a decision that is hard to reverse
3. Verification cost For someone competent to judge the output, how long does checking take? Much less time than doing the task Nearly as long as doing it
4. Required skill Can the people who would use the tool on this task judge its output and ask for a better one? If the answer differs between groups, rate each group separately. They can They cannot
5. Permissions and data access Is the tool approved for the data the task involves, and can it reach that data? Yes on both Either answer is no
6. Workflow effort Does the tool sit where the work happens? Inside the existing system Extra copying, switching between systems, or coordination with other people
Task (and people, if ratings split) 1. Usefulness 2. Error consequence 3. Verification cost 4. Required skill 5. Permissions 6. Workflow effort Decision

Apply the decision rules, in order

  1. Unfavorable permissions means blocked. Fix the access or drop the task. Do not count non-use on a blocked task as resistance.
  2. Someone using the tool must be able to check its output. If required skill is Unfavorable for a group, that group does not use the tool on this task.
  3. Unfavorable error consequence plus unfavorable verification cost means do not pilot now. That is where a fluent wrong answer does the most damage.
  4. Rank the remaining tasks by usefulness, then by the fewest Unfavorable ratings. Pilot one or two. Treat every Mixed rating as something the pilot must measure.

Blocked tasks and what would unblock them: ______________________________________________

Groups excluded from a task by rule 2: ______________________________________________

Tasks not to pilot now, and any non-AI alternative for the same outcome: ______________________________________________

Tasks to pilot (one or two), in rank order: ______________________________________________

Lower-ranked tasks and any quick test to run meanwhile: ______________________________________________

Mixed ratings the pilot must measure: ______________________________________________


Part 2: Match people, tasks and checks

Task Who will use the tool on it Who should not yet, and why How the output is checked (evidence compared, trigger, who checks)

Part 3: Plan the pilot

Task piloted: ______________________________________________

Item Your plan
Who and when (who takes part, how the start is assigned, the comparison period)
Primary outcome: completed work (the unit of finished work, and where it is counted)
Quality (the quality measures, and who reviews a sample without knowing the group)
Serious errors (what counts as one, and what happens when one occurs)
Time, including checking
Usage, as a diagnostic (used as is, edited, or discarded)
Decision rules, set before the pilot starts (expand, revise, drop, each compared with the comparison group)
Non-users to ask, and what to ask them
Spillover to watch (ways the comparison group could pick up the tool’s effects)

What you will do if the result is unclear: ______________________________________________

Next pilot, if any: ______________________________________________


Part 4: Read the result

What the pilot shows What it likely means Next step
Completed work up, quality steady The task fits and the check works Make the task standard practice; keep sampling quality
Completed work up, errors or reopened work up People are trusting outputs they should be checking Delay wider use and redesign the check
Little change, many outputs discarded Poor fit for this task, at least with this tool Drop the task or test a narrower version
Gains for some people, none for others Fit differs by person Make the tool optional where it adds nothing; keep review where it is needed
Low use concentrated among people who never tried it Access, skill, unclear expectations or fear Ask them; fix what they report; then measure again

What the pilot showed: ______________________________________________

Ratings to update in Part 1: ______________________________________________

Decision: ______________________________________________


Filled example (hypothetical)

This example is invented for teaching. The team, people, ratings and thresholds are not from a real organization or study.

Part 1: Choose the tasks

Workflow: Customer support for scheduling software sold to dental clinics. Team of 14 agents.

Outcome the work should improve: Clinics’ problems resolved correctly and quickly.

Tool and what it is approved to access: An AI assistant built into the help desk, approved for ticket text and the public help-center articles. No access to billing data or the reporting export.

Task (and people, if ratings split) 1. Usefulness 2. Error consequence 3. Verification cost 4. Required skill 5. Permissions 6. Workflow effort Decision
A. Draft routine replies, agents past onboarding F M F F F F Pilot first (rule 4)
A. Draft routine replies, new hires F M F M F F Second round, with a daily review sample (rule 4)
B. Summarize ticket history before escalation M M U F F F Third in rank; quick timing test only (rule 4)
C. Calculate prorated refunds M U U M U U Do not pilot (rules 1 and 3)
D. Weekly trends report from ticket tags M M M M U U Blocked (rule 1)
E. Translate replies, bilingual agents M M F F F F Second in rank; pilot after A (rule 4)
E. Translate replies, all other agents M M F U F F Do not use (rule 2)

Blocked tasks and what would unblock them: D. The tool cannot read the reporting export. Ask IT whether the export can be approved and connected. C is also blocked, because the tool cannot see billing data, but rule 3 would still rule it out if access were fixed.

Groups excluded from a task by rule 2: E for all agents except the two bilingual agents, because they cannot read the translated output.

Tasks not to pilot now, and any non-AI alternative for the same outcome: C. A wrong refund costs money and checking means redoing the calculation. A proration calculator inside the billing system may solve the real problem.

Tasks to pilot (one or two), in rank order: A with the 12 agents past onboarding. E with the two bilingual agents, after A’s result.

Lower-ranked tasks and any quick test to run meanwhile: B. Time ten escalations with and without a summary while A runs.

Mixed ratings the pilot must measure: A’s error consequence (reopened tickets and a reviewed sample). New hires’ required skill for A (a daily review sample after the pilot). E’s usefulness and error consequence (time per Spanish-language ticket and the corrections the bilingual agents make).

Part 2: Match people, tasks and checks

Task Who will use the tool on it Who should not yet, and why How the output is checked (evidence compared, trigger, who checks)
A. Routine replies 12 agents past onboarding; experienced agents may ignore drafts The two newest agents, until after the pilot Before sending, the agent confirms that the help-center article behind the draft matches the clinic’s question and product version, then compares the draft with that article. After the pilot, a senior agent reviews a daily sample from new hires.
E. Translations The two bilingual agents Everyone else, because they cannot read the output A bilingual agent reads every translated reply before it is sent.

Part 3: Plan the pilot

Task piloted: A. Draft replies to routine “how do I” tickets.

Item Your plan
Who and when 12 agents past onboarding. Six chosen at random start in week 1; the other six start in week 5 unless the pilot has been stopped. Weeks 1 to 4 are the comparison period.
Primary outcome: completed work Routine tickets resolved per agent-hour, counted from the help desk.
Quality Share of tickets reopened within seven days. Each week a senior agent scores a random sample of ten replies per agent for correctness, completeness and tone, without being told the agent’s group. Record that the reviewer may guess from writing style.
Serious errors Any reply with instructions that could corrupt a clinic’s appointment data. Log every one; stop the pilot for review after the first.
Time, including checking Handle time per ticket.
Usage, as a diagnostic For each ticket: draft sent as is, edited, or discarded.
Decision rules, set before the pilot starts Expand if resolved tickets per hour are at least 10% higher than in the comparison group during weeks 1 to 4, with a reopen rate no more than one percentage point higher. Revise if speed improves but quality falls. If quality improves without a speed change, decide whether that gain alone justifies the cost. Drop the task if neither moves.
Non-users to ask, and what to ask them Any agent who uses drafts on fewer than one in five tickets: “Walk me through the last ticket where you skipped the draft. What would it have taken to use it?”
Spillover to watch Comparison agents may copy AI-drafted replies they see in the shared queue. Note any shared templates that appear during weeks 1 to 4.

What you will do if the result is unclear: Keep the tool optional for task A, continue the quality sample for four more weeks, and decide again.

Next pilot: If A goes well, the two bilingual agents try task E for four weeks, reading every translated reply before sending. Maya tracks time per Spanish-language ticket and their corrections.

Part 4: Read the result

The pilot has not run; this example stops at the plan. When it ends, Maya records what the pilot showed, updates the ratings for task A in Part 1, and writes her decision using the table above.

References