An AI tool prepares a reply almost immediately. Someone reads it, opens the source document, corrects a service detail, removes an unsupported promise, and rewrites the last paragraph.
Did it save time?
You can answer only after counting the work needed to reach an acceptable result. Generation speed is one part of that job. Preparation, checking, correction, and handling failures belong in the same evaluation.
If you have already chosen one useful task for an AI pilot, the next step is deciding how you will judge it.
Define the finish before starting the timer
For an illustrative customer-service task, the finish might be a reply that an employee has checked and considers ready to send. The draft must answer the question, match the current service guide, and avoid making commitments the business has not approved.
That definition should apply to both the current process and the AI-assisted version. Comparing a finished manual reply with an unchecked AI draft leaves out part of the work.
Some questions will require a handoff instead of an answer. If the guide does not cover a requested service, an acceptable result might identify the missing information and route the question to the right person. Count that as its own outcome, not as a completed customer answer.
Choose examples that resemble the work
Gather a small evaluation set using information you have permission to use. Include ordinary questions as well as the cases that make someone pause: vague wording, an outdated reference, missing details, and a request outside the service guide.
Write down what an acceptable result would contain for each example. Keep some examples separate from those used to adjust the tool, so the final check includes work the setup has not already been tuned around.
NIST's Generative AI Profile advises against inferring broad capabilities from narrow, anecdotal assessments. It also recommends checking sources and citations during evaluation and ongoing monitoring. Those are useful reasons to look beyond a polished demonstration. NIST Generative AI Profile, MEASURE 2.5
Record the work around the draft
For each example, capture the time spent preparing the input, waiting for output, checking the result, and making corrections. Distinguish elapsed waiting from active work, especially if the employee can do something else while the tool runs.
Record the outcome alongside the time:
- Ready to use after review.
- Usable after small edits.
- Substantial correction required.
- Rejected or handed off because the tool could not help.
Keep a short description of each meaningful correction. An awkward sentence and an invented service commitment are different problems, even if both take a moment to edit.
Compare with the current process on similar examples. If the same person has already solved a question manually, account for that familiarity before crediting the tool with a faster second attempt. Varying the order or using comparable questions can make the comparison more useful.
Look closely at what the reviewer has to catch
A source reference helps only if the source supports the answer. In the service-reply example, a draft could cite the right guide while adding an installation service the guide explicitly excludes.
Show the question, the relevant source, and the draft together. Make it straightforward to edit, reject, or hand off the result. For this pilot, sending remains a deliberate employee action after review.
Ask the reviewer what felt difficult. Did the wording make an unsupported claim seem plausible? Did finding the source take longer than writing the answer? Was it obvious which details needed another person's judgment?
The interface is part of the test. A better source view or clearer treatment of missing information may improve the complete task more than another prompt revision.
Agree on a decision the evidence can support
Before the trial, define the conditions for continuing. For this example, that could mean less total active work on routine questions, acceptable replies after review, and a reliable handoff for questions the source cannot answer. Treat serious unsupported commitments as a reason to investigate, even if the average completion time looks good.
Include the rejected drafts and difficult cases in the findings. Keep setup effort, service costs, and ongoing supervision visible alongside per-task time. A short trial can guide the next decision; it cannot prove every future outcome.
If the result is mixed, narrow the task and try again. If review consistently removes the benefit, improving the source material or using a simpler workflow may be the more useful next step.
Bring a task and a definition of good
At Code + Carbon, we want an AI pilot to answer a practical question about how the work improves. Start with the task, the information available, and what someone must check before using the result.
Our AI service examples show prepared, fictional outputs you can explore. They illustrate an experience; they do not measure model performance or process your documents.
Tell us which task you want to evaluate and what a good result needs to include. An anonymized example and the checks you already make are a useful starting point.
Put the idea to work
What’s getting in your way?
Bring us a workflow, a question, or an idea. We’ll help you find a useful place to start.
