“We should be using AI” is a starting point for a conversation. To turn it into a useful project, name a piece of work you want to improve.

Perhaps someone repeatedly searches a service guide to answer customer questions. Perhaps they read incoming requests and pull out the same details each time. Perhaps they spend time preparing a first draft that another person needs to check.

For a first project, we would look for a task with a clear input, an answer someone can judge, and a manageable consequence when something goes wrong.

Start with a task you can describe

Consider an illustrative customer-service workflow: a business receives questions about what its installation service includes. An employee checks the current service guide, prepares a reply, and reviews it before sending.

An initial AI-assisted version could prepare a draft from that guide and show the material it used. The employee would still check the answer and decide what to send.

That is a specific job with visible boundaries. It gives you something to evaluate: whether the draft reflects the guide, addresses the question, and helps the employee finish the work.

Before choosing a tool, finish this sentence: “Given this information, prepare this result for this person to review.” If that is difficult, the workflow probably needs a little more definition first.

Put the source material in order

For the service-reply example, start with the guide itself. Is it current? Does it distinguish included work from optional extras? Who updates it when the service changes?

Write down which sources the assistant may use and which information is outside its scope. Decide who may access the source material and the resulting drafts. Choose test examples that you have permission to use, and remove unnecessary personal details.

The interface should make it easy for the reviewer to open the relevant source. A source reference is useful for checking an answer, but its presence alone does not prove that the answer is correct.

If the guide does not answer a question, the desired result may simply be a request for clarification or a handoff to someone who knows.

Define what a good result looks like

Collect a small set of representative questions before testing. Include straightforward requests, ambiguous wording, missing information, and a question the guide cannot answer.

For each one, describe what an acceptable draft should contain. In our example, the reviewer might check:

  • Does the reply accurately describe what the service includes?
  • Does it leave out unsupported prices, dates, and promises?
  • Can the reviewer find the source behind the answer?
  • Does it flag missing information instead of filling the gap?
  • Is the result clear enough to edit and use?

Keep some examples separate from the ones used to tune the setup. That gives you a better check of how it handles questions it has not already been adjusted around.

Make review part of the workflow

Generative AI can confidently produce incorrect information. NIST identifies this as a risk in its Generative AI Profile, including answers and citations that look convincing despite being wrong. NIST Generative AI Profile, section 2.2

For this first pilot, show the customer question, the source, and the draft together. Give the employee a clear way to edit, reject, or approve it. Keep sending as a deliberate action after review.

Also decide what happens when the source is unavailable, the question is outside scope, or the draft is unsuitable. The employee should be able to continue using the normal process.

A review step needs someone with enough context and time to do it. Assign that responsibility before the pilot starts.

Measure the complete task

A fast draft is only useful if the whole task improves. Compare the existing process with the assisted version, including the time spent checking sources, correcting errors, and preparing the final reply.

Record which drafts were useful, which needed substantial changes, and which were rejected. Look for patterns in the questions that caused problems.

Include the cost of the tool, any integration work, and the effort needed to keep the source material current. Decide what would justify continuing, changing the approach, or stopping the pilot.

Sometimes the exercise points to a simpler answer: a clearer guide, a reusable response template, or a better search feature. Those can be worthwhile outcomes too.

Bring one task to the first conversation

You do not need a company-wide AI plan to explore a useful starting point. Bring a recurring task, a few suitable examples, the information used to complete it, and a description of what a good result would look like.

At Code + Carbon, we would use that material to discuss whether AI is a reasonable fit, where review belongs, and what a small version should prove.

The first milestone should be a result you can inspect and compare with the way the work happens today.

Put the idea to work

What’s getting in your way?

Bring us a workflow, a question, or an idea. We’ll help you find a useful place to start.

Start a conversation
Back to all insights