The first question about AI at work may not be which tool to buy. Ask which recurring task takes time, has an output someone can check, and allows mistakes to be corrected before they matter. Such a task lets the team learn through practice without immediately handing over an important decision.
The 30-day plan below is an editorial proposal for an internal trial, not a promise of savings. Its purpose is to support an evidence-based decision to continue, change the workflow or stop.
1. Choose one task with an owner and a boundary
Possible candidates include drafting a response from approved service information, sorting sample messages stripped of identifying details, or outlining a document from an existing template. The owner must know which source is authoritative and whether the data may be used in the chosen tool.
Replace a broad ambition such as “help customer service” with a task description: “Draft a response about service scope from these documents, with staff approval before every send.” State the excluded decisions too. The tool must not invent warranty terms, promise a price or approve an exception on behalf of an authorized person.
| Selection question | A sign the task is ready | If it is not ready |
|---|---|---|
| Who checks the output? | An owner understands the content | Assign a reviewer and review time |
| Which sources are allowed? | Approved, versioned documents exist | Organize the information first |
| What happens if it is wrong? | It can be corrected before leaving the team | Narrow the scope or use sample data |
| What is the comparison? | Examples and baseline completion times exist | Observe the current method first |
2. Measure quality before celebrating speed
A draft generated in ten seconds is not necessarily a task completed in ten seconds. Checking references, correcting facts and rewriting may consume the apparent saving. Measure time from starting the task to a reviewer accepting the output, including corrections.
Create a test set containing ordinary requests, ambiguous questions and cases the supplied information cannot answer. Check accuracy, completeness, unsupported additions and tone. Separate errors that prevent release from minor wording changes, so an average score does not conceal a serious failure.
For example, if a customer asks about a condition absent from the approved documents, the appropriate draft may ask a colleague to verify it. A fluent invented answer is not a successful completion. Define what the system should do when information is missing, then include that situation in the test.

Count AI’s benefit after checking and correcting its output.TIIS Insights
3. Make responsibility part of the trial
NIST’s AI Risk Management Framework groups risk-management activities into Govern, Map, Measure and Manage, with governance spanning the process. [1] For a small trial, use this to ask who owns the work, what context it operates in, how results will be checked and what happens when a problem is found.
NIST also provides a voluntary companion profile for generative AI. [2] Referring to that framework does not certify a project or establish compliance with every applicable requirement. The team must still examine its organization’s rules, tool settings and the data involved in the actual workflow.
Decide who can change the instructions, who approves source documents and what triggers a pause. Repeated material errors or an inability to verify an answer may justify stopping the trial. Keep useful examples of failures without unnecessarily circulating sensitive information.
4. Give each part of the 30 days a purpose
In week one, observe the existing method and agree acceptance criteria. In week two, run a limited test set. In week three, let intended users try the workflow while every output remains subject to review. Use week four to examine quality, total time, reviewer workload and cost. Adjust this suggested schedule to the task’s risk and frequency.
Avoid changing the instructions, reference documents and scoring criteria at once without a record. Otherwise the team cannot explain why results changed. Keep versions and a short change log, and rerun the same difficult cases after a significant revision.


5. Decide from completed work
Less time with more serious errors is not a satisfactory basis for expansion. If quality is adequate but reviewing becomes a bottleneck, narrow the task before adding users. If results are consistent, expand to one further group with a guide and an accountable owner.
The capability gap may concern problem definition, document checking or process design as much as operating a tool. The developing K-XCEL Professional framework is a starting point for discussing capability needs. Confirm the objectives, scope and program availability with the team first.
To connect learning with work after a course, read Did the training change the work?.
Before the trial starts
Start with one task whose result can be checked. Let evidence from completed work determine the next step. Knowing when to pause or decline AI assistance belongs within a capable workflow too.
References and further reading
Use these primary sources to check the supporting information. Examples, tables and trial plans in this article are editorial proposals.
- National Institute of Standards and TechnologyAI Risk Management Framework 1.0: Core (2023) ↗
Govern, Map, Measure and Manage across the AI system lifecycle.
- National Institute of Standards and TechnologyArtificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (2024) ↗
A companion risk profile focused on generative AI.
Discuss capability and workflow development
Start with the service information, then discuss a scope suited to your setting, users and goals.
Explore the related service


