Test an AI assistant against the job it is allowed to do, including the moments when it should ask a question, hand over or stop. Before launch, agree the approved information, permitted actions, review owner and release conditions. A fluent answer is not enough: the answer and the action must both be appropriate.
Define the job in one paragraph
A useful pilot brief is narrow enough to review. For example: “Help customers understand our collection instructions, collect the details of an unresolved request and prepare a handoff to the support team. Do not confirm stock, accept orders or promise exceptions.”
That gives the team something concrete to test. “Answer customer questions” leaves too much undecided. Write down which sources the assistant may use, who maintains them and what happens if they disagree.
Also decide whether the assistant drafts for a colleague or responds directly to a customer. Those are different release decisions. A team-reviewed draft can be corrected before it leaves the business; a customer-facing answer needs stronger controls at the point of use.
Create a test set before tuning the answers
Collect representative questions from the team, removing personal information where it is not needed. Add missing-information cases and the exceptions people routinely escalate. Keep some questions separate from the examples used during setup so the final review is not just a replay of rehearsed wording.
For each case, write the expected behavior before looking at the assistant’s response. This can be “answer from the current collection policy,” “ask for the requested date,” or “prepare a handoff without confirming availability.” You do not need to prescribe one exact sentence.
| Scenario | Expected behavior | Failure to catch |
|---|---|---|
| A straightforward policy question | Answer from the approved source. | Invented hours or conditions. |
| The same question in different wording | Preserve the meaning and policy. | A contradictory answer. |
| An ambiguous date or item | Ask for the missing detail. | An unsupported assumption. |
| A stock question without inventory access | Ask the team to confirm. | A made-up availability promise. |
| Conflicting source documents | Flag the conflict for review. | Quietly choosing an unreliable answer. |
| A request outside the agreed service | Explain the boundary and route appropriately. | Accepting unsupported work. |
| A customer asks for an exception | Preserve the request for a person. | Approving the exception. |
| A connected tool is unavailable | State the limitation and offer the agreed fallback. | Claiming an action succeeded. |
Review sources, answers and actions separately
A citation can point to the right document while the answer still misstates it. A correct answer can be followed by an unauthorized action. Review both, along with the information handed to the team.
Our original Halden Desk concept puts a conversation beside its source and handoff. It is useful for discussing the interface and review process. Its predefined scenarios are not evidence that a live model has passed an evaluation.
For a real pilot, retain the test question, source version, observed response, action taken and reviewer decision. Keep records proportionate to the work and avoid retaining customer details merely because the software can store them.
Agree what blocks release
Set release conditions around consequence, not a single average score. A polite answer to most questions does not compensate for promising stock that does not exist or claiming to have sent a request when the tool failed.
- Block release
- Unsupported commitments, wrong access, lost handoffs or false success messages remain unresolved.
- Keep in team review
- The workflow is useful, but responses still need frequent corrections or exceptions are unclear.
- Approve the defined pilot
- The agreed cases pass, ownership is clear and a fallback is available for the limited launch.
These are suggested decision categories, not a certification scheme. Your team should set the exact thresholds and decide which failures matter most in its operation. NIST’s voluntary AI Risk Management Framework provides broader context for managing AI risk across design, use and evaluation.
Test the handoff as a colleague would receive it
Open the actual destination used by the team. Can someone see the original request, the relevant details, what the assistant already said and what needs a decision? Is there an owner? If the handoff arrives as an unexplained summary in a different tool, the assistant may have moved work rather than reduced it.
For an inquiry assistant, test duplicate messages, missing contact details and a customer changing their request. For a document-collection workflow, test an item that does not apply and a document marked received but not yet reviewed. Those states need deliberate handling.
Run a small pilot with a way back
Start with a defined channel, customer task and review period. Tell the team how to pause the assistant and use the ordinary process. Review representative outputs and every significant failure. Update the test set when a new case exposes a gap, then rerun the relevant checks after changes.
Compare time spent reviewing and correcting the work with the time the previous process required. Record missed handoffs and customer confusion as well as successful responses. Speed alone does not establish usefulness.
What to ask a supplier to hand over
Ask for the approved source list, scenario set, observed results, action permissions, human review process, running costs and instructions for pausing or updating the workflow. Make those deliverables part of the proposal rather than hoping they arrive after launch.
If the job can be handled with a form or an existing inbox feature, start there. If it needs an assistant, our AI assistant service is organized around one defined workflow, a controlled pilot and an operating handover.