A polished answer can hide an upstream failure
My AI assistant analyzed the wrong Reddit post this week. I sent it a shortlink. The link resolved to a real OpenClaw post, the page looked relevant, and the assistant produced a coherent explanation. It was still the wrong source.
Nothing crashed. A review focused on grammar or plausibility could have approved the answer. The useful part of the test was finding the failure before I treated a clean output as evidence that the workflow worked.
A pilot has to test the path to the answer. The output is one piece of that test.
Count the checking and correction work
Time saved is easy to understand, which is probably why it shows up first. It also gets flattering when the measurement stops as soon as the draft appears.
I would count the full cycle: preparation, source checking, reviewer time, corrections, reruns, handoffs, and closeout. If a five-minute draft creates twenty minutes of quiet detective work, the pilot should show that. The detective work may still be worthwhile. It just belongs in the result.
The measures I would put in a Pilot Review
- Source integrity — Did the workflow use the intended source, version, jurisdiction, record, and time period?
- Output quality — Were the important facts supported, and were missing or conflicting facts visible?
- Corrections — What did reviewers change, how often, and did the correction survive the next run?
- Stop conditions — Did the workflow pause on weak sources, missing facts, sensitive topics, or actions outside its authority?
- Reviewer effort — How much time went into checking, correcting, rerunning, and explaining the work?
- Failure modes — Which errors looked plausible enough to escape an ordinary output review?
- Prohibited uses — Did people try to use the workflow for decisions or data that were outside the approved test?
- Follow-through — After approval, did the action reach the right owner and leave evidence that it happened?
End with a scale, change, or stop decision
A pilot review should finish with a decision about the workflow, not a celebration of the demo. Some parts may be ready for a larger test. Some need tighter sources or a stronger human gate. Some should stop because the data, decision, or failure cost is wrong for the tool.
Uncertainty belongs in that decision. A limited local test can show that a workflow behaved as expected under those conditions. It cannot prove complete coverage, fairness, legal correctness, security, or production readiness.