A timer makes many AI demonstrations look successful. It stops before the reviewer opens the sources, repairs the format, handles the exception, or redoes a rejected result.
The measurement that matters runs from intake to accepted handoff. Count review, correction, exceptions, rework, and downstream repair. A workflow earns continued use when the completed work meets its quality condition with less burden. It can also earn continued use by delivering a better result for a burden the team has chosen to bear.
This article provides general operational education. It is not legal advice or a compliance determination. No attorney-client relationship exists, and none of its protections apply.
Define the result that cannot be traded away
Start with the quality condition that the workflow cannot trade away. It might be accurate extraction from approved documents, complete issue spotting against a playbook, or correct routing to the responsible person.
Write that condition as an observable test. “High quality” is too vague. “Every extracted date matches the source document, and every uncertain field routes to review” can be checked.
Treat this condition as a gate. A workflow that fails it does not earn a positive result because it saved minutes elsewhere.
This approach follows a useful principle in the NIST AI Risk Management Framework Core. NIST calls for measurement with quantitative, qualitative, or mixed methods. It also calls for monitoring in conditions similar to deployment and for comparing performance with benchmarks. The framework is voluntary, cross-sector guidance, rather than a legal standard for a specific team.
Count the work that reaches acceptance
Choose a bounded unit that both processes can complete. One reviewed contract summary, one invoice intake, or one matter-opening record works better than an abstract hour of “AI use.”
Run a paired comparison where practical. Give the current process and the proposed process comparable inputs, output requirements, and reviewers. A small pilot will not prove a universal effect, but it can expose where the work moves.
Record five categories for each completed unit:
- Quality: Did the output meet the stated acceptance conditions? Which critical or noncritical errors appeared?
- Total effort: How much active time did intake, generation, review, correction, and handoff require?
- Rework: Did anyone reopen the work because of an omission, wrong answer, format problem, or downstream mismatch?
- Exceptions: Which inputs fell outside the intended workflow, and what happened to them?
- Operating cost: What tool, implementation, training, support, and oversight cost belongs to the measured volume?
Track adoption separately. Low use can reveal poor fit, weak training, missing access, or a process that people do not trust. It cannot establish the cause by itself.
Include effort spent on rejected attempts in the batch's total cost. Otherwise, the successful items can hide work that produced no accepted result.
A calculation that exposes where the time went
Consider a fictional contract-summary pilot. The numbers below are synthetic and illustrate the calculation only.
For one illustrative completed summary, the current process uses 52 minutes of active work, including input preparation. The proposed process takes 4 minutes for input preparation and generation, 22 for review, 12 for correction, and 5 for handoff. These non-overlapping stages total 43 minutes of active work. This single-unit example assumes no rejected attempts.
Use this one-unit arithmetic only to understand the calculation. A decision about the workflow belongs at batch level, where the denominator includes every started unit, including rejected attempts, and reports minutes per accepted unit. The workflow evidence card uses that batch denominator.
The difference is 9 minutes, or 17.3 percent of the illustrative 52-minute baseline. At a loaded labor rate of $180 per hour, those 9 minutes equal $27 of capacity per completed summary.
Assume $10 per summary in other incremental costs. This allocation includes software, implementation, training, support, and oversight costs not already counted in the active work. It adds no second charge for timed review. The tentative net capacity value is $17.
Review and correction account for 34 of the proposed process's 43 minutes, or about 79 percent. In a workflow with that profile, reviewer pace and judgment will move the result more than generation speed. Record who reviewed the work, their relevant experience, and enough units to see whether the apparent saving survives normal variation between reviewers.
Record waiting time separately to understand elapsed delivery time. For a batch, calculate each completed unit's total before comparing medians.
That result remains conditional. Capacity is not cash. The team realizes economic value only if it redirects the time to useful work, increases completed volume, or avoids a real expense.
The quality gate also comes first. If the proposed process misses a required clause or creates a material source error, the pilot fails that unit even if it is faster. The right next step may be a narrower task, better source controls, or no AI for that workflow.
Averages hide who benefited
AI effects can differ by task, user experience, and measurement date. A September 2026 NBER field experiment in patent drafting illustrates why one average should not become a universal promise.
The preregistered study enrolled 133 patent lawyers at 11 United States firms; 91 completed the full protocol. Blinded expert raters found gains in benchmark drafting quality during AI-assisted work. A later unaided redline exercise found an average advantage concentrated among senior participants, while junior participants showed no average gain and wider variation.
This is a working paper, not peer-reviewed evidence. Google funded the experiment's direct costs, and 6 of its 7 authors were Google employees. All 11 participating firms had ongoing patent-drafting relationships with Google that facilitated recruitment. The tool, InFlow, was an unreleased Google Labs product. Those relationships do not invalidate the study, but they belong in any assessment of how far to generalize it. The experiment covers one patent-drafting setting and does not predict results for contract intake, legal research, compliance work, or another tool.
Current industry evidence also warns against assuming that adoption equals value. The 2026 Thomson Reuters AI in Professional Services report surveyed 1,514 respondents across several professional sectors. Eighteen percent said their organizations collected AI return-on-investment metrics, while 40 percent did not know whether they did.
The survey was fielded in October and November 2025. It drew from Thomson Reuters lists and screened for AI familiarity. The 2026 report label therefore describes the publication, while the responses are from late 2025. The survey describes reported measurement practice in that sample, not the success rate of legal AI projects.
The measurement must change a decision
Do not end a pilot with a dashboard and no operating choice. Use the evidence card to choose a route.
- Continue: quality gates passed, the comparison shows a useful result, and the workflow has a named owner for monitoring.
- Redesign: the task still appears useful, but review, sources, exceptions, adoption, or cost prevents a reliable result.
- Stop: the workflow repeatedly fails a critical condition, shifts too much work into review, or has no credible path to realized value.
A stop decision can be a good result. It prevents a polished demo from becoming a permanent operating burden.
What a good batch looks like
A batch that supports continued use reads something like this, with fictional numbers. Twenty invoice-intake records, all 20 accepted, 2 returned once for a missing field. A median of 11 minutes of active time per accepted record against a 19-minute baseline, with the analyst's review counted in full. One exception, a foreign-currency invoice, routed to the manual queue as designed. Operating cost allocated per record and stated as an assumption.
The same batch with 3 records accepted only after a second correction is a redesign, whatever the median says. So is 1 wrong vendor posted before anyone caught it.
The workflow evidence card holds these fields in that order. Fill it for 1 baseline batch and 1 pilot batch, and write down the number that would reverse the decision before you see the result.
The Testing & Success Metrics framework in Ortaire's AI Governance Toolkit carries the card into a maintained record. It holds test scenarios, success measures, acceptance decisions, and ongoing monitoring.
Sources used
- NIST AI Risk Management Framework Core, accessed September 18, 2026. Used for benchmark, measurement, deployment-context, and monitoring principles.
- Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting, NBER Working Paper 35720, September 2026. Used for a bounded example of heterogeneous, task-specific effects.
- 2026 AI in Professional Services Report, Thomson Reuters Institute, 2026; survey fielded October and November 2025. Used for reported AI ROI-measurement practice and its disclosed sample limits.




