Where should human review sit in an AI workflow?

A review step can exist on paper and still be useless. If the output has already changed a record, reached a client, or shaped a filing, the reviewer is investigating a consequence rather than controlling it.

Place human judgment at the last point where the named person can reject the work or change its route. Some workflows need more than one such point. Others need a narrow checkpoint and later sampling. The choice depends on the decision being made, not on a blanket promise that every output receives human review.

This article provides general operational education. It is not legal advice or a compliance determination. No attorney-client relationship exists, and none of its protections apply.

Follow the decisions, not the boxes

Begin with the trigger and trace the work to its destination. Mark each point where the AI receives sensitive information, creates a proposition, or causes an action.

Then ask two questions. Can a mistake still be corrected here? Does the reviewer have the context and authority to reject it?

The NIST AI Risk Management Framework says organizations should define roles for human-AI configurations and oversight. Its appendix also recognizes that some systems may not need human oversight, while others may require it.

This supports a proportionate approach. Human attention should follow the decisions and risks in the workflow.

Three places where review can still work

Consider 3 review moments when mapping a legal-team workflow. They serve different purposes.

1. Before sensitive access or expanded authority

Place a gate before the AI receives information or permissions beyond its approved boundary. The gate checks the matter, data class, tool, account, and requested action.

This moment answers whether the system may act at all. It is distinct from checking the quality of an output.

For example, an intake assistant may process approved public submissions. Route privileged matter files through the firm's approved data-handling process before any model receives them.

2. Before reliance or an external consequence

Place review after the AI produces work but before someone relies on it. This includes advice, a sent message, a filing, a record update, or another hard-to-reverse action.

The reviewer needs the source, the proposed output, and the acceptance criteria. A bare approval button is not enough.

The NIST Generative AI Profile recommends comparing outputs with known ground truth. It also recommends documented fact-checking when generated information has multiple or unknown sources.

The profile further suggests independent evaluations that match the identified risk. It does not prescribe one universal review method for every workflow.

3. After operation, through sampling and monitoring

Item-by-item approval cannot detect every pattern. A workflow also needs periodic review of errors, overrides, misses, and changing inputs.

Sampling can reveal changing errors or approvals made without enough time. Its results can inform a decision about review scope, subject to the team's applicable rules.

Post-use review does not cure a missed pre-action gate. It answers a different question: does the workflow still perform within its approved bounds?

Match the reviewer to the decision

The person checking format may not be qualified to assess legal analysis. The person who understands the law may not control system access.

Name the role for each checkpoint. Record what that person sees, what they decide, and what happens when they are unavailable.

For a consequential action, absence should usually block or queue the work. “Proceed if nobody responds” defeats the checkpoint.

The ABA's Formal Opinion 512 illustrates why review design must be task-specific. It discusses testing a tool on a smaller document subset before relying on larger-scale summaries.

The opinion also says lawyers cannot rely on generative AI alone for work requiring professional judgment. Lawyers remain responsible for client work.

The opinion applies the ABA Model Rules. Teams must consult the rules and authorities that govern their own work.

How the checkpoint moves with the consequence

These examples are fictional. They illustrate placement and do not report client results.

An approved system proposes a category and destination for a new request. Routine requests remain in a draft queue until an intake coordinator confirms the destination.

Ambiguous, urgent, or sensitive requests route to a lawyer. In this fictional design, a weekly sample checks false routing and missed urgency.

The before-reliance checkpoint protects the request. The sample tests whether the process remains dependable.

Contract drafting support

An AI assistant proposes clauses using an approved playbook. A lawyer reviews the source facts and each material clause before the draft leaves the team.

Any request outside the playbook stops and escalates. The system cannot send, negotiate, or accept language.

The main checkpoint sits before external use. An earlier gate controls which matters and data may enter the tool.

Court submission

An AI-assisted draft cannot move directly to filing. The responsible lawyer checks every cited authority, quotation, factual statement, and required disclosure.

This is more than a final read. The reviewer traces propositions to authoritative sources and corrects the draft before submission.

Formal Opinion 512 states that lawyers should review generative AI outputs for accuracy. It specifically addresses analysis and citations submitted to a court.

Review that looks stronger than it is

First, check whether the assigned review volume leaves enough time for each decision. Redirect low-stakes items only within the team's approved review rules.

Second, do not treat seniority as sufficient. The reviewer needs relevant expertise, source access, time, and rejection authority.

Third, do not rely on a policy label. “Human supervised” describes an intention until the workflow records who reviewed what and what changed.

Mark the point of no return

Take one workflow and draw five boxes: input, AI step, proposed output, human decision, and destination. Add a gate before each sensitive or consequential transition.

For every gate, name the reviewer, evidence, acceptance rule, and unavailable-person route. Then test one normal case and one exception.

Checkpoint worksheet

Use the checkpoint and ownership card to record the reviewer, source evidence, rejection route, and backup.

The Human Oversight & Accountability framework in Ortaire's AI Governance Toolkit is the record version of the checkpoint map above. It covers review responsibilities, permissions, escalation, and intervention arrangements in an editable file.

Does every AI task need a formal approval record?

No. A private brainstorming exercise and a proposed filing do not need identical records. The useful question is whether another responsible person could later understand what was proposed, what evidence mattered, who decided, and what remained unresolved.

Your firm's policies and applicable obligations may require particular records. This article does not determine those requirements.

Preserve the decision, not a second copy of the work

For a consequential checkpoint, identify the proposed output, relevant source, reviewer, decision, and time. Include unresolved questions and the correction or escalation route when needed.

Use an existing ticket, document history, or approval system if it already captures that information. A second form may add work without making the decision clearer.

For a fictional internal routing suggestion, the existing request record could show the proposed category and coordinator's correction. The team would not need a separate narrative for every suggestion.

For a fictional court-submission workflow, a routing record alone would be inadequate. The responsible lawyer needs a process appropriate to verifying the actual authorities, facts, and work being submitted.

The ABA's Formal Opinion 512 discusses review that depends on the tool and task. It applies the ABA Model Rules; it does not determine a particular jurisdiction's obligations.

Avoid creating another sensitive copy

Consider whether the record can point to the approved source instead of duplicating its contents. Use the firm's permitted storage, access, and retention practices.

Try the record on one normal case and one rejected output. Someone taking over the workflow should be able to understand what was accepted and what still needs attention.

Sources used

  1. NIST AI RMF 1.0, January 26, 2023. Used for defined human roles, oversight, and context-dependent configurations.
  2. NIST Generative AI Profile, July 26, 2024, NIST page updated April 8, 2026. Used for risk-proportionate evaluation, ground-truth comparison, and fact-checking.
  3. ABA Formal Opinion 512, July 29, 2024. Used for task-specific verification, professional judgment, and court-submission review.
Check Your AI Readiness

Take ten questions to identify strengths and gaps in how your legal team adopts and governs AI.