A feature list makes almost any workflow look possible. It says little about whether the work will survive contact with real inputs, real permissions, and a busy reviewer.
Choose a complete piece of work instead: one trigger, a known set of inputs, an observable output, a named owner, and a destination. Once those facts are on paper, the useful question becomes specific. Where could AI help without obscuring responsibility or making a bad result harder to contain?
This article provides general operational education. It is not legal advice or a compliance determination. No attorney-client relationship exists, and none of its protections apply.
The unit of analysis is the work
AI can produce a plausible result even when the surrounding process is poorly defined. A polished answer can therefore mask a weak intake, an unknown source, or a missing owner.
The NIST AI Risk Management Framework offers a useful operating principle. Map the context, measure performance, manage risk, and govern the work throughout. NIST describes the framework as voluntary and use-case agnostic. Its current page says version 1.0 is being revised as part of the White House AI Action Plan.
For a legal team, context begins with the work itself. A contract summary used for triage differs from a clause recommendation sent to a business owner. The same model may touch both, but the decisions and consequences differ.
Six questions that expose weak candidates
These questions do not create a maturity score. One serious boundary can outweigh several attractive features.
1. What outcome does the workflow produce?
Write the outcome in observable terms. “Help with contracts” is too broad. “Extract renewal dates into a draft tracker for human confirmation” is specific enough to test.
Record the current method as a baseline. Measure quality, time, rework, and exceptions before claiming improvement. A faster first draft may still create more review work.
2. Is the work repeatable enough to test?
Strong candidates have recurring inputs and a stable output shape. They also have enough examples to test normal cases and exceptions.
Variation needs a defined test boundary. If each matter requires a new objective and new professional judgment, narrow the proposed AI task before testing it.
3. Can you bound the AI's role?
Define what the system may read, produce, and change. Prefer a draft, suggestion, or classification that remains inside an approved process.
A narrow role gives the team specific behavior to test. Check that the configured tools and permissions enforce the intended boundary.
4. Can a qualified person verify the result?
Name the reviewer and the evidence they will see. “Human in the loop” is not a review design.
The reviewer needs time, relevant source material, and authority to reject the output. If nobody can tell whether the answer is right, the workflow is not ready.
The ABA's Formal Opinion 512 addresses lawyers using generative AI. It says review depends on the tool and task. It also says lawyers cannot surrender work that calls for professional judgment.
The opinion concerns professional duties under the ABA Model Rules. It does not decide a reader's duties in a specific jurisdiction or situation.
5. Are the information and tools approved?
List the information the workflow will expose. Include personal, confidential, privileged, proprietary, and regulated information.
Then confirm the approved tool, account, settings, contract terms, retention, and access controls. A good task in an unapproved environment is still a no-go.
Formal Opinion 512 also warns that generative AI can raise confidentiality, intellectual property, and security issues. Those concerns require tool-specific and matter-specific review.
6. Can the team contain a bad result?
Ask what happens when the output is wrong. Consider who sees it, what system receives it, and whether the action can be reversed.
A draft saved for review is easier to contain than an external message. A suggestion is easier to contain than a filed document or binding decision.
What the screen changes
These examples are fictional. They illustrate the screen and do not describe client work.
Candidate: outside-counsel invoice intake
The AI extracts matter number, firm, date, and line-item categories into a draft record. A legal operations analyst checks the invoice and source fields before posting.
The system cannot approve payment or decide whether a charge is reasonable. Exceptions route to the existing review process.
This is a plausible pilot because the output is structured, reviewable, and reversible before posting. It still needs approved data handling and a measured baseline.
Conditional candidate: NDA comparison
The AI compares an incoming NDA with an approved playbook. It identifies clauses for review and links each flag to the source text.
A qualified reviewer checks every flag and makes the legal judgment. The system does not negotiate, send language, or decide acceptable risk.
This candidate depends on a current playbook, reliable source links, and a reviewer who can assess the clause. Without those controls, it needs redesign.
No-AI decision: settlement authority
An AI system should not decide whether an organization accepts a settlement. The decision is externally consequential, fact-specific, and difficult to reverse.
AI may assist a separate, bounded task such as organizing approved facts. The responsible lawyer and client still make the decision.
Choose a route, not a score
The screen should produce one of three routes.
- Bounded pilot: the task, inputs, output, reviewer, and containment plan are clear.
- Redesign first: the use case may help, but a data, source, ownership, or review boundary remains unresolved.
- No AI for this decision: the proposed role would replace professional judgment or create an unacceptable consequence.
The no-AI route is useful. It protects attention for workflows where testing can produce evidence.
Where the screen tends to fail
Expect the screen to fail late rather than early. Most teams can name an outcome, find enough examples, and bound the AI's role. The hard questions are the fourth and fifth: what the reviewer will actually look at, and whether the folder the tool needs is approved. Approving the tool is not approving the folder.
Those 2 answers decide the route. Missing reviewer evidence makes the candidate a redesign, however good the demo. An unresolved information boundary makes it a no-go until the right person has decided.
Record the answers for 1 workflow on the workflow-selection worksheet, and keep the version that said no. A dated no-AI decision is easier to defend, and easier to revisit, than a pilot nobody can explain.
If you are unsure which workflow to screen first, start with the free AI readiness assessment. Its 10 questions return your priority gaps and a recommended next step. Begin with the workflow behind the widest gap.
Sources used
- NIST AI Risk Management Framework, accessed September 18, 2026. Used for voluntary, use-case-specific risk management and the stated White House AI Action Plan revision driver.
- NIST AI RMF 1.0, January 26, 2023. Used for the GOVERN, MAP, MEASURE, and MANAGE operating frame.
- ABA Formal Opinion 512, July 29, 2024. Used for task-specific review, professional judgment, and confidentiality limits.




