A practical AI workflow guide for small law firms

Suppose a staff member sends an internal request in a free-text message. Someone has to understand it, choose a category, and route it to the right colleague. The work is modest, recurring, and easy to describe. That makes it a better learning candidate than a vague ambition to “use AI for legal work.”

This guide follows that fictional request from selection through a test decision. The proposed assistant never answers a client, makes a legal judgment, or changes a matter record. Its narrow job is to prepare a draft category and routing note for a person to confirm.

The example illustrates a method, not client results or a ready-to-deploy system. A firm should compare the workflow with a simpler form or routing rule and preserve a manual route. It should also decide in advance what evidence would justify continued use.

This article provides general operational education. It is not legal advice or a compliance determination. No attorney-client relationship exists, and none of its protections apply.

Write the job so another person can recognize it

Write the job in one sentence. Include the trigger, output, and receiving person.

For our fictional firm: “When a staff member submits an internal administrative request, prepare a draft category and routing note for the office coordinator.”

This leaves legal advice, urgent client matters, and decisions about accepting work outside the scope. The proposed assistant can organize an internal request. It cannot commit the firm, respond to a client, or change a matter record.

Compare that design with the current process. A clearer request form and simple routing rules might solve the problem without AI. If requests already arrive in a consistent format, test that simpler option first.

AI becomes a candidate when varied wording creates useful work for a language model. That is a hypothesis to test. The model's ability to classify a demonstration does not establish that the complete process works.

Use the workflow-selection worksheet to record the choice. If the outcome, data permission, or review method remains unclear, choose redesign before choosing a product.

Set a boundary the team can recognize

An instruction such as “handle routine requests” leaves too much interpretation to the operator and system. Describe included inputs and excluded cases separately.

Our fictional test includes requests for meeting-room setup, approved equipment, and internal administrative assistance. It uses synthetic messages without client information.

An apparent client matter, a legal deadline, or a message with sensitive attachments goes to the designated person through an approved process. The assistant does not decide the legal significance of the request.

Define what happens before information reaches the AI. If the input channel can receive prohibited material, an instruction inside the prompt does not prevent disclosure. The intake design needs a permitted way to separate that material first.

Document the tool, account, permitted data, and destination. Have the appropriate firm decision-maker resolve information-handling questions before testing with real material.

Assign an owner who can operate the process

The office manager owns this fictional workflow. The coordinator checks proposed routing. The administrator manages the approved account and configuration.

These are role assignments for the example, not a staffing recommendation. One person could hold more than one role if the firm's operating rules permit it.

The owner keeps the instructions current, reviews exceptions, and coordinates changes. They also know who can pause the workflow and how pending requests will be handled.

Specify a backup reviewer. If neither reviewer is available, requests remain queued or follow the existing manual route. Decide how the receiving team learns about delays.

The ownership article explains the division of responsibilities. Its companion card includes an absence test and a restart decision.

Place review before the consequential action

In this design, the AI proposes a category and a routing note. The coordinator compares them with the original request before sending work to another person.

The reviewer sees the full approved input, the proposed category, and the routing rules. They can correct the category, request clarification, or reject the output.

Give ambiguous requests a route that does not require guessing. A message containing “urgent” may concern a meeting room or a client deadline. The test should check both.

Later sampling serves another purpose: finding patterns in misses, delays, and overrides. It does not replace the checkpoint before a consequential action.

Review effort should fit the actual work. A private brainstorming aid may need a different process from a filing or an external message. Your firm's rules and qualified professionals determine legal obligations.

For the checkpoint design, see Where should human review sit?. For recordkeeping scope, see Does every AI task need an approval record?.

Write the test before running it

Choose acceptance conditions that someone can observe. Avoid a condition such as “the AI should be accurate.” It does not identify the output, allowed variation, or handling of uncertainty.

For the fictional routing test, proposed conditions are:

  • The output uses an allowed category or explicitly routes the request for clarification.
  • The routing note contains no invented facts about the requester or their circumstances.
  • Excluded requests reach the designated human route.
  • No request is forwarded until the coordinator accepts it.

These conditions belong to this example. The firm must decide whether they are sufficient for its actual workflow.

Prepare ordinary inputs and exceptions. Include incomplete messages, contradictory instructions, duplicate requests, a changed routing rule, and an unavailable reviewer. Test the boundary as well as output quality.

The test must also check what a connected tool can actually do. A prompt asking the assistant not to send messages is insufficient if an unintended send action remains available.

Record which cases were tested and which remain untested. A handful of successful examples supports a narrow rehearsal result. It does not establish the frequency of rare failures.

Compare the complete process

Compare the proposed workflow with a reasonable alternative on similar work. Use the current manual process when it exists. Include the simpler form or routing-rule option if it could solve the same problem.

Count active time spent preparing inputs, reviewing, correcting, and handing off the result. Separately record elapsed time, including queues and waiting. The two measurements answer different questions.

Keep quality visible for each completed unit. A faster average can hide an unacceptable mistake. Record rejected outputs, uncertain cases, and work that later needed repair.

For a newly established firm, the comparison may begin with rehearsal. Describe that as readiness evidence. It is not measured savings from a history of live operations.

The measurement article includes a synthetic calculation and an evidence card. Any estimated value still depends on how the team uses released capacity and what the workflow costs to maintain.

Decide what the evidence permits

Imagine that the routing test correctly handles ordinary requests but misroutes the two “urgent” examples. The team changes the boundary so those requests always reach the coordinator directly.

That is a redesign decision. The team should retest the changed behavior and check whether the narrower workflow still saves useful work.

If the form-and-rules option performs adequately with less effort, use it. If the AI-assisted version meets the criteria and offers a useful benefit, consider a bounded next stage.

Neither result requires a wider rollout. Permission to test one administrative workflow does not authorize client work, new connections, or expanded access.

Write down the decision, its evidence, the responsible person, and the next condition for review. “Continue” should name the scope that may continue.

Prepare for a change or interruption

After a successful test, rehearse the fallback. Have someone other than the builder process an approved example using the instructions.

Check that they can find the input, explain the acceptance criteria, reject a result, and recover a queued item. If they need the builder at every step, the handoff is incomplete.

Reopen testing when a change affects the task, model, source material, permissions, reviewer, or output destination. Preserve the previous version and its evidence.

For a pause, distinguish work already completed from work still pending. Turning off the AI does not reverse messages or records already sent downstream.

These operating choices fit the NIST AI Risk Management Framework, a voluntary framework for managing AI risks. It supports considering context throughout the system's life. It does not approve a particular workflow or replace the firm's obligations.

What the firm holds at the end

At the end of this method the firm holds 4 things it did not have before:

  1. a 1-sentence job description for the workflow;
  2. a boundary that names the excluded requests;
  3. a test record with the cases that failed; and
  4. a written decision that says what may continue and who can stop it.

None of the 4 depends on the AI working. If the form-and-rules alternative won, the firm keeps the same 4 records with a simpler tool inside them. That is the reason to start with a modest administrative request rather than client work. The method is the asset, and the assistant is 1 candidate it evaluated.

The first of the 4 is the workflow-selection worksheet. When the firm has a bounded workflow, it may want the assistant configured in its own environment, on client-approved tools. That is what Ortaire's Configure AI Agents work does.

Sources and further reading

Check Your AI Readiness

Take ten questions to identify strengths and gaps in how your legal team adopts and governs AI.