
Agentic AI can move legal work forward, but only when authority, evidence, and accountability are designed into the workflow.
Most conversations about artificial intelligence begin with capability. Can the system summarize a deposition, classify a document, extract a deadline, draft a response, or identify privileged material? Those questions matter, but capability alone should not determine readiness for high-risk legal and litigation workflows.
The more important questions that must be asked are operational. What happens when the system is wrong? Who notices? Can the error be contained or reversed? Does the reviewer have enough context to challenge it? Who is accountable when a bad recommendation becomes legal action?
Legal workflows do not treat mistakes as isolated events. A misclassified request can distort collection scope. A missed deadline can become a procedural failure. A faulty privilege decision or redaction can expose protected information. In an agentic system, the larger risk is that a bad answer becomes an action and propagates through the workflow.
Courts are recording the cost. Damien Charlotin’s database tracked roughly 200 decisions involving AI-fabricated citations in mid-2025 and more than 1,500 by June 2026, including over 1,100 in the United States. Penalties have ranged from the $5,000 Mata v. Avianca fine to $15,000 per attorney in Sixth Circuit sanctions and an interim license suspension in Nebraska. (Norton Rose Fulbright)
The standard cannot be zero error, because neither people nor software can meet it. The practical standard is controlled failure: errors should be visible, contained, reversible, and prevented from silently becoming consequential actions. That shifts the design goal from impressive demonstrations to reliable performance under real legal and operational pressure.
The most dangerous AI errors look reasonable
Cory Doctorow describes discernment as a prerequisite for effective AI use. How, he asks, can someone fact-check a system that is explaining something they do not already understand?
Doctorow’s point is not that AI is useless. Its value depends on the user’s ability to evaluate it. An expert may recognize a misleading answer, faulty assumption, or missing source. Someone without that expertise may mistake fluency for accuracy. (Medium)
Stanford RegLab tested 2023-era general-purpose models against more than 800,000 verifiable legal questions and found hallucination rates from 58 percent for GPT-4 to 88 percent for Llama 2. Legal research tools performed better, but retrieval-augmented generation narrowed rather than closed the gap. (Stanford RegLab)
These errors rarely announce themselves. They appear as plausible deadlines, reasonable interpretations, professionally drafted documents, or confident privilege recommendations. The Stanford typology distinguishes fabrication from misgrounding, where a system cites a real source for a proposition the source does not support. The citation checks out. The reasoning does not.
Legal AI should therefore help users exercise discernment, not merely deliver conclusions. For litigation and sensitive document workflows, the system should show:
- Sources supporting material conclusions and the precise language behind extracted obligations
- Documents considered, excluded, or unavailable
- Instructions and policies applied
- Conflicting information, uncertainty, and reasons for escalation
- Changes made between versions
The goal is not to make AI sound more certain. It is to make the basis of its work inspectable. A citation is not decoration. In a consequential workflow, it is part of the control system.
Human in the loop is not enough
The standard response to concerns about legal AI is that a human will remain in the loop. The phrase reassures, but says little about how the system operates.
Doctorow distinguishes a centaur, where technology augments a skilled person, from a reverse centaur, where the person assists the machine. In the first, the human directs the work. In the second, the machine produces at scale while a person is expected to catch mistakes. (Pluralistic)
Imagine AI preparing hundreds of documents for production: identifying responsive material, proposing redactions, applying naming conventions, and assembling the package. A reviewer is asked to approve it. Technically, a human is in the loop. Operationally, the reviewer may face more material than anyone can evaluate carefully. The system has placed a human signature at the end of a machine-driven process.
People are poorly suited to perfect vigilance over large volumes of plausible machine output. Automation-bias research distinguishes errors of commission, when an operator follows a recommendation despite contrary evidence, from errors of omission, when the operator misses what the system failed to flag. A second reviewer did not reliably eliminate omission errors. (Pluralistic; PubMed)
A 2024 CSET review concluded that human-in-the-loop design cannot, by itself, prevent all accidents or errors. Meaningful review requires time, expertise, evidence, an explanation of system actions, authority to reject or stop the workflow, and a manageable scope focused on exceptions. (CSET)
Human oversight should not require legal professionals to reproduce the machine’s analysis by hand or turn them into rubber stamps. Automation should focus their attention where legal judgment is most valuable.
Do not turn legal professionals into accountability sinks
Doctorow, drawing on Dan Davies, calls this an accountability sink: responsibility without meaningful control. A person may be named as reviewer, but workload or interface prevents real judgment. When something fails, the organization points to whoever clicked approve even though the system dictated the result. (Pluralistic)
The conditions already exist. In the 8am 2026 Legal Industry Report, 43 percent of more than 1,300 legal professionals reported no formal AI policy and no plans for one. Only 9 percent had an enforced written policy, while 54 percent reported no responsible-use training or plans to provide it. (American Bar Association) Adoption is outpacing defensible controls.
A defensible system must distinguish the agent’s authority, the legal professional’s judgment, the administrator’s configuration role, the business owner’s process responsibility, the vendor’s obligations, and the organization’s monitoring duties. These cannot be collapsed into saying a human approved it.
For each consequential action, the audit record should show the inputs; the agent, model, rule, and instruction involved; supporting sources and uncertainty; the action taken; the reviewer; changes after approval; reversibility; and responsibility for correction. Audit trails should preserve reasoning context, not merely activity.
Agentic does not have to mean unconstrained
Agentic AI is often defined as AI that plans and executes work independently. In legal workflows, useful agency should mean bounded authority within defined permissions, approved instructions, accessible data sources, and escalation rules.
For a litigation document request, an agent might capture and classify it, identify parties and dates, route tasks, send reminders, assemble documents, flag sensitive information, propose redactions, draft a response, and record actions.
It should not determine final legal scope, resolve disputed interpretation, waive privilege, release sensitive information, alter preservation duties, submit a production without approval, or conceal uncertainty. The issue is not whether AI acts, but whether it acts within an enforceable authority model. Permissions belong in architecture, not merely in a prompt.
Gartner predicts that more than 40 percent of agentic AI projects will be canceled by the end of 2027 because of escalating costs, unclear business value, and inadequate risk controls. Model capability is not on the list. (Gartner)
Seven principles for workflows that cannot afford silent failure
1. Require evidence before eloquence
A polished answer without support is a liability. Legal AI should lead with sources, document references, and extracted language. Generated narrative should explain evidence, not substitute for it.
2. Separate recommendations from actions
The ability to recommend a step does not justify authority to execute it. Systems should distinguish actions they may perform automatically, actions requiring approval, and actions they must never perform.
3. Place review points around legal judgment
Review everywhere creates fatigue without improving control. Concentrate it on scope, privilege, privacy, disclosure, legal interpretation, exceptions, and irreversible external actions.
4. Design for refusal and escalation
An agent should be able to report missing information, conflicting evidence, or the limit of its authority. A system that always answers is less trustworthy than one that knows when to stop.
5. Make recovery a standard workflow
Errors will occur. Systems should support versioning, rollback, quarantine, reprocessing, correction, and notification so recovery does not depend on email, spreadsheets, and memory.
6. Test the workflow, not only the model
Step-level errors compound across agents, integrations, approvals, and downstream actions. Test incomplete requests, conflicting instructions, unusual documents, bad metadata, privilege edge cases, unavailable sources, and attempts to exceed permissions. The unit of evaluation is the complete legal process.
7. Preserve the ability to leave
Doctorow’s enshittification describes how platforms deteriorate after users become locked in. Legal departments should retain data, instructions, workflow records, audit logs, and decision histories; know which models are used and how they may change; and ensure workflows can continue if a model or vendor is replaced. Portability is a governance control. (Pluralistic)
A practical path to cautious adoption
Cautious adoption does not mean waiting for infallible AI. It means introducing agency in proportion to the organization’s ability to observe, govern, and reverse the system’s actions.
Begin with a narrow, repeatable workflow with clear inputs, owners, and outcomes. Litigation intake, document request coordination, production preparation, and deadline extraction are better starting points than open-ended legal analysis.
Progress through three stages. In observation mode, AI analyzes work without affecting the live process. In assisted execution, it prepares work, identifies exceptions, and proposes actions while people retain authority. Bounded autonomy comes last, limited to low-risk, reversible, well-understood tasks. Expand authority only after demonstrating that failures can be detected and contained.
Measurement is often skipped. Thomson Reuters found that 82 percent of legal departments do not measure AI return on investment or do not know whether they do. Beyond completed tasks, measure exceptions detected, errors contained, reviewer attention preserved, rework reduced, traceable decisions, correct escalations, successful reversals, and documented responsibility. (Modern Counsel)
The purpose of agentic AI is not to remove people from legal workflows. It is to remove unnecessary coordination, repetition, and administrative effort so people can concentrate on decisions for which legal judgment is indispensable.
Trust must be engineered
Doctorow’s critiques offer a corrective to the prevailing AI narrative. The central question is not whether AI is powerful. It is who the system serves, what incentives shape it, who bears the cost of mistakes, and whether the people supposedly overseeing it have meaningful control.
For legal technology professionals, this is not an argument against agentic AI. It is an argument for a more demanding standard. A trustworthy system should not ask users to believe it is correct; it should make its work inspectable. It should preserve genuine human authority, treat uncertainty as a reason to pause, and measure autonomy by how safely it operates within its assigned role.
In workflows that do not forgive mistakes, intelligence is not enough. The system must know its limits, show its work, preserve accountability, and stop before uncertainty becomes action.



