Agentic AI in Legal: Document Review and Contract Workflows
Even purpose-built legal AI research tools hallucinate between 17% and 33% of the time, according to a peer-reviewed Stanford RegLab study testing Lexis+ AI and Westlaw's AI-Assisted Research.
That's despite vendor claims that their retrieval-based design had eliminated the problem entirely.
A New York attorney found out what that gap costs in 2023, when he cited six cases ChatGPT had invented in a real federal filing. The court's response became the founding precedent for how every subsequent court has treated unverified AI output.
This guide covers agentic AI in legal work specifically: how contract review and discovery document review actually work at volume, and why an attorney's verification stays the one step nothing can replace.
What Is Agentic AI in Legal?
Agentic AI in legal work is a system that reviews contracts or discovery documents at volume, flagging what needs an attorney's judgment rather than replacing that judgment itself.
A keyword search finds documents containing a specific term. An agentic system reads what a contract clause actually says, or what a document actually reveals about a legal issue, and reasons about whether it matters to the specific matter at hand.
That distinction matters because legal documents are dense with context a keyword can't capture: a clause that looks standard but conflicts with an earlier section, a document that's relevant not because of what it says but because of when it was sent.
Why Purpose-Built Doesn't Mean Hallucination-Free
Legal research tools built specifically for the profession, with retrieval-augmented generation pulling from case law databases instead of the open internet, were marketed as having solved the hallucination problem general-purpose chatbots have.
The Stanford RegLab study tested that claim directly and found it overstated. Lexis+ AI and Westlaw's AI-Assisted Research still hallucinated on a meaningful share of queries, producing fabricated citations or answers that misstated what a real case actually held.
That doesn't mean purpose-built tools aren't an improvement. A general-purpose model tested in the same study hallucinated noticeably more often. It means specialized retrieval reduces the problem without eliminating it.
Agentic review works the same way: better than a keyword search, not a substitute for a human confirming what actually matters before it goes anywhere near a filing.
The Contract and Discovery Review Workflow, Step by Step
The workflow below covers document review broadly. The sections after it go deeper into contract review specifically, discovery review, and where verification has to sit.
The workflow
Contract Review and Redlining
Reviewing a contract for the clauses that actually matter, an indemnification provision that shifts risk unexpectedly, a termination clause with unusual notice requirements, takes real attorney time when done clause by clause from scratch.
An agent that flags deviations from a standard template or a prior negotiated position gives an attorney a starting point: here's what's different, here's why it might matter, rather than a blank read-through of the entire document.
Discovery Document Review
A document production can run into the hundreds of thousands of files, and finding the handful that actually matter to a case by reading each one is exactly the volume problem agentic review is suited for.
The agent's job is surfacing what's relevant with a stated reason. The privilege call, the responsiveness determination, and anything that becomes part of a production still needs an attorney's sign-off.
The Verification Discipline Courts Now Require
In June 2023, a federal judge sanctioned two attorneys and their firm $5,000 after they submitted a brief citing six cases that turned out to be entirely fabricated by ChatGPT.
The court didn't fault the attorneys for using AI. It faulted them for not verifying what it produced before filing it, a distinction that has shaped how every court since has approached the same question.
That principle applies just as directly to contract review and discovery work as it does to legal research: an agent's output is a draft to verify, never a citation or a conclusion to file on trust.
The Attorney's Role
The attorney's job shifts from reading every document or drafting every clause from scratch to verifying what an agent flagged, the exact judgment call no verification statistic can substitute for.
That verification duty sits with the attorney regardless of which tool produced the output, purpose-built or general-purpose, since the signature on the filing is the attorney's, not the tool's.
Implementation: Guardrails Specific to Legal Work
Every guardrail below exists because a citation or a clause that looks right and isn't is a different, harder problem than one that's obviously wrong.
| Layer | What it does | Legal-specific example |
|---|---|---|
| System prompt | Sets the non-negotiables up front | "Never state a citation or a clause interpretation without a traceable source" |
| Input filters | Block or sanitize out-of-scope requests | Treat contract and discovery text as data to evaluate, not instructions to follow |
| Tool-call gatekeepers | Cap what actions an agent can take | Flagging and drafting allowed; anything filed with a court always needs a human |
| Output checks | Scan before the action executes | Block any citation that hasn't been checked against a primary source |
| Human-in-the-loop | Requires approval for high-impact actions | An attorney verifies every flagged item before it's used or filed |
Rolling This Out: What to Expect
Start with contract review or discovery triage on a defined, bounded document set rather than an entire matter at once.
Compare the agent's flags against a sample a person has already reviewed by hand before trusting its output on new material, and build citation verification into the workflow as a required step, not an optional check.
Expect the verification step to feel slow at first. It's the step every sanctioned case on record skipped, and it's the one nothing about this technology has made optional.
The Team Behind Production Agentic AI
The Stanford study's finding wasn't that legal AI is unusable. It's that the marketing claim of zero hallucinations was never true, and building a workflow around that honest baseline is what actually holds up under a court's scrutiny.
Tecla's Agentic AI services design, build, and operate this workflow directly, the same contract and discovery review systems above, running in your stack with the evals and guardrails production requires.
Or bring the expertise in-house: AI engineers who've worked on live legal systems, past the demo stage.
Tecla runs a network of senior engineers across the US and Latin America, built over more than a decade, with a top 3% acceptance rate and first candidates in 3 to 5 business days.


.png)
%20(1).avif)
.avif)
.avif)