Select Page

Attorneys can do a lot with new generative AI tools beyond earn sanctions for faking citations. But despite the AI industry’s curiously apocalyptic advertising claims — like Anthropic’s World Cup ad that sort of implied AI would kill us all? — the fact of the matter is that large language models can’t do everything right now. Indeed, for a lot of tasks, they’re affirmatively worse than the technology and processes that we’ve refined over the past decade or so.

A new report from Casepoint titled “From AI Hype to AI Accountability: What You Can Learn From Legal and FOIA Teams About AI Modernization,” built on practitioner interviews with legal, records, and FOIA teams, covers a broad array of AI topics. One that jumped out is the admonition that lawyers need to understand that this new technology may excite the legal community, but when it comes to eDiscovery, the old ways are still the best.

“Privilege review is one example,” the report explains. “The commercial appeal is obvious: privilege review is expensive, time-consuming, and often pattern-driven.” The problem is that mistakes can be devastating and users all too often trust AI’s confident, mansplaining-as-a-service output without performing needed diligence. “Al may still have a role in privilege review, responsiveness calls, redactions, FOIA exemptions, legal advice, and production decisions,” the report continues, but teams “should know how output will be checked, what human remains accountable, what documentation will be retained, and whether the process can be explained later.”

Conveniently, we already have these processes for existing technologies:

GenAl also lacks the mature validation history that helped technology-assisted review gain acceptance. TAR became defensible because practitioners developed ways to explain and test the process, including recall, precision, seed sets, sampling, and quality-control protocols. GenAl review has not yet reached the same level of accepted process maturity.

As a user at a large federal civilian agency explained, “There’s no point using AI in the manner I described if we get to a production, go to court, and they say: this is not verifiable, this is not TAR 2.0, and I don’t have any of the statistics and measures I need.”

We’ve built technology to perform these tasks that doesn’t have the hankering to take black box flights of hallucinatory fancy and spent years making sure it had the reliability and accountability to hold up in court. The Casepoint interviewees describe hallucinations not as a bug to be patched but as a permanent “workflow reality.” TAR had plenty of flaws. Confidently inventing a case that does not exist was not among them. Vendors are building generative AI into their systems to handle what it’s suited to handle, and metric tons of digital ink have been spilled describing the efforts legal tech providers have made to build out genuine audit processes, but part of adopting AI is understanding its limits.

It’s been 14 years since Magistrate Judge Andrew Peck issued the first judicial opinion approving predictive coding, in Da Silva Moore v. Publicis Groupe. The crux of the case wasn’t “computers are good,” but that defensibility lives in the process. TAR didn’t become boringly trustworthy because the technology was magic. We just poured years into forging a system we could test, measure, and explain to a judge. Part of selling the courts on the rules was using (for the most part) rules-based technology.

But the selling point of the “agentic” era is that it comes up with its own rules to match the goal. That doesn’t mean it’s always wrong — the rules it comes up with may well be defensible — but it’s a different task to explain that to a judge.

Though so far, the courts don’t seem to think so. In Schulte v. LinkedIn, decided in the Northern District of California this June, Judge Eumi Lee declined to invent a special legal framework for generative AI in discovery and simply analyzed it under the principles courts already apply to TAR. At a high enough level this is true — it has to consistently display recall, precision, and a testable process. But it also hamstrings the technology a bit. It’s not quite like demanding that cars be built as mechanical horses, but some of the old processes just won’t get the most out of the new technologies.

The practical advice is simple: do not treat Al governance as policy language alone. Build it into the workflow. Legal teams need to define what Al may do, what information it may access, what records it creates, how output is reviewed, and how decisions are documented. Without that structure, organizations may not discover the accountability gap until a matter, FOIA request, production, court challenge, or oversight review forces the issue.

None of this means generative AI won’t take on more and more discovery tasks or that courts won’t eventually develop rules optimized to the technology… but it also means everyone needs to be a little honest about what it’s not capable of doing.


HeadshotJoe Patrice is a senior editor at Above the Law and co-host of Thinking Like A Lawyer. Feel free to email any tips, questions, or comments. Follow him on Twitter or Bluesky if you’re interested in law, politics, and a healthy dose of college sports news.

The post Lawyers Learning You Can’t Square Peg AI Into eDiscovery Round Holes appeared first on Above the Law.

Attorneys can do a lot with new generative AI tools beyond earn sanctions for faking citations. But despite the AI industry’s curiously apocalyptic advertising claims — like Anthropic’s World Cup ad that sort of implied AI would kill us all? — the fact of the matter is that large language models can’t do everything right now. Indeed, for a lot of tasks, they’re affirmatively worse than the technology and processes that we’ve refined over the past decade or so.

A new report from Casepoint titled “From AI Hype to AI Accountability: What You Can Learn From Legal and FOIA Teams About AI Modernization,” built on practitioner interviews with legal, records, and FOIA teams, covers a broad array of AI topics. One that jumped out is the admonition that lawyers need to understand that this new technology may excite the legal community, but when it comes to eDiscovery, the old ways are still the best.

“Privilege review is one example,” the report explains. “The commercial appeal is obvious: privilege review is expensive, time-consuming, and often pattern-driven.” The problem is that mistakes can be devastating and users all too often trust AI’s confident, mansplaining-as-a-service output without performing needed diligence. “Al may still have a role in privilege review, responsiveness calls, redactions, FOIA exemptions, legal advice, and production decisions,” the report continues, but teams “should know how output will be checked, what human remains accountable, what documentation will be retained, and whether the process can be explained later.”

Conveniently, we already have these processes for existing technologies:

GenAl also lacks the mature validation history that helped technology-assisted review gain acceptance. TAR became defensible because practitioners developed ways to explain and test the process, including recall, precision, seed sets, sampling, and quality-control protocols. GenAl review has not yet reached the same level of accepted process maturity.

As a user at a large federal civilian agency explained, “There’s no point using AI in the manner I described if we get to a production, go to court, and they say: this is not verifiable, this is not TAR 2.0, and I don’t have any of the statistics and measures I need.”

We’ve built technology to perform these tasks that doesn’t have the hankering to take black box flights of hallucinatory fancy and spent years making sure it had the reliability and accountability to hold up in court. The Casepoint interviewees describe hallucinations not as a bug to be patched but as a permanent “workflow reality.” TAR had plenty of flaws. Confidently inventing a case that does not exist was not among them. Vendors are building generative AI into their systems to handle what it’s suited to handle, and metric tons of digital ink have been spilled describing the efforts legal tech providers have made to build out genuine audit processes, but part of adopting AI is understanding its limits.

It’s been 14 years since Magistrate Judge Andrew Peck issued the first judicial opinion approving predictive coding, in Da Silva Moore v. Publicis Groupe. The crux of the case wasn’t “computers are good,” but that defensibility lives in the process. TAR didn’t become boringly trustworthy because the technology was magic. We just poured years into forging a system we could test, measure, and explain to a judge. Part of selling the courts on the rules was using (for the most part) rules-based technology.

But the selling point of the “agentic” era is that it comes up with its own rules to match the goal. That doesn’t mean it’s always wrong — the rules it comes up with may well be defensible — but it’s a different task to explain that to a judge.

Though so far, the courts don’t seem to think so. In Schulte v. LinkedIn, decided in the Northern District of California this June, Judge Eumi Lee declined to invent a special legal framework for generative AI in discovery and simply analyzed it under the principles courts already apply to TAR. At a high enough level this is true — it has to consistently display recall, precision, and a testable process. But it also hamstrings the technology a bit. It’s not quite like demanding that cars be built as mechanical horses, but some of the old processes just won’t get the most out of the new technologies.

The practical advice is simple: do not treat Al governance as policy language alone. Build it into the workflow. Legal teams need to define what Al may do, what information it may access, what records it creates, how output is reviewed, and how decisions are documented. Without that structure, organizations may not discover the accountability gap until a matter, FOIA request, production, court challenge, or oversight review forces the issue.

None of this means generative AI won’t take on more and more discovery tasks or that courts won’t eventually develop rules optimized to the technology… but it also means everyone needs to be a little honest about what it’s not capable of doing.


HeadshotJoe Patrice is a senior editor at Above the Law and co-host of Thinking Like A Lawyer. Feel free to email any tips, questions, or comments. Follow him on Twitter or Bluesky if you’re interested in law, politics, and a healthy dose of college sports news.

The post Lawyers Learning You Can’t Square Peg AI Into eDiscovery Round Holes appeared first on Above the Law.