When a law firm evaluates a legal AI tool, the questions are almost always about the things you can see in a demo. How fast is it? Do the citations look right? Can it draft something usable in under a minute? These are fair questions but not what I would lead with. I want to know if this tool will tell me when it doesn’t know the answer.
An LLM is a brain, and almost every legal AI tool on the market today is built on top of a general model trained on the open web. AI has a long memory and an earnest desire to please. The moment it can’t find an answer in the material you’ve provided, the model performs exactly as trained, reaching back through everything else it knows in effort to help. The open web is full of material not rooted in fact, or in other words made up, and AI doesn’t know the difference.
Context is always underestimated
Most of what you can pull for free online is primary source law. Legislation. Rules. Official missives from agencies like the IRS. Primary source law is an essential component of effective legal AI, but taken alone it’s still like trying to assemble a 1,000 piece puzzle without being able to look at the image on the box for context. Secondary source material — the treatises, the analysis, the practitioner guidance that explains how a rule actually works, why it’s important and what it relates to — is that bigger picture.
My go-to example is elections. Search “elections” in almost any state and you’ll get a mountain of provisions, and most of them have nothing to do about the rule that prompted your original search. Absent the larger context provided by authoritative secondary source content, a large language model has nothing to which it can anchor. The result is an answer unmoored from any verifiable body of information, such as a table of contents, drifting closer and closer towards potential hallucination.
Guardrails don’t stay put
A new round of testing should accompany every model update and feature release. Each iteration of an LLM reasons differently, and a guardrail that previously kept the system anchored to your vetted content may now have a loophole you never anticipated.
Ask the system to build a contract and lay out the appropriate clauses. You can watch to see if the eager-to-please model will out-think your guardrails in an effort to deliver an answer.
An automated test will more than likely pass a fault output so long as the citations provided appear genuine. A human reviewer needs to physically click the citation and verify that it does not answer the question in order to catch the problem.
Saying “I don’t know” is harder than it sounds
Nobody likes admitting they don’t have all the answers. Teaching a model to say “I don’t know” is both a matter of training and a statement of purpose. While it’s true that a tool which sometimes comes back with “I don’t have a source for this” may seem less impactful in a sales demo, practicing attorneys will actually recognize it as solid ground.
When I practiced law, my greatest fear was missing something critical. If there’s no reputable authority that can be cited, that would have been genuinely useful for me to know. It empowers me to begin taking steps in the right direction. On the other hand, a paragraph buoyed by false confidence and built from dubious sourcing takes much longer to untangle.
Remember that transparency is a vital component of trust. Accuracy isn’t just about pinpointing the correct answer. It’s also about being able to acknowledge when there is no answer to be found. Establishing that level of visibility as table stakes is essential to building a strong future for legal AI.
Nicole Stone is Director of AI & Agentic Solutions Product Management at Wolters Kluwer Legal & Regulatory U.S., where she leads product strategy and development for digital legal content and technology solutions. With over 22 years of experience in legal technology and a background as a practicing attorney, she focuses on integrating emerging technologies, including generative AI, into products that serve legal professionals.
The post The One Thing You Didn’t Think To Look For In Legal AI appeared first on Above the Law.
When a law firm evaluates a legal AI tool, the questions are almost always about the things you can see in a demo. How fast is it? Do the citations look right? Can it draft something usable in under a minute? These are fair questions but not what I would lead with. I want to know if this tool will tell me when it doesn’t know the answer.
An LLM is a brain, and almost every legal AI tool on the market today is built on top of a general model trained on the open web. AI has a long memory and an earnest desire to please. The moment it can’t find an answer in the material you’ve provided, the model performs exactly as trained, reaching back through everything else it knows in effort to help. The open web is full of material not rooted in fact, or in other words made up, and AI doesn’t know the difference.
Context is always underestimated
Most of what you can pull for free online is primary source law. Legislation. Rules. Official missives from agencies like the IRS. Primary source law is an essential component of effective legal AI, but taken alone it’s still like trying to assemble a 1,000 piece puzzle without being able to look at the image on the box for context. Secondary source material — the treatises, the analysis, the practitioner guidance that explains how a rule actually works, why it’s important and what it relates to — is that bigger picture.
My go-to example is elections. Search “elections” in almost any state and you’ll get a mountain of provisions, and most of them have nothing to do about the rule that prompted your original search. Absent the larger context provided by authoritative secondary source content, a large language model has nothing to which it can anchor. The result is an answer unmoored from any verifiable body of information, such as a table of contents, drifting closer and closer towards potential hallucination.
Guardrails don’t stay put
A new round of testing should accompany every model update and feature release. Each iteration of an LLM reasons differently, and a guardrail that previously kept the system anchored to your vetted content may now have a loophole you never anticipated.
Ask the system to build a contract and lay out the appropriate clauses. You can watch to see if the eager-to-please model will out-think your guardrails in an effort to deliver an answer.
An automated test will more than likely pass a fault output so long as the citations provided appear genuine. A human reviewer needs to physically click the citation and verify that it does not answer the question in order to catch the problem.
Saying “I don’t know” is harder than it sounds
Nobody likes admitting they don’t have all the answers. Teaching a model to say “I don’t know” is both a matter of training and a statement of purpose. While it’s true that a tool which sometimes comes back with “I don’t have a source for this” may seem less impactful in a sales demo, practicing attorneys will actually recognize it as solid ground.
When I practiced law, my greatest fear was missing something critical. If there’s no reputable authority that can be cited, that would have been genuinely useful for me to know. It empowers me to begin taking steps in the right direction. On the other hand, a paragraph buoyed by false confidence and built from dubious sourcing takes much longer to untangle.
Remember that transparency is a vital component of trust. Accuracy isn’t just about pinpointing the correct answer. It’s also about being able to acknowledge when there is no answer to be found. Establishing that level of visibility as table stakes is essential to building a strong future for legal AI.
Nicole Stone is Director of AI & Agentic Solutions Product Management at Wolters Kluwer Legal & Regulatory U.S., where she leads product strategy and development for digital legal content and technology solutions. With over 22 years of experience in legal technology and a background as a practicing attorney, she focuses on integrating emerging technologies, including generative AI, into products that serve legal professionals.
The post The One Thing You Didn’t Think To Look For In Legal AI appeared first on Above the Law.

