What AI Actually Does When You Ask It a Legal Question

What AI Actually Does When You Ask It a Legal Question

Abstract

When you ask AI a legal question, it predicts the next likely word based on patterns in its training data. It does not look up your state’s law. That gap between “sounds right” and “is right” has already cost real people real money, including a security deposit case where the wrong appeal deadline meant a lost claim. Here is how to protect yourself.

Dilemma

You typed your question into a chatbot because you needed a fast, clear answer, and you got one. It read like it came from someone who knew exactly what they were talking about. No hedging, no “consult a professional” disclaimer buried at the bottom, just a confident answer to your landlord dispute, your custody question, or your small claims filing.

Here is the problem: that confidence has nothing to do with whether the answer is correct. AI chatbots are built to sound sure of themselves. Whether the underlying legal information is accurate for your state, your situation, and your deadline is a completely separate question, and one the chatbot is not designed to flag for you.

This is not a small technical footnote. People use AI for legal questions constantly now, asking about lease terms before signing, appeal deadlines before a hearing, or custody rules before a court date. In every one of those moments, the person asking needs the answer to be right, not just readable. Understanding what is actually happening inside the chatbot when it answers you is the first step to protecting yourself.

How AI Actually Generates Answers

An AI chatbot like ChatGPT, Gemini, or Claude is a large language model. A large language model learns by reading enormous amounts of text, including websites, books, and legal documents, and studying the statistical patterns in how words follow other words.

When you ask a legal question, the model does not search a legal database and pull out the correct statute. It generates a response one word at a time, each word chosen because it is statistically likely to come next given everything before it. This is why AI answers read so smoothly. The model is optimized to produce fluent, natural-sounding text, not to verify that a cited law or case actually exists.

✽ Some AI tools now connect to live web search or legal databases, which can improve accuracy for basic questions. Even then, the model still decides what to include and how to phrase it, and it can still blend real information with invented details in a single answer.

Think about the difference between a search engine and a language model this way. A search engine finds an existing document and shows it to you. A language model writes a new document, word by word, based on patterns it learned. When you ask a search engine for a case name, it either finds a real one or returns nothing. When you ask a language model for a case name, it produces the most statistically plausible-sounding case name, whether or not that case actually exists. Both can look identical on your screen. Only one of them is reliably tied to reality.

This distinction matters most when a legal answer needs to include a specific detail: a filing deadline, a statute number, a court name, or a case citation. Those are exactly the details a language model is most likely to get wrong, because they require precise recall rather than general pattern-matching, and precise recall is not what the underlying technology is built to do.

Why It Sounds So Confident

Language models are trained to produce the response a human rater would rate highly. Confident, direct answers score better in training than answers full of qualifications and uncertainty. The result is a system that defaults to sounding sure, even when the underlying information is thin, outdated, or wrong.

This is not the model lying to you. It has no concept of “I don’t actually know this.” It generates the most probable-sounding response to your question, and a probable-sounding legal answer usually reads like a confident one. You are left holding an answer that sounds like it came from a lawyer, with none of the accountability a real lawyer carries.

A real attorney who is not sure of a state’s deadline will tell you so, and will go check before advising you. A chatbot has no equivalent instinct. It does not experience uncertainty the way a person does, so it cannot reliably signal uncertainty to you unless it has been specifically designed and tested to do that. Most consumer-facing chatbots have not been built or evaluated with that standard in mind, which is exactly why the confident, wrong answer is so common and so hard to spot from the outside.

What Hallucination Means in Practice

“Hallucination” is the term for when an AI model generates information that sounds real but is not, including invented statutes, fake case citations, or legal rules that do not exist in the form the AI describes. In legal contexts, this has moved from theoretical risk to documented, repeated fact.

The clearest example is Mata v. Avianca, a 2023 federal case in New York. Attorney Steven Schwartz used ChatGPT to research a personal injury claim and submitted a filing citing several cases as precedent. The airline’s lawyers could not locate the cases, and the judge ultimately confirmed the citations were fabricated. Schwartz later admitted he had not realized ChatGPT was a text-generation tool rather than a legal search engine.

That was not an isolated incident. In 2025, a California appeals court fined an attorney $10,000 after finding that most of the quotes in his filing were invented by ChatGPT, and published the ruling as a warning to the profession. A Utah attorney was separately sanctioned for citing a case that does not exist. By May 2026, a federal judge in Oregon issued a combined $110,000 penalty against two lawyers over fabricated citations and invented quotations, the largest fine of its kind on record. According to attorneys tracking the problem, close to 900 AI hallucinations have now been documented in US court filings since 2023, and the pace shows no sign of slowing.

These are trained attorneys, with law degrees and bar licenses, who trusted an AI-generated citation without checking it first. If it can happen to them, it can happen to you when you ask AI whether your landlord followed the right eviction procedure or whether your appeal deadline has already passed.

The court cases get attention because a judge caught the error and put it on the public record. For every fabricated citation caught by a judge, an unknown number of similar answers go unchecked in ordinary consumer situations: a renter who accepts a wrong move-out date, a small claims plaintiff who misses a filing window, a parent who acts on invented custody guidance. There is no judge reviewing those conversations. You are the only check in the system, which is exactly why verifying the answer yourself matters so much.

Three Legal Question Types and How AI Handles Each

Not all legal questions carry equal risk. It helps to know which category yours falls into before you decide how much to trust the answer.

  • Factual questions. “What does ‘eviction’ mean?” or “What is a security deposit?” AI handles these reasonably well because the definitions are stable and widely documented across its training data.
  • Procedural questions. “How many days do I have to appeal a small claims decision in my state?” This is where AI is weakest. Procedures vary by state, county, and sometimes by court, and the model often gives a generic answer as if one national rule applied everywhere.
  • Judgment-based questions. “Should I accept this settlement?” or “Should I sign this lease addendum?” These require weighing your specific circumstances, and AI has no way to know your full situation. An answer here is a guess dressed up as advice.

✽ The riskiest pattern is a procedural or judgment question answered with the same flat confidence as a factual one. That flattening, treating “your state’s deadline” like a fixed fact, is one of the most common ways AI legal guidance goes wrong.

Here is why the procedural category trips AI up so often. Training data is not evenly distributed across all fifty states. Some states publish far more court opinions, legal aid articles, and news coverage online than others, so the model has seen far more examples of, say, California eviction procedure than Wyoming eviction procedure. When you ask about a less-documented state, the model still answers, often by defaulting to the most common pattern it has seen elsewhere and presenting it as if it were universal. You have no way to know, from the answer alone, whether you got a well-supported response or an educated guess wearing a confident tone.

What to Verify Before You Act

  • Confirm the jurisdiction. Ask directly: “Does this answer apply specifically to [your state/county]?” A generic answer is not a wrong answer, but it is an incomplete one.
  • Check the deadline independently. If AI gives you a number of days to respond, appeal, or file, verify it against your court’s website or your official notice. Deadlines are exactly where hallucination causes the most damage.
  • Ask AI to name its source. If it cannot point you to a specific statute, court rule, or official government page, treat the answer as a starting point, not a conclusion.
  • Watch for missing caveats. A trustworthy answer to a judgment-based question should tell you when to get a lawyer, not just tell you what to do.
  • Re-read the question you actually asked versus the question it actually answered. AI sometimes answers a simpler, more general version of your question without saying so.

➜ For a full six-question framework you can use on any AI legal answer, see “How to Tell If the AI Advice You Got Is Actually Good.”

When to Talk to a Real Lawyer

Some situations are not appropriate for AI-only guidance, no matter how the answer sounds. Talk to a licensed attorney, or at minimum a free legal aid clinic, when:

  • Money changes hands based on your decision, like a settlement offer or a lease with financial penalties.
  • A deadline has legal consequences if missed, like an appeal window or a statute of limitations.
  • The other side already has a lawyer.
  • You are unsure whether your situation is even the kind of case AI described.

Many states have free or low-cost legal aid organizations for exactly these situations. A short call to one of them costs far less than a missed deadline or an accepted settlement you should not have taken.

None of this means you should stop using AI for legal questions. It means you should use it the way you would use a knowledgeable friend who has read a lot but has no law degree: a good starting point for understanding the landscape, and never the last stop before you act. The lawyers in the cases above did not lose their sanctions hearings because they used AI. They lost them because they treated an unverified AI answer as if it were already verified.

Where this comes from

CAVA’s evaluation methodology grew out of the AskAFriend Consumer Guides, a book series covering legal literacy, insurance, and other areas where everyday Americans get hurt by systems they do not fully understand. That research kept turning up the same problem: AI-generated guidance that sounded right and was not. No tool existed to measure that gap systematically, which is why CAVA was built.

Next step

Before you act on any AI legal answer, run it through the checklist above. If you want to understand the deeper evaluation framework behind it, read “How to Tell If the AI Advice You Got Is Actually Good,” or explore the CAVA Methodology section to see how these failures are scored dimension by dimension.