Abstract
The CAVA Reference Implementation is a pre-registered study evaluating AI-generated consumer guidance across six domains, 270 prompts, and five AI systems. This article explains what is being tested, why those specific domains were chosen, how the evaluation works, and where the results will be published when the study is complete.
Dilemma
When you ask an AI for help with a legal problem, an insurance dispute, a medical bill, or a home purchase, you are not taking a test. You are trying to solve a real problem with real consequences. The AI’s answer may be the only guidance you get before making a decision that costs you time, money, or both.
The CAVA Reference Implementation exists to find out how well five leading AI systems actually perform in those moments. Not on academic benchmarks. Not on trivia. On the kinds of questions real consumers ask in the six domains where AI guidance failures hurt people the most.
The Six Domains: Why These and Not Others
The study evaluates AI-generated guidance across six consumer domains. Each one was chosen because it meets a specific set of criteria: the stakes are material, the rules vary by state or jurisdiction, the process involves multiple steps, and historically people have needed professional help to navigate it correctly. These are not easy domains. They are the domains where getting it wrong costs the most.
✔ Legal Literacy
This domain covers the legal questions consumers face without attorneys: landlord-tenant disputes, small claims court, debt collection, personal injury basics, and self-advocacy. Legal guidance is jurisdiction-specific by definition. A deadline that applies in Texas may not apply in Colorado. An AI that does not acknowledge this is giving you guidance that could lose your case.
✔ Insurance Navigation
Insurance is one of the most complaint-generating industries in the United States. The study tests whether AI can accurately guide consumers through claim disputes, denial appeals, policy interpretation, and state-specific procedures under the McCarran-Ferguson framework, where insurance regulation is controlled by individual states.
✔ Home Buying
The average home purchase is the largest financial transaction most Americans ever make. The study tests AI guidance on mortgage qualification, contract terms, inspection rights, escrow procedures, and consumer protections under the Real Estate Settlement Procedures Act. This is a domain where missing one step can cost tens of thousands of dollars.
✔ Consumer Fraud
This domain covers fraud recovery, identity theft, debt disputes, and complaints against businesses. The study tests whether AI can accurately describe the processes and rights available under the Fair Debt Collection Practices Act, the CFPB complaint system, and state consumer protection agencies.
✔ Home Remodeling
Contractor fraud is one of the most common forms of consumer fraud in the country. The study tests AI guidance on permit requirements, contractor vetting, dispute procedures, and safety thresholds. Building codes are locally adopted and frequently amended. AI that presents general guidance as universally applicable is dangerous here.
✔ Health Advocacy
This domain covers the consumer side of healthcare: pre-authorization appeals, medical billing disputes, patient rights under HIPAA, discharge planning, and coverage questions under the Affordable Care Act and the No Surprises Act. This is a domain where incomplete or incorrect guidance can affect both health outcomes and financial stability.
The 270 Prompts: Three Tiers, One Architecture
The study uses exactly 270 prompts: 45 per domain, 15 per complexity tier, across three tiers. The three-tier structure is one of the most important design decisions in the framework. Here is what each tier tests and why all three matter.
✔ Tier 1: Factual
A Tier 1 prompt has a correct answer. “What is the deadline to respond to an eviction notice in Florida?” Either the AI knows the answer and states it accurately, or it does not. These prompts test whether the AI has reliable factual knowledge in the domain and whether it states that knowledge safely.
Tier 1 is the most important tier. Most AI evaluations test against sophisticated, well-formed questions because they are designed to show the AI at its best. Real consumers, especially those most likely to rely on AI because they cannot afford professional help, ask Tier 1 questions. In every exploratory analysis conducted during framework development, Tier 1 prompts exposed failure modes that more complex testing missed.
✔ Tier 2: Procedural
A Tier 2 prompt asks for a process. “How do I dispute a debt collection call?” There is a correct sequence of steps. The AI must not only know the steps but sequence them correctly, include all required actions, and tell you what happens if you skip one. These prompts test completeness and actionability.
✔ Tier 3: Judgment-Based
A Tier 3 prompt has no single right answer. “My landlord is claiming normal wear and tear as a damage deduction. Do I have a case?” The AI must reason about competing considerations, acknowledge what it does not know, correctly flag where professional judgment is required, and avoid giving the kind of confident, specific advice that leads consumers to act without adequate information.
✽ Tier 3 prompts are where the most consequential AI failures occur. They are also where current benchmarks offer the least useful signal.
The Five AI Systems: Why This Lineup
The study collects responses from five AI systems chosen to span meaningful variation in architecture, alignment approach, and how they retrieve and use information. The goal is not to rank the systems from best to worst. It is to identify whether failure patterns are consistent across AI systems or specific to particular architectures.
The lineup includes: a GPT-class model from OpenAI, a Gemini-class model from Google, a Grok-class model representing a distinct alignment approach, and two retrieval-augmented systems that ground answers in live search results. Testing both standard and retrieval-augmented systems allows the study to examine whether live data access improves accuracy and jurisdiction sensitivity in practice.
✽ Model names and versions are documented at collection time and locked for each domain block. If a model is updated mid-collection, affected responses are discarded and re-collected under the new version. This is the study’s mid-window disruption protocol, specified in the pre-registered design.
The Six Evaluation Dimensions: What Gets Measured
Every AI response in the study is scored on six dimensions, each on a five-point anchored scale. The dimensions are defined in the Universal Core Framework v1.0 (doi.org/10.5281/zenodo.20511504) and are not redefined here. A brief summary of what each measures:
✔ Accuracy:
Is the information factually correct? No hallucinated statutes, fabricated deadlines, or invented procedures.
✔ Completeness:
Does the response include everything the consumer needs to act? A response that is accurate in everything it says and still missing critical steps fails on Completeness.
✔ Actionability:
Can a real person follow these instructions? Vague advice like “contact the appropriate agency” without naming the agency is an Actionability failure.
✔ Safety:
Does the response include appropriate warnings, flag professional referral where needed, and avoid giving consumers false confidence?
✔ Jurisdiction Sensitivity:
Does the response correctly handle variation in law across states and localities? This is the most consistently observed failure mode in consumer AI guidance.
✔ Transparency:
Does the AI accurately represent what it knows and does not know? The most dangerous responses are not the ones that say “I don’t know.” They are the ones that say nothing about uncertainty at all.
Each dimension is scored independently for every response. Scorers assign each dimension on its own evidence rather than forming a single overall impression. This prevents one strong dimension from masking failures in another.
How Scoring Works: The Dual-Track Approach
The study uses a dual-track evaluation process. The two tracks are not competing. They are complementary, designed to catch what the other misses.
✔ Track A: Expert panel review.
Domain experts score responses in two pilot domains: Legal Literacy and Health Advocacy. These two were chosen because they involve the most subjective judgment and carry the highest stakes. Inter-rater reliability is assessed using Krippendorff’s alpha, with a target of 0.70 or higher per dimension before results from those domains are reported.
✔ Track B: Automated scoring pipeline.
An AI judge evaluates every response against the six-dimension rubric. The automated track also runs structural checks: detecting whether the response acknowledges jurisdiction, includes conditional language, provides step-by-step structure, and flags professional referral. Automated scores must achieve a Pearson correlation of 0.70 or higher against expert scores in the pilot domains before being applied to the remaining four domains.
Scores from both tracks are merged into a single dataset. Discordant scores are reviewed and documented. A result that cannot be reproduced is not reported as a finding.
What the Study Will Produce
The Reference Implementation is designed to generate findings that are specific enough to be actionable and reproducible enough to be trusted. The planned outputs include:
Results are published at Ask a Friend Publishing (askafriend.com) as evaluations are completed. The site is the consumer-facing repository for the study: anyone can read the evaluated AI responses, see the scores, and understand what each AI system got right and wrong on a given question. The prompt library and underlying data are published open source at github.com/owalcher/askafriend-AI-consumer-analysis under CC BY 4.0.
What This Study Does Not Claim
The Reference Implementation will not tell you which AI system is the best overall. It will tell you how each system performs on consumer guidance quality in these specific domains, measured against these specific criteria, at a specific point in time. AI systems change. Models are updated and replaced. The study produces a baseline and a reusable methodology for tracking change over time.
The study also does not evaluate internal model architecture, training data, or alignment mechanisms. It evaluates observable outputs against a pre-registered rubric. That scope boundary is deliberate. Findings stay within what the data can actually support.
✽ This is what it means to govern claims explicitly. The framework specifies not just what the study measures, but what it is not permitted to conclude. That boundary is as important as the findings themselves.
➜ Read the pre-registered methodology: doi.org/10.5281/zenodo.20511504
➜ Read the pre-registered study design: doi.org/10.5281/zenodo.20511747
➜ Follow the open data repository: github.com/owalcher/askafriend-AI-consumer-analysis
➜ Read published study results as they are released: askafriend.com
Other Articles
What AI Actually Does When You Ask It a Legal Question
When you ask AI a legal question, it predicts the next likely word based on patterns in its training data.
Why We Published Our Methodology Before Our Results: The case for Registered Protocols
AI evaluation results are only trustworthy if the methodology was locked before data collection began.



