Abstract
AI evaluation results are only trustworthy if the methodology was locked before data collection began. CAVA published the Universal Core Framework , the complete six-dimension rubric, study design, and behavioral detection rules, as a registered protocol before testing a single AI system. Here is why that decision matters and what it protects you from.
Background
You have probably read an AI evaluation report. A technology publication releases findings on which AI chatbot performs best. A company publishes internal testing showing their model outperforms the competition. The scoring criteria appear in a footnote, if at all. You have no way to know whether those criteria were designed before or after the researchers saw the results.
That matters more than it sounds. When the people running an AI evaluation also design the rubric, select the prompts, and decide what counts as a good answer, and they do all of this while looking at the data, the results cannot be trusted. Not because anyone cheated. Because the method is broken by design.
The Invisible Problem in AI Evaluation
Most AI evaluations are not reproducible. The researchers who conduct them typically design the scoring rubric, collect the data, and publish the findings in one continuous process, with no external record of what the methodology was before they saw what the AI produced. The technical term for this is post-hoc methodology, and it corrupts findings even when everyone involved is acting in good faith.
Here is how it works in practice. A researcher notices that one AI system tends to give more detailed answers than another. They design a scoring dimension that rewards detail. The detailed system scores better. The finding looks like a discovery, but the evaluation did not discover that pattern. It constructed it. The rubric was built around what the researcher already observed.
This problem is not limited to careless researchers. It affects well-funded labs, academic institutions, and commercial bench-marking services. It is structural. The only fix is to lock the methodology before data collection begins, and prove it.
Why This Problem Is Worse for Consumer AI Evaluation
Standard AI benchmarks — tests that measure whether a model can pass a bar exam or solve math problems — are relatively resistant to post-hoc methodology because the answers are objectively correct or incorrect. Either the model got the right answer or it did not.
Consumer guidance is different. When you ask an AI how to dispute a medical bill, what your rights are when a landlord keeps your security deposit, or how to appeal an insurance denial, there is no single right answer. The quality of the guidance depends on procedural completeness, jurisdictional accuracy, appropriate safety caveats, and honest acknowledgment of uncertainty. These are judgment calls.
Judgment calls are exactly where post-hoc methodology does the most damage. A rubric designed to evaluate consumer guidance must define what good looks like before seeing what the AI produces, otherwise the rubric bends toward what the AI already does well, and the evaluation measures performance against a standard the AI helped create.
What a Registered Protocol Is
A registered protocol is a methodology document published in a permanent, timestamped repository before data collection begins. It locks in, publicly and verifiably, the complete study design: the evaluation dimensions, the scoring criteria, the prompt structure, the model selection rules, the reliability thresholds, and the hypotheses the study will test.
Registered protocols are not a bureaucratic formality. They are the minimum bar for producing evaluation findings that deserve to be believed. Any AI evaluation that cannot point to a pre-registered methodology is asking you to take its results on faith.
What CAVA Did — and Where It Came From
The Universal Core Framework v1.0 was published on Zenodo as a registered protocol in May 2026, before a single AI response was collected for evaluation. The complete methodology: six evaluation dimensions, anchored scoring descriptors, behavioral signature detection rules, inter-rater reliability requirements, and explicit governance over what claims the data can and cannot support, is available at https://doi.org/10.5281/zenodo.20511504.
A companion Reference Implementation: the specific study design that deploys the framework across six consumer domains, 270 prompts, and five AI systems, was published separately at https://doi.org/10.5281/zenodo.20511747. Both are open access under CC BY 4.0. Anyone can read them, replicate them, or build on them.
That decision, methodology first, results second, did not happen because it was required. It happened because of where CAVA came from.
The Universal Core Framework was not built in an AI lab. It was built by a researcher who spent years writing consumer guides — books on navigating the legal system, understanding insurance, advocating for yourself in healthcare, protecting yourself from fraud. Books written for the Americans who get hurt by systems they do not understand.
In the course of that research, the same AI systems that millions of consumers now consult produced guidance that sounded authoritative and was wrong in ways with lasting consequences: jurisdictionally incorrect legal procedures, insurance appeal deadlines that do not exist, patient rights that vary by state presented as universal rules. No measurement tool existed to evaluate this systematically.
Those two decisions are what separate CAVA’s findings, when published, from the AI evaluation reports you cannot verify.
The Open Research Stack
The registered protocol is one layer of a fully open research architecture. Every component of the CAVA study is public before the next layer is built:
This means you do not have to take CAVA’s word for anything. You can read the methodology before the results are published. You can examine the prompts before the AI is tested. When results are published, you can verify that the scoring followed the pre-registered criteria, or read the documented deviation notices where it did not.
What to Ask About Any AI Evaluation You Encounter
The next time you read an AI evaluation report, whether from a technology publication, a model developer, or a third-party bench-marking service, these four questions will tell you whether the findings are worth trusting:
The Bottom Line
The Universal Core Framework exists because consumer AI guidance is being evaluated — when it is evaluated at all — with methods that would not survive scrutiny in any field where the stakes are measured in human consequences. A registered protocol is not extra credit. It is the difference between findings that can be trusted and findings that cannot be verified.
CAVA published the methodology before the results because that is the only way the results mean anything. That standard is available for any researcher, organization, or regulator who wants to apply it.


