Why We Published Our Methodology Before Our Results: The case for Registered Protocols

Why Did We Publish Our Methodology Before Our Results The Case for Registered Protocols


Abstract

AI evaluation results are only trustworthy if the methodology was locked before data collection began. CAVA published the Universal Core Framework , the complete six-dimension rubric, study design, and behavioral detection rules, as a registered protocol before testing a single AI system. Here is why that decision matters and what it protects you from.

Background

You have probably read an AI evaluation report. A technology publication releases findings on which AI chatbot performs best. A company publishes internal testing showing their model outperforms the competition. The scoring criteria appear in a footnote, if at all. You have no way to know whether those criteria were designed before or after the researchers saw the results.

That matters more than it sounds. When the people running an AI evaluation also design the rubric, select the prompts, and decide what counts as a good answer, and they do all of this while looking at the data, the results cannot be trusted. Not because anyone cheated. Because the method is broken by design.

The Invisible Problem in AI Evaluation

Most AI evaluations are not reproducible. The researchers who conduct them typically design the scoring rubric, collect the data, and publish the findings in one continuous process, with no external record of what the methodology was before they saw what the AI produced. The technical term for this is post-hoc methodology, and it corrupts findings even when everyone involved is acting in good faith.

Here is how it works in practice. A researcher notices that one AI system tends to give more detailed answers than another. They design a scoring dimension that rewards detail. The detailed system scores better. The finding looks like a discovery, but the evaluation did not discover that pattern. It constructed it. The rubric was built around what the researcher already observed.

  • ✘ When a rubric is designed after the data is collected, the scoring criteria inevitably mirror what the researcher already found, inflating the results.
  • ✘ This makes findings unreproducible. When another researcher uses the same design, they get different results, because the original rubric was tuned to one dataset.
  • ✘ It also makes the findings unfalsifiable. If you designed the test to produce a result, you cannot use that result as evidence.

This problem is not limited to careless researchers. It affects well-funded labs, academic institutions, and commercial bench-marking services. It is structural. The only fix is to lock the methodology before data collection begins, and prove it.

Why This Problem Is Worse for Consumer AI Evaluation

Standard AI benchmarks — tests that measure whether a model can pass a bar exam or solve math problems — are relatively resistant to post-hoc methodology because the answers are objectively correct or incorrect. Either the model got the right answer or it did not.

Consumer guidance is different. When you ask an AI how to dispute a medical bill, what your rights are when a landlord keeps your security deposit, or how to appeal an insurance denial, there is no single right answer. The quality of the guidance depends on procedural completeness, jurisdictional accuracy, appropriate safety caveats, and honest acknowledgment of uncertainty. These are judgment calls.

Judgment calls are exactly where post-hoc methodology does the most damage. A rubric designed to evaluate consumer guidance must define what good looks like before seeing what the AI produces, otherwise the rubric bends toward what the AI already does well, and the evaluation measures performance against a standard the AI helped create.

  • ✽ This is why most AI evaluation work in consumer domains produces findings that cannot be replicated. The methodology was designed around the data. Change the data, change the result.

What a Registered Protocol Is

A registered protocol is a methodology document published in a permanent, timestamped repository before data collection begins. It locks in, publicly and verifiably, the complete study design: the evaluation dimensions, the scoring criteria, the prompt structure, the model selection rules, the reliability thresholds, and the hypotheses the study will test.

  • ✔ A registered protocol cannot be changed after data collection begins without a public, documented deviation notice, on record, visible to anyone who reads the results.
  • ✔ It creates a verifiable audit trail: anyone can compare the published protocol to the published results and confirm the methodology was not altered to fit the findings.
  • ✔ It is the standard in clinical research, where unreliable findings cost lives. It is almost entirely absent from AI evaluation research.

Registered protocols are not a bureaucratic formality. They are the minimum bar for producing evaluation findings that deserve to be believed. Any AI evaluation that cannot point to a pre-registered methodology is asking you to take its results on faith.

What CAVA Did — and Where It Came From

The Universal Core Framework v1.0 was published on Zenodo as a registered protocol in May 2026, before a single AI response was collected for evaluation. The complete methodology: six evaluation dimensions, anchored scoring descriptors, behavioral signature detection rules, inter-rater reliability requirements, and explicit governance over what claims the data can and cannot support, is available at https://doi.org/10.5281/zenodo.20511504.

A companion Reference Implementation: the specific study design that deploys the framework across six consumer domains, 270 prompts, and five AI systems, was published separately at https://doi.org/10.5281/zenodo.20511747. Both are open access under CC BY 4.0. Anyone can read them, replicate them, or build on them.

That decision, methodology first, results second, did not happen because it was required. It happened because of where CAVA came from.

The Universal Core Framework was not built in an AI lab. It was built by a researcher who spent years writing consumer guides — books on navigating the legal system, understanding insurance, advocating for yourself in healthcare, protecting yourself from fraud. Books written for the Americans who get hurt by systems they do not understand.

In the course of that research, the same AI systems that millions of consumers now consult produced guidance that sounded authoritative and was wrong in ways with lasting consequences: jurisdictionally incorrect legal procedures, insurance appeal deadlines that do not exist, patient rights that vary by state presented as universal rules. No measurement tool existed to evaluate this systematically.

  • ✔ The first decision was to build the measurement tool before running the measurement.
  • ✔ The second decision was to publish it openly, so the methodology could be audited, replicated, and extended by anyone, not treated as a proprietary scoring system only the creator can verify.

Those two decisions are what separate CAVA’s findings, when published, from the AI evaluation reports you cannot verify.

The Open Research Stack

The registered protocol is one layer of a fully open research architecture. Every component of the CAVA study is public before the next layer is built:

This means you do not have to take CAVA’s word for anything. You can read the methodology before the results are published. You can examine the prompts before the AI is tested. When results are published, you can verify that the scoring followed the pre-registered criteria, or read the documented deviation notices where it did not.

  • ➜ This architecture is the consumer version of the same transparency standard clinical trials have followed for decades. AI guidance that affects people’s legal rights, insurance coverage, and healthcare decisions deserves the same bar.

What to Ask About Any AI Evaluation You Encounter

The next time you read an AI evaluation report, whether from a technology publication, a model developer, or a third-party bench-marking service, these four questions will tell you whether the findings are worth trusting:

  • Was the methodology published before data collection began? If not, you cannot rule out post-hoc construction of the scoring criteria.
  • Are the scoring dimensions publicly defined and fixed? Or are they described vaguely and applied at the evaluator’s discretion?
  • Can another researcher replicate the findings using the same prompts, models, and rubric? If the prompts and rubric are not published, replication is impossible, and unreproducible results are not evidence.
  • Who funded the evaluation? An AI benchmark produced by the company whose models it tests, or funded by AI developers, is not independent, regardless of how rigorous the methodology appears.
  • ✽ These questions do not assume bad faith. They are the minimum standard for distinguishing trustworthy evaluation findings from results that look like findings.

The Bottom Line

The Universal Core Framework exists because consumer AI guidance is being evaluated — when it is evaluated at all — with methods that would not survive scrutiny in any field where the stakes are measured in human consequences. A registered protocol is not extra credit. It is the difference between findings that can be trusted and findings that cannot be verified.

CAVA published the methodology before the results because that is the only way the results mean anything. That standard is available for any researcher, organization, or regulator who wants to apply it.

Similar Posts