Gauntletscore · The Gauntlet

You did not write it. You still have to act on it.

Filings from opposing counsel, claim narratives from contractors, research from vendors: AI now drafts the documents that land on a desk, and nobody upstream was required to prove them true. This page is the gate. It runs a sample document through the Gauntlet as you scroll.

Scroll to run. Scroll back to rerun. Same evidence, same score.

Scroll
01 · The document arrives

It reads with full confidence. So would the errors.

Below is a sample: an AI-generated company profile of the kind our validation study examined. The planted errors mirror documented catch classes from the study. On a read, they are indistinguishable from the correct material around them.

SAMPLE
Received · Company profile · Author unknown · Drafting model unknown

Northgate Instruments plc, Overview

Northgate Instruments designs precision measurement systems for industrial process control. The company reported revenue of $412M in fiscal 2025, with research and development spending of $176M, reflecting its stated commitment to platform renewal.

Following the $340M settlement of the Halworth litigation, the company reorganized its compliance function under Chief Growth Officer Daniel Reeves, who joined from a competitor in 2024. Management attributes the 9% margin improvement to the automation program completed in March.

The balance of the profile continues in the same register: fluent, specific, and unsourced.

Four claims are underlined. The Gauntlet has not examined any of them yet.

02 · How unchecked claims travel

The last person to accept a document inherits every unchecked claim inside it.

In a 2026 California appellate matter, a fabricated legal citation began its life in a Reddit post. It moved into a sworn declaration, then into an attorney's filings, then into opposing counsel's proposed order, and then under a judge's signature. Nobody in that chain invented it. Every one of them received it, and every one of them assumed someone upstream had checked.

FORUM POSTDECLARATIONFILINGPROPOSED ORDERSIGNATURE

The court sanctioned the attorney who introduced it $5,000 and referred the matter to the State Bar. A public database now tracks more than 1,000 court and tribunal decisions worldwide in which a judge addressed fabricated AI content in a filing.

Verification is not a courtesy you owe the sender. It is how you keep their unchecked claims from becoming your record.

03 · The obvious check fails

Asking another model is not an independent check. We measured it.

The natural reflex with a document you received is to paste it into a chatbot and ask whether it looks right. Models trained on overlapping data share blind spots, so a second model confirms the first one's error with equal confidence. We ran the experiment: twenty AI-generated company profiles, three independent frontier checkers, identical instructions.

9.6 : 1

Variance in self-check error counts across frontier models on the same document. Which checker you happen to pick moves the count more than the errors do.

0 vs 16

The harshest single-model checker returned zero errors on documents containing sixteen verifiable ones.

43%

Share of organizations with more than $1 billion in revenue that already override AI outputs when a confidence score falls below a threshold. The score already runs the decision.

A second checker cannot be another model's opinion. It has to be deterministic math, or it inherits the blind spots of whatever wrote the document.

04 · The Gauntlet runs

The AI is the witness. The math is the judge.

Seven adversarial agents from independent model providers examine the sample. They query the primary record, challenge each other's findings across four rounds, and forward evidence, never opinions, to a deterministic scoring layer. Watch the twelve extracted claims resolve as you scroll. Scroll up and the debate runs in reverse, to the same result.

ROUND 0 / 4
VERIFIEDDEBUNKEDUNVERIFIABLEOPEN
COURTLISTENERECFRPUBMEDSEC EDGARCLINICALTRIALS.GOVCROSSREF

Each analysis triggers between 100 and 200 queries to authoritative databases. A claim no public source can confirm is reported as unverifiable, not as false.

05 · The verdict is arithmetic

No model sits in the verdict. The same evidence always produces the same score.

posterior log-odds = prior + Σ log LR(evidence i)   →   score = 100 · σ(posterior)

The shape of the scoring rule. Each verified claim adds evidence, each debunked claim subtracts it, and Bayes' rule does the rest.

50/ 100 · DEMONSTRATION RUN
0100
95% interval, narrowing as evidence accumulates. It never touches 0 or 100. Cromwell's rule keeps certainty off the dial.
Deterministic

The agents gather the evidence; the score is computed, not voted.

Auditable

Every point on the dial traces to a claim, a source, and a verdict.

Bounded

The interval states what the run can and cannot support.

06 · The standard

We ran the system on our own paper. It caught one of our own errors.

Our validation study originally reported 27 tool-verified catches. When we ran GauntletScore on the draft manuscript itself, it flagged one of our own examples as a false positive. We retracted the example, corrected the public count to 26, and documented the change in the study record.

2726correction logged,
not overwritten

A verification company that cannot survive its own gauntlet has no business selling one.

Pre-registered validation study of 20 public companies, conducted with independent academic oversight. 1,829 claims examined in Phase 2. Study in progress; manuscript in preparation.

07 · The artifact

Change histories are editable. Signatures are not.

Every real run ends in a cryptographically signed, tamper-evident certificate: a portable record that the verification occurred, what it found, and when. Attach it to the file, cite it in the record, or send it back upstream with the document. Anyone can check the signature. No account is required.

DEMONSTRATION
Gauntlet Verification Certificate
DocumentNorthgate Instruments plc, profile (demonstration)
Score23 / 100 · 95% interval 17 to 31
Claims examined12 · 6 verified · 3 debunked · 3 unverifiable
Signature schemeEd25519
SignatureDEMONSTRATION-0000 . this demonstration certificate signs nothing . any real run yields a verifiable one

GauntletScore does not decide for you, and it does not verify what no public source can confirm. A claim it cannot check is reported as unverifiable, not as false. The decision stays yours; the record that it was checked is the product.

08 · The one on your desk

Run the document you received.

Upload the file. Get a score, a 95% interval, every flagged claim with its source, and a signed certificate.

Analyze a Document
Three free credits. No card. No call.
SourceKPMG AI Pulse survey, Q2 2026. 204 organizations with more than $1 billion in annual revenue. 43% report overriding AI outputs when a confidence or quality score falls below a defined threshold; 33% report overriding ad hoc with no formal criteria.
RegistrationThe validation study methodology was pre-registered on the Open Science Framework in March 2026, before Phase 2 analysis. Six locked hypotheses. The study is in progress and the manuscript is in preparation; no result on this page is presented as peer-reviewed.
SourceA public database of court and tribunal decisions worldwide, maintained by a legal researcher, in which judges explicitly addressed fabricated AI content in filings. It passed 1,000 documented decisions in 2026. Those are only the cases where the fabrication was caught and put in writing.