You did not write it. You still have to act on it.
Filings from opposing counsel, claim narratives from contractors, research from vendors: AI now drafts the documents that land on a desk, and nobody upstream was required to prove them true. GauntletScore checks every claim in the document you received against primary sources and returns a calibrated probability you can defend, with the evidence behind every point.
The last person to accept a document inherits every unchecked claim inside it.
In a 2026 California appellate matter, a fabricated legal citation began its life in a Reddit post. It moved into a sworn declaration, then into an attorney's filings, then into opposing counsel's proposed order, and then under a judge's signature. Nobody in that chain invented it. Every one of them received it, and every one of them assumed someone upstream had checked. The court sanctioned the attorney who introduced it $5,000 and referred the matter to the State Bar.
A public database now tracks more than 1,000 court and tribunal decisions worldwide in which a judge addressed fabricated AI content in a filing. Those are only the cases where the fabrication was caught and put in writing.
Verification is not a courtesy you owe the sender. It is how you keep their unchecked claims from becoming your record.
Asking another model is not an independent check. We measured it.
The natural reflex with a document you received is to paste it into a chatbot and ask whether it looks right. Models trained on overlapping data share blind spots, so a second model confirms the first one's error with equal confidence. A second checker cannot be another model's opinion. It has to be deterministic math, or it inherits the blind spots of whatever wrote the document. We ran the experiment: twenty AI-generated company profiles, three independent frontier checkers, identical instructions.
On the same document, one frontier model flagged 48 errors and another flagged 5. Neither has a claim to being correct. The choice of checker is an unmanaged variable.
The harshest checker returned zero errors on three documents. Adversarial verification against primary sources then found 16 verifiable errors in them. A clean self-check report is not evidence that a document is clean.
Their answers survive an argument. Our claims survive the evidence.
Enterprises already govern AI by score. In KPMG's Q2 2026 AI Pulse survey of 204 organizations with more than $1 billion in revenue, 43% override AI outputs when a confidence or quality score falls below a threshold, and 33% override ad hoc with no formal criteria at all. The score already runs the decision. The question is whether it is computed independently and backed by evidence. (Source: KPMG AI Pulse, Q2 2026.)
The AI is the witness. The math is the judge.
A second model asked to check the first produces an opinion that cannot be audited, with errors correlated to the ones it is checking. GauntletScore removes the model from the verdict entirely.
The sample run: twelve claims, each addressed by every agent. Only evidence moves forward; opinions do not.
Seven agents, four model providers. Each claim is checked against primary sources: SEC and EDGAR filings, court records, regulatory databases, the medical literature. Agents disagree, escalate, and produce verdicts. This is the sensor array, and it is the part anyone can build.
Deterministic math only. The verdicts become a Bayesian posterior. No language model runs here. An automated integrity test fails the build if anyone wires one in.
Surviving a debate is not the same as being verified. Other tools end at the answer that won the argument. GauntletScore continues, to an auditable probability with the evidence behind it.
Grounded tells you where it came from. Verified tells you whether it's true.
The score is a probability, derived in one line of math.
Every verified claim contributes one term. The prior is the baseline for the document type. Each term is a likelihood ratio weighted by the quality of the source that produced it. The sum, passed through a logistic function, is the score.
The same evidence always produces the same score, to the digit. The math has no moods and no temperature setting.
Every point traces to a specific claim, verdict, source, and likelihood ratio. "Why this score" has a real answer.
Cromwell's rule: no finite evidence reaches certainty. The score reports a 95% interval that widens when evidence is thin.
From document to signed certificate in minutes.
Upload a document or send it through the API.
Seven agents from four providers check each claim against primary sources, then defend their findings against adversarial challenge. The agents gather evidence; they do not set the score.
The verdicts are combined by a deterministic Bayesian engine into a trust score, per-claim verdicts, and a cryptographically signed certificate. The report arrives in minutes, with per-claim verdicts and the signed certificate.
It tests whether the reasoning holds, not just whether the facts check out.
A document can be built from individually true facts and still argue something false. GauntletScore runs a dedicated analysis over every cause-and-effect claim, testing temporal order, proportionality, confounders, and logical structure. The result enters the same posterior with a deliberate asymmetry: sound reasoning earns only a small positive weight, because internal consistency is not external proof. Broken reasoning counts heavily against the document. Most AI operates at the level of association, what correlates with what. The causal pass operates at the levels of intervention and counterfactual: what changes what, and what would have happened otherwise.
Errors that survive a single-model check.
Every example below arrived reading as authoritative. Every one is a documented catch from our pre-registered validation study, verified against primary sources.
A generated profile named a Chief Growth Officer who does not exist. Caught against corporate filings.
A profile described a nine-figure settlement between two companies that never occurred. Caught against court records.
Research and development spending asserted at roughly four times the documented figure. Caught against the filings.
A CFO named who had been succeeded; the real appointment is in the company's press releases. Caught against the corporate record.
A financial walk asserting 1.2 + 6.6 = 12.5. Caught by the deterministic math verifier.
We ran the system on our own paper. It caught one of our own errors.
not overwritten
Our validation study originally reported 27 tool-verified catches. When we ran GauntletScore on the draft manuscript itself, it flagged one of our own examples as a false positive. We retracted the example, corrected the public count to 26, and documented the change in the study record. A verification company that cannot survive its own gauntlet has no business selling one.
Change histories are editable. Signatures are not.
Frontier model cards now document models concealing actions and editing change histories so the changes would not appear in the record. The audit trail you need is one that neither a human nor a model can quietly rewrite. Every Gauntlet Report ships with a cryptographically signed, tamper-evident certificate recording what was checked, against which sources, and with what result. Hand it to a reviewer, a regulator, or opposing counsel; anyone can verify it has not been altered.
Attach it to the file, cite it in the record, or send it back upstream with the document. Anyone can check the signature. No account is required.
Pick the desk the document landed on.
Run the document you received.
Upload the file. Get a score, a 95% interval, every flagged claim with its source, and a signed certificate. Three free credits. No card. No call.