Test the exam and the rating decision against the record, before the appeal is built on them.
A veterans practice runs on documents other people drafted: the C&P exam, the rating decision, the treatment summaries, the inherited file. AI now drafts and summarizes them at every step. A wrong diagnostic code, a misdated symptom onset, or a combined-rating arithmetic error in these documents is not a typo. It is the difference between a granted claim and years of appeals.
The errors hide in the specifics the case turns on: a 38 CFR reference that does not say what the decision says, a service-connection timeline that does not hold, a treatment detail the record contradicts, combined ratings that do not compute. The rating decision reads as authority, and the exam reads as evidence, whether or not the specifics hold. Asking another AI to check the draft is not a fix. In our pre-registered study of twenty company profiles, three frontier checkers disagreed by as much as 9.6 to 1 on the same document, and the harshest one returned zero errors on documents containing sixteen verifiable ones.
The AI is the witness. The math is the judge. GauntletScore verifies each claim against primary sources, then computes the trust score with deterministic math. The agents gather the evidence; no model sits in the verdict. The same evidence always produces the same score.
GauntletScore runs a dedicated analysis of cause-and-effect claims, testing service-connection chains, symptom-to-diagnosis attribution, and temporal ordering. A nexus argument whose causal chain fails the test counts heavily against the document, which is exactly the argument a rating decision will reject.
Upload the document at gauntletscore.com. Minutes later you are reading the Gauntlet Report: a trust score with a 95% interval, every flagged claim with its verdict and the primary source behind it, and a cryptographically signed, tamper-evident certificate. The error in the exam that denied your client is exactly the finding the appeal is built on, and it comes with the source that contradicts it. In a live production run, the engine examined nineteen claims in one fluent, credible document and returned two debunked against the court record, each with the source that contradicts it. Every voting agent independently recommended against proceeding.
LIMITSIt does not assess the merits of the claim and it does not verify what no source can confirm, including private medical records it is not given. A claim it cannot check is reported as unverifiable, not as false.
Test the specifics the denial turns on against the record before the appeal strategy is set.
Run the checkable specifics of an inherited file before committing the practice's time.
Verify every citation, code, date, and calculation before the appeal or supplemental claim is filed.
There is a rating decision on your desk right now that turned on specifics nobody has tested. Run it. Either they hold, or one does not, and that finding is the appeal.
Validation study pre-registered on the Open Science Framework, March 2026; in progress, manuscript in preparation.