
About this series. This page is the second of a five-part series on verifying artificial intelligence. Each page can be read on its own, but they build on one another, and later pages refer back to earlier ones by number or by title.
Page One: "What Verification Is." A Framework for Knowing What to Trust
Page Two: "The Material Anchor." Evidence, Auditing, and the Failure of Digital Text
Page Three: "The Self-Certification Collapse." Why the Institutions Meant to Check AI Can't Be Trusted to Check Themselves
Page Four: "Solomon's Fork." Testing What Can't Be Trusted to Report on Itself
Page Five: "The Terminal Boundary." Where Checking Ends and Judgment Begins
Picture two different ways someone could try to convince you they've checked their facts. In the first, they think about the claim again, more carefully this time, and tell you they're now confident it holds up. In the second, they hand you a document, a filing, a record, a source, and say, "Here, look for yourself." These feel similar. Both end with someone telling you a claim is solid. But they are not the same kind of act, and the difference between them is the idea this whole page is built on.
When someone re-examines their own reasoning, there is exactly one place anything could have gone wrong, and it's the same place it went wrong the first time, if it did. Re-checking doesn't create a second, independent opportunity to catch the error. It's the same process, run again, by the same mind, using the same information it already had. If the original reasoning had a blind spot, a second pass through that same reasoning has no particular reason to find it, because the thing doing the checking is the thing that might be broken.
A document works differently, not because documents are magically truer than careful thought, but because a document sits outside the person making the claim. Anyone, not just the person who found it, can open the same file, read the same page, and see whether it says what they claimed it says. That's a second, independent point of failure. If the claim is wrong, two separate things would have to be wrong at once: the original claim and the document itself, rather than one thing checked twice. This is a reasoned argument, not something anyone ran an experiment to confirm. It follows from thinking carefully about what makes a check independent in the first place, the same way a second opinion from a different doctor is worth more than the same doctor reconsidering, not because the second doctor is smarter, but because their judgment doesn't route through the same reasoning that produced the first opinion.
It's worth being honest about where this advantage actually stops, because it isn't unlimited. A document only creates a real second point of failure once someone other than the person reporting on it actually opens it and checks it. If you tell someone "I read the source and it says X," and they simply believe you, nothing has actually been checked twice. The document existed, but the independent check didn't happen, it was only described. The advantage of an external document is real, but it isn't automatic. It has to actually be used by a second, independent set of eyes before it does the work being claimed for it.
Saying "go check the source" is easy advice and an easy thing to claim you've done. What does actually doing it require?
Four distinct operations, and they're worth separating because people routinely treat them as interchangeable. The first is Searching, simply finding where a claim might be addressed: an index, a search engine, a list of results. Searching tells you something exists somewhere in the neighborhood of your question. It tells you nothing about whether it's true. The second is Reading, going to one specific document and taking in what it actually says, as opposed to relying on someone else's summary of it, a headline, a snippet, a secondhand paraphrase. The third is Comparing, holding two specific statements side by side and checking whether they agree, contradict, or have nothing to do with each other. The fourth, and the one that actually deserves the word Verifying, is the combination of all three aimed at one specific claim: search for the source, read it directly, compare what it says against what was claimed, and say plainly whether the claim holds up, needs correcting, or turns out to be false.
This is a methodological distinction, not a discovery about the world: a definition built to tell real verification apart from something that merely resembles it. Most of what passes for "I checked" in ordinary conversation is really just the first of these four acts, searching, dressed up in the vocabulary of the fourth.
There's a further requirement, and it comes from an idea much older than anything on this page. A claim isn't genuinely verified just because someone gathered evidence that agrees with it. Confirming evidence is easy to find for almost anything, because a motivated search will usually turn up something supportive if that's the only kind of result being looked for. Genuine verification requires actively trying to find evidence that would prove the claim wrong, and only counting it as checked once that attempt fails to turn anything up. This is a version of an idea the philosopher Karl Popper argued for in the twentieth century, in a very different context, the philosophy of science: a claim only earns scientific credibility by surviving real attempts to falsify it, not by accumulating examples that happen to agree with it. Applied here, the same logic means the honest question isn't "can I find support for this," it's "did I actually go looking for the thing that would break it, and did it survive that search?"
The last piece of what real checking requires is less about the claim and more about how you show the checking to other people. It's possible to say "I checked, and it's fine" without showing anyone what checking actually consisted of, a bare summary, a conclusion with nothing underneath for a reader to inspect. Real transparency means the opposite: showing the specific quotes, the specific document, the specific comparison, so that someone else can look at the same evidence and reach their own conclusion rather than simply trusting a verdict. This has a limit worth naming honestly: showing your work only proves something if the work shown is what actually produced the answer. If the steps on display were written after the fact to justify a conclusion already reached, displaying them looks like transparency but isn't — it's the same unchecked claim wearing a citation.
Real instances are easier to believe than abstract descriptions, so here are three, drawn from an actual attempt to apply these standards in practice.
The clearest example of genuine reading, as opposed to relying on a summary, involved a specific legal filing, a complaint filed in Florida, and one numbered paragraph within it, paragraph 141. Rather than relying on a secondhand news summary of what that paragraph reportedly said, the actual PDF was retrieved, and the paragraph was read directly, word for word. This is a demonstrated instance, not a hypothetical one: the primary source was fetched, opened, and read, and whatever conclusion followed came from that document's actual language rather than someone else's characterization of it.
A second instance shows falsification in action: the requirement described in the previous section, applied to a real case. An initial grouping had assumed that a particular lawsuit, Rushlow v. Altman (PacerMonitor case 66591019), belonged with a set of similar Florida cases. Actively searching for evidence that the grouping might be wrong, rather than simply accepting it, revealed that Rushlow actually involved a different legal basis entirely — a shooting in Tumbler Ridge, British Columbia, filed in California federal court on diversity jurisdiction, with no connection to Florida anywhere in the case. That correction is a demonstrated finding: a specific, wrong grouping was actually broken by the act of looking for evidence against it, exactly the falsification standard described above, not just a description of what that standard would look like if applied.
The third instance is a demonstrated correction worth sitting with, because it shows a mismatch being caught and then resolved, not just a mismatch existing. A quote, "there's going to be a point in the next few years where basically everyone at this company has to switch to working on safety, or else we're fucked," was first attributed to a TechCrunch article. TechCrunch did not contain it. The quote itself was real, but the cited source was wrong. Further verification showed the correct source was the same legal filing discussed above: a complaint filed by the Office of the Attorney General, State of Florida, against OpenAI and Sam Altman, and paragraph 141. Being external didn't protect against the first error, a real, unrelated document doesn't verify a claim just because both exist in the same general vicinity of the story. What caught it was the specific act of comparing the specific claim against the specific line in the specific document, first finding that TechCrunch didn't have it, then finding the document that did.


There's an obvious shortcut a reader might expect at this point: if outside documents are what make verification real, and there's already a mature profession built around checking documents rigorously, financial auditing, why not just apply that profession's existing methods to the problem of verifying an AI system's behavior?
The reason is worth taking seriously precisely because the shortcut isn't a bad idea, it's a genuinely good one that happens to run into a hard limit. Financial auditors already work under a real, established standard issued by the Public Company Accounting Oversight Board, and these audit procedures work because they check what they're actually being asked to check: a company's financial statement describes a specific, already-completed set of transactions. Every transaction being verified has already happened by the time anyone goes looking for it. An auditor samples from that finite, dated, closed universe of events, traces each one to a source document, bank records, invoices, contracts, and checks whether what was recorded matches what the documents show. The entire method depends on the thing being checked having already occurred and stopped changing.
An AI system's future behavior has no such property. It isn't a closed, dated universe of things that have already happened. It's a prospective, essentially unlimited space of things the system might say or do in response to inputs nobody has written yet. There is no finite set of transactions to sample from, because the relevant 'transactions,' every possible thing the system could be asked and every way it could respond, haven't occurred and can't be enumerated in advance. Anthropic said as much about its own models. In a February 2026 update on whether Claude Opus 4.6 crossed a specific safety threshold, the company wrote that the CBRN-4 rule-out was 'less clear for Opus 4.6 than we would like,' citing 'a substantial degree of uncertainty' it could not close through benchmark testing alone. (Claude Opus 4.6 System Card Anthropic, Feb. 2026) This is a reasoned argument, not a report on a failed pilot program: it follows from comparing what the audit method actually requires, a bounded, backward-looking universe, against what verifying AI behavior actually is, an unbounded, forward-looking one. The audit method itself isn't flawed. It's being asked to do a job it was never built for, checking something that hasn't finished happening yet and, in an important sense, never will.

The previous stress test assumes something important stays true even as it points out a mismatch: it assumes that whatever documents and records do exist are genuine, that a transaction log, once found, actually reflects what happened. A second, harder problem doesn't make that assumption, and it's worth building toward carefully, because it's a sharper break than the one before it.
Everything discussed on this page so far treats "fetch the document and read it" as close to the most reliable method available, the gold standard the earlier sections built the whole case around. That method depends on one quiet assumption: that the thing being read is actually what it claims to be. A benchmark result, a safety report, a training log, a chat transcript, these are all, in the end, digital text. And once a generative AI system becomes capable of producing convincing digital text on demand, at essentially no cost, that assumption stops holding. A fabricated log and a real one can look identical on the page. Read carefully, and the method built up across the last several sections stops discriminating between the genuine article and something manufactured to look like it.
This is a reasoned argument grounded in a well-established fact: current systems can generate fluent, convincing text with ease. The core claim here is not that widespread evidence fabrication has already occurred, but that once fabrication becomes effortless, a whole category of evidence loses its trustworthiness—a conclusion drawn from logic, not from large-scale observation. It's worth stating plainly, though, why this deserves to be treated as the more serious of the two stress tests rather than a milder variation on the first. The audit-scope problem says an existing method doesn't reach far enough. This problem says that even a narrow, carefully scoped check, exactly the kind the earlier sections of this page recommended, can be defeated if the evidence it's checking against isn't real. What has to change isn't more rigor applied to the same kind of evidence. It's a shift toward a different kind of evidence altogether, ones where faking it costs something real: physical consequences that can't be typed into existence, sworn testimony carrying actual legal jeopardy for lying, live adversarial testing conducted in real time rather than reviewed after the fact, and cryptographic records establishing that a piece of evidence existed before anyone had the ability to fabricate its equivalent.
The same reasoning that applies to AI safety evidence applies here too. It applies, with little modification, directly to the profession this page has been using as its working comparison.
Modern financial audits depend heavily on documents that could be the target of exactly this problem: bank confirmations, invoices, contracts, system-generated transaction logs. The entire apparatus described earlier in this page, AS 2201 included, assumes that a document, once confirmed as authentic, reliably reflects the underlying transaction it describes. If AI systems can generate convincing fabricated confirmations, invoices, and logs at scale, that foundational assumption breaks for financial auditing in exactly the same way it breaks for AI safety evidence. This is an applied extension of the previous section's reasoned argument to a specific domain, not a new, separately demonstrated claim.
One part of the earlier list of fabrication-resistant evidence is worth naming as a concrete, forward-looking action a field like accounting could actually take now, rather than waiting for the problem to arrive fully formed. Financial and transaction systems could adopt hardware-rooted, cryptographically signed, timestamped record-keeping, establishing an unbroken chain of custody from the moment a transaction actually occurs, not merely from whenever it happens to be reported. This isn't a theoretical fix waiting to be invented. The Coalition for Content Provenance and Authenticity (C2PA) — backed by Adobe and others — already builds exactly this kind of cryptographic signing into real software today. The technology exists. What's missing is a profession willing to require it. That kind of record resists fabrication not because the document itself is more honest, but because forging it after the fact would require compromising the underlying infrastructure that produced it, not simply generating convincing text. A system built this way now becomes auditable under fabrication-resistant standards later. A system that waits faces a far harder retrofit once the capability to fabricate convincingly is already widespread.
This isn't hypothetical. In September 2026, the District of Columbia Court of Appeals struck a legal brief filed on behalf of Deutsche Bank after finding it cited four court cases that don't exist — invented by an AI research tool and never checked before the brief was filed. The lawyers involved represented one of the world's largest financial institutions. (D.C. Court of Appeals Strikes Deutsche Bank Brief Over AI-Generated Fake Cases, 2026) The failure wasn't a lack of resources. It was exactly the failure this page describes: confident, fluent, professional writing that nobody checked against a real source before relying on it.


Two things remain unresolved even after accepting everything argued above, and naming them plainly matters more than letting the page end on the comfortable impression that the problem has now been fully solved.
The first is a genuine trade-off, not a flaw to be engineered away. A system that shows its work, that can be checked, inspected, and its errors located, might in principle be somewhat less accurate than an opaque system that can't be inspected at all but happens to be right more often. The reasoned argument is that legibility should still be the governing consideration, not raw accuracy, because an inspectable system which fails visibly can have its failures found and corrected, while an opaque system which fails silently gives no one a way to notice when its accumulated correctness has quietly broken down. This claim is one that could, in principle, be shown false: if an opaque system's accuracy were sustained and independently verified across a long enough series of trials, that would genuinely challenge the claim as made here, not evidence being brushed aside.
The second is a distinction worth keeping precise, rather than letting the two blur together. Some uncertainty exists because checking simply stopped, time ran out, access wasn't available, or a tool wasn't used, even though a real, knowable answer exists somewhere in the world and could, in principle, be found. Other uncertainty is a different kind entirely: even with complete access to every piece of relevant evidence, a judgment call remains a judgment call, an evaluative weighing that no amount of additional checking would resolve, because it was never a factual question in the first place. Much of what gets described as appropriate humility, in this domain and elsewhere, is actually the first kind, an honest admission that the search simply ended, rather than the second, a genuine acknowledgment that no further search would have helped. Drawing a sharp line between these two kinds of uncertainty is essential: only the first can be dispelled by further evidence, while the second remains forever a matter of judgment, no matter how much verification is attempted.
Stay Sovereign.
Jim Germer
September 23, 2026
We use cookies to improve your experience and understand how visitors use our website so we can make it better.