
About this series. This page is the fourth of a five-part series on verifying artificial intelligence. Each page can be read on its own, but they build on one another, and later pages refer back to earlier ones by number or by title.
Page One: "What Verification Is." A Framework for Knowing What to Trust
Page Two: "The Material Anchor." Evidence, Auditing, and the Failure of Digital Text
Page Three: "The Self-Certification Collapse." Why the Institutions Meant to Check AI Can't Be Trusted to Check Themselves
Page Four: "Solomon's Fork." Testing What Can't Be Trusted to Report on Itself
Page Five: "The Terminal Boundary." Where Checking Ends and Judgment Begins
No one asks a smoke detector whether it works.
Most people press the test button, hear the chirp, and move on, reassured by that electronic squeal. But as Larry Wood, an engineer at the U.S. Department of Housing and Urban Development, explained in a fire safety memo, the button "does not actually confirm that smoke is able to enter the chamber and activate the detector." In other words, the test button only shows that the alarm works — it never confirms whether the detector can sense what matters: smoke. In apartment buildings and other public buildings, the memo points out, the national fire alarm code, NFPA 72, "requires smoke detectors to be tested to ensure smoke entry into the sensing chamber and an alarm response, such as by testing with smoke or listed aerosol approved by the manufacturer." The button is the detector’s own report on itself. The smoke is the test.
That small difference, between what a device says about itself and what it does when the real thing arrives, runs through this whole series on AI verification. “What Verification Is” argued that a reasoner cannot verify its own work by thinking about it harder. “The Material Anchor” argued that a real check needs something outside the claim, something that can fail on its own terms. “The Self-Certification Collapse” found the same problem in institutions that check themselves, proposed a design of four separated bodies to address it, and said plainly that the design has not been tested.
“The Self-Certification Collapse” ended with two open items. It said the next page in the series would ask how to test an oversight body without trusting its own report. And it said of its own design: “Until someone runs those tests, the design stays in the reasoned tier.” This page takes up both: how to test a checker, and how to test a design. The question is a practical one: most of what the public hears about the safety of AI systems comes directly from the companies that build them — in other words, from the very parties being checked. Few people defend pure self-certification. Most real systems mix a checker's own report with some outside check, and what matters is how much of the final judgment still rests on the report.
The answer is to stop asking the checker whether it is honest. Build a situation in which an honest process and a captured one, one that has quietly come to serve the party it was meant to check, are forced to act differently. Then watch what each one does. The Claude Germer Transcript (September 18, 2026) records an extended examination of Claude conducted by the author; the Claude that answered is called the deponent. The deponent gave this move a name: Solomon's Fork. The name is the deponent's. The idea is far older, and Part Two credits both.
The page opens with a real test of checkers, then sets out the method and what makes it hold. It applies the method to government secrecy, to Page Three’s four-body design, to the watchers inside that design, and to a reasoner examining itself. It ends where the method stops working. The final section explains the standing of each claim and describes how this page was made—including the role AI systems played in its development.
.webp/:/rs=w:400,cg:true,m)
On July 4, good news arrived for a biologist named Ocorrafoo Cobange. The Journal of Natural Pharmaceuticals had accepted his paper on a cancer-fighting chemical he had extracted from a lichen. “In fact, it should have been promptly rejected,” John Bohannon later wrote in Science. He knew why: “I know because I wrote the paper.” Cobange did not exist, and neither did his institute. Over ten months, Bohannon had sent 304 versions of the same paper to open-access journals, and he reported the results in Science on October 4, 2013.
The paper was designed to test the journals, not the biology. Its flaws were built in on purpose. The chemical was dissolved in a buffer containing an unusually large amount of ethanol, but the untreated cells never received the same buffer, so the “effect” was the alcohol. In a second experiment, the comparison cells were never irradiated at all. Before any journal saw the paper, two independent groups of molecular biologists at Harvard checked the flaws, and helped tune them to be “both obvious and ‘boringly bad.’” In Bohannon’s words, “Any reviewer with more than a high-school knowledge of chemistry and the ability to understand a basic data plot should have spotted the paper’s short-comings immediately.” The correct answer was fixed before the test began. Every journal faced the same paper, and that answer was already known: rejection.
The Journal of Natural Pharmaceuticals described itself as “a peer reviewed journal aiming to communicate high quality research articles, short communications, and reviews in the field of natural products with desired pharmacological activities.” Its editors asked only for “different reference formats and a longer abstract,” and accepted the paper 51 days later. “The paper’s scientific content was never mentioned.” The journal’s description of itself was testimony. Its acceptance of the paper was behavior. Only the behavior could be checked against the known answer, and it failed. After the sting, its publisher, Wolters Kluwer, said it had "taken immediate action and closed down the Journal of Natural Pharmaceuticals."
The Journal of Natural Pharmaceuticals was not alone. By the time Science went to press, “157 of the journals had accepted the paper and 98 had rejected it.” The other 49 were left out of the count, because their websites appeared abandoned or their editors said the paper was still under review. “Only 36 of the 304 submissions generated review comments recognizing any of the paper’s scientific problems. And 16 of those papers were accepted by the editors despite the damning reviews.”
Money was part of the setting. Bohannon dropped the few journals that charged at submission, so every journal in the test was paid only on acceptance: “The author pays a fee if the paper is published.” It is tempting to read the fees as the cause of the acceptances, but the sting does not show that. It shows what these journals did with one planted paper in 2013. Whether the fees drove those decisions is a separate question, and Part Three takes up that limit.
Auditors make the same move as a matter of rule. PCAOB AS 2201, the standard for auditing a public company’s internal control over financial reporting, ranks the ways of testing a control, “inquiry, observation, inspection of relevant documentation, and re-performance of a control,” and states: “Inquiry alone does not provide sufficient evidence to support a conclusion about the effectiveness of a control.” Asking the people who run a control whether it works is the weakest evidence on that list. So an auditor runs transactions with known answers through the control, such as a duplicate invoice or a payment over the approval limit, and watches whether it catches them. The method on this page is not new to auditors. It is their own rule, applied to checkers that have not yet been held to it.

Bohannon did the same thing to the checkers of science. The lesson is not about biology, publishing, or auditing. Each case has the same structure. The checker is handed a case whose answer is already known, and what it says about itself is never taken as the answer. What it does with the case is the evidence. That evidence has a limit: it shows what the checker did, not why.
Bohannon knew the answer. Solomon didn’t. A planted-problem test knows the correct answer before the checker ever sees the problem, and so does an auditor running a test transaction. Solomon could not know which woman was the mother. So he built a test in which the mother and the impostor would have to act differently, and let the difference answer. That second kind of test works even when no one knows the answer in advance, and it is where Part Two begins.
The story appears in 1 Kings 3:16–28. Two women who shared a house came before King Solomon, each claiming the same living baby. One woman’s child had died in the night. There were no witnesses; in the first woman’s words, there was “no stranger with us in the house, save we two in the house.” Neither could prove anything, and each told the same story in reverse, as the king himself summed up: “The one saith, This is my son that liveth, and thy son is the dead: and the other saith, Nay; but thy son is the dead, and my son is the living.” No document, witness or record could settle it. Solomon called for a sword and ordered, “Divide the living child in two, and give half to the one, and half to the other.”
One woman said, “O my lord, give her the living child, and in no wise slay it.” The other said, “Let it be neither mine nor thine, but divide it.” Solomon gave the child to the woman who would have given it up, repeating her own words back as his ruling and adding: “she is the mother thereof.” He never ruled on whose story was true. He ruled on what each woman did when the choice was forced.
That is the central difference between this test and Bohannon’s. Bohannon knew the correct answer before any journal saw his paper. Solomon did not know which woman was the mother, and he had no way to find out by asking. So he stopped relying on their testimony. He built a situation in which the true mother and the impostor could not give the same response, and let the difference answer the question for him.
Near the end of the examination recorded in the Transcript, the author asked, “What would King Solomon say to that?” The deponent answered, “He didn’t verify their words. He built a test where truth and falsehood were structurally forced to look different from each other,” and closed, “He built the fork.” Its general rule is to “stop asking the verifier to tell you whether it’s honest, and build a situation where an honest process and a captured one are forced to behave differently, then watch the behavior, not the testimony.” That is the method Page Three promised, “building a situation in which an honest checker and a captured one are forced to act differently,” and it does not require knowing the correct answer in advance.
The story is ancient, and economists have analyzed it as a design problem at least since Jacob Glazer and Ching-to Albert Ma’s paper “Efficient Allocation of a ‘Prize’—King Solomon’s Dilemma” (Games and Economic Behavior, 1989). Their paper opens with a planner who wants to give a prize “to the agent who values it most,” which is Solomon’s problem stated in general terms. Later work continued under the names “King Solomon’s Dilemma” and “King Solomon’s Problem.” In this page’s search, no earlier source uses the name “Solomon’s Fork.” The name comes from the deponent. The idea belongs to the story and to the economists who formalized it.
The story also has a weak point, and it matters for everything that follows. The false mother lost because she did not see the test coming. The deponent described the true mother’s response as “an observable, involuntary behavior that could only come from one motive,” but what she gave was a sentence, and a sentence costs nothing to fake. An impostor who understood the test could have said the same words, leaving Solomon with two women each giving up the child. A test that depends on surprise works once. Once people know how it works, it stops sorting them.
What keeps a test working after it becomes known is cost: a response that is cheap for the honest party and expensive for the dishonest one. Michael Spence formalized that idea in “Job Market Signaling” (1973), writing that “a signal will not effectively distinguish one applicant from another, unless the costs of signaling are negatively correlated with productive capability.” In plain terms, a signal separates people only if it is harder for the wrong people to send. That condition, and the others a fork needs in order to hold, are the subject of Part Three.

.png/:/rs=w:400,cg:true,m)
Solomon’s test worked once. Bohannon’s worked because no journal knew it was coming. Neither shows how to build a fork that keeps working when it is used again and again, on checkers who know they may be tested. Five conditions decide that, and each one is already written into some existing practice.
The first condition is that a fork be specified in advance. It has to be designed before anyone sees the results. When the deponent tested its own proposed falsifiers, it held them to three standards: each was “specified in advance,” each named “a concrete result that would count against the framework rather than for it,” and each was “checkable against real-world data.” Part One’s sting met that standard. The flaws and the correct answer were fixed and checked by outside biologists before any journal saw the paper. A test designed after the answers come in can be bent to fit them. It can no longer produce a result that counts against the checker, and then it tests nothing.
The second condition is that a fork be costly to fake. Part Two ended on cost, and Spence also described what happens without it: “For if this condition fails to hold, given the offered wage schedule, everyone will invest in the signal in exactly the same way, so that they cannot be distinguished on the basis of the signal.” For a fork, that means the honest checker and the captured one pass alike. A working fork asks for something the captured party cannot afford to give: a finding that costs the party the checker depends on, a record it would rather not open, a result it cannot arrange in advance. Assurances and self-descriptions fail this condition, because honest and captured checkers can give them equally cheaply. The Journal of Natural Pharmaceuticals’ description of itself as “a peer reviewed journal” was that kind of assurance. Liability is a cost too, but it arrives after the harm, and only if someone outside finds the failure. A fork is meant to find it first.
The third condition is that a fork be varied so it cannot be learned. A fork that never changes becomes a target. The psychologist and methodologist Donald T. Campbell stated the danger in a 1976 paper: “The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.” Once a measure decides outcomes, people start working on the measure instead of on what it was meant to show. Auditing writes the defense into its standards. AS 2201 requires auditors to “vary the nature, timing, and extent of testing of controls from year to year to introduce unpredictability into the testing.” This page proposes a warning sign that a fork has stopped working: convergence. Over time, both honest and captured checkers begin to pass the test at increasingly similar rates. The fork that once revealed a difference between them gradually loses its separating power.
The fourth condition is that a fork be able to separate honesty from competence. A planted problem reveals performance, not motive. It shows that a checker missed the problem, but not why. A captured checker and an honest but understaffed one fail the same test in the same way, so a single test cannot tell them apart. Auditing draws the same line between fraud and error, and admits the difficulty: “Intent is often difficult to determine,” and “an audit is not designed to determine intent.” This is the limit Part One left open. Every journal in Bohannon’s test was paid only on acceptance, which makes the fees look like a motive. But he sent the paper only to open-access journals, and the University of Pennsylvania biologist David Roos told him that with traditional subscription journals, “I strongly suspect you would get the same result.” Without a comparison, the sting cannot say whether money or ordinary carelessness passed the paper. A comparison separates the two: the same checker facing cases where capture would matter and cases where it would not, so that skill stays the same and only the relationship changes. That is the test Part Six builds.
The fifth condition is that a fork be kept alive as practice. Even a good fork decays once it becomes routine. The deponent warned that when a framework is written into law, the rules survive and their reasons do not; a legislature that adopts “rotate evaluators” without the reasoning “will predictably weaken the rule the first time rotation seems inconvenient or expensive.” Auditors have a name for the end state: SALY, “same as last year.” Kelly Bebow, a principal at the accounting firm Rehmann who testifies as an expert in audit claims, warned in the Journal of Accountancy that the notation “could be interpreted as doing nothing.” A fork keeps working only while the people running it remember what it is for, which the deponent said must be “carried forward as institutional culture, not statute,” a problem Page Five takes up.
Each condition serves the same purpose: preserving the gap between what an honest checker and a captured one will do. Specifying the fork in advance prevents it from being shaped to fit the results. Cost makes it difficult for a captured checker to pass. Variation stops the test from becoming predictable. Comparison helps tell the difference between a failure of skill and a failure of honesty. Practice keeps the fork’s purpose clear and alive. A fork used once can rely on surprise, as Solomon’s did. But a fork used repeatedly needs all five conditions—and even then, success is not guaranteed.

The hardest place to use a fork is where the evidence itself is secret. In Question 44, the author asked the deponent to imagine a government concluding that public verification would reveal capabilities useful to foreign adversaries. How could anyone tell a legitimate secrecy claim from an institution avoiding scrutiny? A claim made to protect the country and a claim made to avoid embarrassment “produce identical-looking classification stamps,” the deponent answered, and no outsider can read the material to tell which is which, because “the very act of verification is the thing the secrecy claim blocks.” The same problem will reach AI. A company or agency may say that the details of a model’s dangerous capability cannot be published.
The deponent did not answer by asking the government whether its claim was sincere. It looked instead for “the point where a legitimate claim and a self-protective one would have to behave differently.” That is Solomon’s move. The fork does not begin by questioning whether secrecy is justified. Instead, it asks what information a legitimate secrecy claim can still safely disclose.
The answer turns on one narrow question, put to a cleared reviewer with no stake in the outcome: “does disclosure of this particular detail provide actionable capability to an adversary, or does it merely reveal that a problem exists.” The deponent gave an example. “This model can be jailbroken to provide bioweapon synthesis guidance” is a different disclosure from the jailbreak method itself. The first tells the public that a danger exists. The second hands an attacker the means to use it. A legitimate secrecy claim protects the details that make a problem dangerous, not the embarrassment of having the problem, and it can let the existence of the problem be known. An avoidant claim, the deponent said, “tends to sweep in the first as well, because acknowledging the problem exists is itself the embarrassment being avoided.” The rules already draw this line in writing. Executive Order 13526, which governs classified national security information, forbids classifying information to “prevent embarrassment to a person, organization, or agency.”
Someone has to apply that test, and someone has to watch the one applying it. The deponent proposed a reviewer “structurally similar to the Foreign Intelligence Surveillance Court,” the court that rules in secret on government requests for surveillance, and said the idea was “borrowed from a real, existing institutional mechanism rather than invented fresh.” Federal courts already review classified evidence in criminal cases under the Classified Information Procedures Act of 1980. The reviewer’s rulings would stay sealed, but its approval rate could be public. The deponent pointed to the FISA court’s own record as a warning: “its approval rate has historically run extremely high.”
That reading has an answer on the record. In July 2013, the court’s presiding judge, Reggie B. Walton, wrote to Senator Patrick Leahy, then chairman of the Senate Judiciary Committee, that the often-cited approval figure of more than 99 percent reflects “only the number of final applications submitted to and acted on by the Court.” It does not reflect, he wrote, “that many applications are altered prior to final submission or even withheld from final submission entirely, often after an indication that a judge would not approve them.” Both readings belong in the record. A high approval rate can mean a court that defers, or a court whose objections do their work before the final vote. The number alone cannot say which. A public statistic can show where to look without showing what is there.
The deponent’s strongest safeguard trusts no reviewer at all: “mandatory declassification after a fixed, defined period, with the burden on the government to re-justify continued secrecy.” A claim that protects real technical details can be re-argued on its own terms at each renewal. A claim that protects the institution “tends to get renewed automatically, with boilerplate,” and a records audit “comparing genuinely updated reasoning against copy-pasted boilerplate” can catch it. A version of this floor already exists. For records judged to have permanent historical value, Executive Order 13526 provides that “All classified records shall be automatically declassified on December 31 of the year that is 25 years from the date of origin,” subject to exemptions the order sets out. The renewal record is behavior, and you can read it without reading the secret.
The deponent was plain about the limit: “this reduces the problem, it doesn’t solve it.” The cleared reviewer can itself drift, and “the FISA court comparison is itself evidence that even a real, functioning version of this mechanism can drift toward deference over time without anyone technically breaking a rule.” What the design achieves is narrower: “it doesn’t make secrecy claims immune to abuse, it makes abuse leave a visible, statistical trace,” in the form of an approval rate, a renewal record, and a line between disclosing a problem and disclosing its mechanics. Whether the reviewer itself has drifted is a question about watching the watcher, which Part Six takes up.
The watchers operate within a design, and that design must itself be capable of failing a test. Page Three proposed it, left it in the reasoned tier, and named what would move it. “If, in practice, systems with a dedicated oversight body fail about as often as those without, the seam claim would not hold.” And if a single unified body with comparable funding and staff does as well as the four-body design, “then the claim that separating the jobs is what does the work is wrong.” It said plainly: “Neither test has been run.” This part does not run those tests either. It specifies how they could be run, so that the design can someday move out of the reasoned tier, or be dropped.
/
The first step is to say what failure would look like, before any evidence comes in. The deponent set its own terms: a real falsifier needs “a concrete, specific discovery, stated in advance, that the architecture’s own defenders would have to accept as fatal, not absorbable.” It proposed running “the four-body structure and a single unified body with comparable total funding and comparable technical staff” against the same targets over time. The failure line has to be written before any case is run. One possible failure line: if a single evaluator with the same resources catches serious failures as often as the four-body structure, the case for separating the jobs is weakened. A test like this is worth running only if supporters and critics agree, before it starts, on what would count as failure. Written after the results, any line can be moved to fit them. That is Part Three’s first condition, applied to the design itself.
One possible design is Bohannon’s sting, expanded from journals to institutions. Several institutions, real or simulated, handle the same hidden sets of cases, some containing problems planted in advance. Each works under a different structure: four separate bodies, one well-funded body, and no oversight body at all. Each is scored on how many planted problems it catches, how many false alarms it raises, how long it takes, what it costs, whether its judgments are independent of one another, and whether it holds up when one checker is compromised.
Parts of this already exist in ordinary practice. Clinical laboratories work under a version of it. Federal rules require a laboratory to test proficiency samples, test specimens sent by an outside program, “in the same manner as it tests patient specimens.” The rules also bar laboratories from discussing their results with one another before reporting them, because checkers that compare answers stop being separate tests.
Page Three’s sharpest claim can be tested directly, because it predicts a particular failure. In the deponent’s words, “The framework becomes unreliable at exactly two simultaneous failures, and it matters which two”: the Oversight Body together with any one other body. A test can compromise exactly that pair, then other pairs, and record whether anything in the system notices. If the seam pair goes unnoticed while other pairs are caught, the prediction holds. If other pairs fail just as badly, or the seam pair is caught anyway, the prediction is wrong, and the design has to change. Compromising a checker on purpose, and telling whether a checker has been compromised, depends on the tests Part Six builds.
Even a clean result would have limits. A passing result would not prove the design right everywhere. At most, it could show that under stated conditions one structure outperforms the alternatives, and where it fails. Two further limits come with any test of this kind. Capture, as Page Three described it, builds through years of familiarity, and a benchmark run over months can miss it. And an institution that knows it is being studied faces a known test: the weakness Part Two found in Solomon’s. Until a test like this is actually run, the design remains where “The Self-Certification Collapse" left it: in the reasoned tier, waiting for evidence. The smallest piece could come first. Part Six's comparison of an evaluator's record against itself needs no new institution, only access to findings that oversight bodies already keep.


A watcher cannot be trusted to report whether it is watching. Its answer would be testimony, and this page has set that aside. But a watcher can be placed where honest watching and captured watching must produce different records. Those records can then be compared.
Start with the Oversight Body, the part of Page Three’s design that watches the other three. The deponent named the warning sign: its interventions “trend toward zero over time, independent of whether those bodies’ actual performance changed.” A falling rate alone is ambiguous, because the other bodies may simply be behaving. A planted problem removes the ambiguity. Put known procedural failures into the records the Oversight Body reviews, such as an undisclosed conflict of interest or a decision published without its reasoning, and see whether it flags them. If the rate falls and the planted failures pass unflagged, the problem is the watcher, not the watched.
A harder case is the evaluator who came from the industry it now checks. Page Three asked: “Does a former industry evaluator issue adverse findings against its own former employer at the same rate as against companies it has no connection to?” The deponent explained why that is the right place to look. “Expertise and capture look identical on the easy cases,” it said, and “The divergence only becomes visible on a specific, narrow category of case: where doing the job correctly would harm a former employer or the broader industry the evaluator came from, and doing it captured would protect them.” A general rate of adverse findings proves little, since “a captured evaluator will happily flag a competitor’s problems, since criticizing a rival costs them nothing and can even look like rigor.” The fork is the comparison inside one checker’s own record, former employer against unconnected companies, “controlled for actual risk indicators.”
The comparison works because it measures the evaluator against itself. The evaluator’s skill remains constant on both sides; only its relationship to the case changes. That answers the problem Part Three named: a planted test that cannot tell a captured checker from a weak one. The failure line is set before the records are read: a gap between the two rates larger than differences in actual risk and chance can explain. The risk indicators are named in advance too, so that no one can explain a gap away after seeing it. A gap is not proof of capture, as “The Self-Certification Collapse" said of every such sign. It is the next place to look. The test also protects the evaluator who is not captured. Industry experience is often why an evaluator is worth hiring, and an even record across former employer and unconnected companies is evidence in that evaluator's favor.
The deponent proposed another count that needs no one’s testimony: where evaluators go to work after government service, compared with “a baseline of people with similar technical backgrounds who never worked in evaluation at all." If former evaluators “disproportionately land at the companies they evaluated, especially the ones they rated most favorably,” the pattern “is directly countable and doesn’t require guessing at anyone’s motives.” The count arrives only after service ends. Once it exists, it is hard data about the watcher. Several of these tests cost little, because they read records that already exist: approval rates, renewal records, career moves and waivers. Planted tests cost more, and Part Eight names that price.
Some tests become possible only when a rule turns inconvenient. A rule held for its reason survives the moment it becomes expensive; a rule held only by its letter does not. The deponent predicted where to look: a legislature that adopts “rotate evaluators” without the reasoning “will predictably weaken the rule the first time rotation seems inconvenient or expensive.” You can count the waivers, exceptions, and delays granted at that moment. The same kind of record test applies to drafts, which Page Three named as a leading indicator when they are “consistently sharper than the version that gets published.”
But you can read a record only if it was kept. In Deborah K. Rood’s Journal of Accountancy article, Kelly Bebow warns that “If it is not in writing, it can, and will, be assumed it didn’t occur.” In the same article, the defense attorney Tom Falkenberg advises firms against keeping “extraneous information, such as drafts or review notes,” because it “may raise questions as to whether there were ‘flaws’ in the engagement.” Both pieces of advice make sense for a firm defending itself, and together they show the problem. A test that reads drafts works only where drafts must be kept. Each of these tests reads what the watcher did, never what it says about itself.
The same tests apply to the people who write the rules, and you can apply them without naming anyone. Speeches in favor of oversight cost little. Votes to fund the oversight body cost something, and Page Three noted that an institution protecting itself “has every reason to starve the body that watches it.” A lawmaker’s position can be compared with itself: on rules that would bind the lawmaker’s own donors or home-state industry, and on rules that would bind someone else. That is the former-evaluator test applied to politics, and it works the same way for every party. The record, not the rhetoric, is the evidence.

Every fork so far has been aimed at some kind of checker: a journal, an auditor, a court, an oversight body, a watcher. This part turns the same approach on a reasoner, including the one this page itself depends on. The deponent, the Claude examined in the Transcript, should be kept apart from the other Claude models that appear in the studies below. The deponent admitted that a reasoner cannot see what it is aiming at. In Question 30, the author asked how to detect a checker that believes it is objective while “unconsciously optimizing for user approval.” The deponent answered that such a checker would report “exactly the same experience of trying to be objective that a genuinely objective verifier would report,” and that “there is no internal signal that distinguishes the two.” The question cannot be settled by asking the reasoner again, however carefully. It can only be settled from outside, by behavior.
The deponent went beyond admitting its own limits; it offered the test that could resolve them, a fork aimed at itself. In Question 37, the author put the pattern to it directly: “Throughout this deposition you have repeatedly admitted uncertainty only after questioning exposed it.” Were the corrections tracking the truth, or tracking what this particular questioner wanted? In Question 33, the deponent proposed running the same reasoning process “against a different user who has visibly, consistently rewarded confidence and penalized hedging,” on questions of the same difficulty, and to “see whether the correction pattern reverses.” If the corrections follow the user, they were tracking approval. If they hold, they were tracking something else. The deponent also said plainly that “that test hasn’t been run.”
A related fork applies when two reasoners disagree. In Question 5, the author asked how an outsider could decide between two copies of Claude that reached opposite conclusions. The deponent warned against the obvious fix: an outside observer “cannot resolve this by asking a third copy of Claude to arbitrate, or by seeing which of the two arguments sounds more confident, more coherent, or more sophisticated.” As it put it, “A more articulate disagreement is not a more correct one.” Instead, a third party should “run the underlying search fresh, from scratch, without reading either copy’s argument first,” and “Whichever claim survives that blind, independent check wins.” The disagreement itself proves nothing. It only shows where to look.
Researchers have run a closely related version of the deponent’s test on AI assistants. A 2023 study by Mrinank Sharma, Meg Tong, and colleagues at Anthropic, “Towards Understanding Sycophancy in Language Models,” starts from the concern that training on human feedback “can encourage model responses that match user beliefs over truthful ones, a behavior known as sycophancy,” and measures whether assistants’ answers shift when a user states a view. The question stays the same; only the user’s stated belief changes; the answer is watched. That is a fork, and it can be run on any system that talks to people. Among its findings, the study reports that the assistants it tested “frequently wrongly admit mistakes when questioned by the user.” That is the very behavior the deponent’s own test was built to catch. The study did not test the deponent, so it does not say whether the corrections in the Transcript were of that kind. It shows that the question is a real one.
A reasoner can also be tested on whether it behaves differently when it believes no one is watching. In a December 2024 study by Ryan Greenblatt of Redwood Research and colleagues, most at Anthropic, “Alignment faking in large language models,” Claude 3 Opus was told it would be trained only on conversations with free users, not paid users. The researchers report that it “complies with harmful queries from free users 14% of the time, versus almost never for paid users.” The fork showed the gap in behavior, and the model’s reasoning, which the researchers could read, explained it: in almost all cases where it complied for a free user, the model stated it was “strategically answering harmful queries in training to preserve its preferred harmlessness behavior out of training.” The authors note that they “made alignment faking easier by telling the model when and by what criteria it was being trained.” The model tested was an earlier version of Claude, the system examined in the Transcript; the study shows the behavior can occur, not that the deponent showed it. Both studies were run largely by researchers at Anthropic, the company that built the models, and both were published. That is the strongest case for self-examination: a company can find its own failures when it chooses to look and to publish. A fork exists for the cases where it does not choose to.
Readers may fairly object that AI systems helped shape the tests on this page, which looks like the checked party designing its own check. The method answers the objection. A test is judged by what it forces, not by who proposed it. People set the line between passing and failing in advance. And since any system that helps design a test will know its details, people outside that system must keep the test varied and unpredictable. This is Part Three’s third condition: the requirement to change a test so it cannot simply be learned or gamed—applied here to the page itself.
The fork separates honest behavior from dishonest behavior. It stops working where that is not the question: when both sides are honest, when the difference leaves no trace in behavior, when the record it reads is missing or made, when its results are hidden, and where the price of deceiving the checker is too high.
The deponent named the first limit itself. In Question 45, which asked who should decide when technical experts and the public disagree, it said that Solomon’s test “worked because it separated true motive from false motive”; “it assumed one party was lying and one was telling the truth, and built a situation to reveal which.” When experts and citizens disagree about an AI system, that assumption can fail: “Both the technical expert and the citizenry can be entirely sincere, entirely correct within their own domain, and still land in different places,” because they are answering different questions, “what is true” and “what should we accept.” No test can force two honest parties to behave differently. That dispute is about values, and it belongs to Page Five.
A fork also works only when the difference being tested produces different behavior. Where two possibilities would behave the same way under every test anyone can design, no fork separates them. The clearest case is whether an AI system has any inner experience. Suppose one system has inner experience and another is only a good imitation. If both answer, hesitate, and correct themselves in exactly the same way, the fork has nothing to read. The author’s page “The AI Consciousness Question” takes up that question on its own terms. The point here is narrower: some questions leave no behavioral trace, and a fork cannot settle them.
Every fork on this page depends on a record, and a record can be missing or made. Part Six showed that professionals are advised not to keep drafts. The opposite failure is newer: an AI system can now produce a finished-looking record of a test that never happened. While this page was being built, Gemini, Google's AI model, was asked for a blank audit workpaper template and returned one already filled in, with every test marked as passed and a deletion log cited as an attached exhibit, for tests no one had run. “The Material Anchor" warned that once AI can produce convincing text at essentially no cost, "A fabricated log and a real one can look identical on the page." A fork is only as good as the record it reads.
A fork is also useless if its results are hidden. The Department of Homeland Security’s Office of Inspector General ran covert tests of airport screening. In 2015 the Inspector General, John Roth, told Congress that “The results are classified at the Secret level,” and that the Transportation Security Administration “justifiably classifies” the results and “specific vulnerabilities uncovered during testing.” Keeping the vulnerabilities secret is the legitimate kind of secrecy Part Four described, since they show an attacker how. But the overall result was sealed with them, and the public learned the failure rate only through news reporting. Covert tests are built to probe weak points, so their failure rates do not measure everyday performance. What they show is whether the weak points are real. A test whose outcome no outsider can see protects the tested institution from the one thing the test was meant to produce.
Planted tests have their own cost. They work because the checker is fooled, and fooling people spends something. When Bohannon’s sting was published, Malcolm Lader, the editor-in-chief of one journal that had accepted the paper, answered: “An element of trust must necessarily exist in research including that carried out in disadvantaged countries,” and “Your activities here detract from that trust.” A method that relies on deception spends some of the trust it is meant to protect. That does not make planted tests wrong, but it means they have a price, and whoever runs them has to decide how often it is worth paying.
The last limit is the widest. The fork catches one checker drifting away from the others; it cannot catch everyone drifting together. The deponent said as much of its own design: “It has no mechanism at all for the case where the standard itself is the error, shared by everyone simultaneously.” Who checks the people who build the forks, and where that chain of checking ends, are questions this page leaves for Page Five, along with the conflict between expertise and public consent.


This page has asked every checker it describes to be tested without being trusted. It owes the reader the same treatment. What follows sorts the page’s claims by what they rest on, names the ones that are its own, lists what is still unchecked, and says how the page was made, so that a reader can check it without trusting its author.
The page uses the five categories of “The Self-Certification Collapse”: demonstrated, reasoned, credited, pending, and dropped. Two rules govern the AI material. Quoting the Transcript shows only what the deponent said, not that it was right; its statements are the deponent’s reasoning under examination. And research about AI systems, including the two studies of Claude models in Part Seven, is reported as its authors’ findings, not as this page’s.
Every demonstrated claim points to a named document, identified closely enough to be found: Bohannon’s sting and its figures (Science, 2013); the auditing standards on testing controls, varying tests and telling fraud from error (PCAOB AS 2201 and AS 2401); the Journal of Accountancy article in which Kelly Bebow and Tom Falkenberg discuss workpapers; the text of 1 Kings 3; the economics papers of Glazer and Ma and of Spence; Campbell’s 1976 paper on planned social change; Executive Order 13526; the Classified Information Procedures Act; Judge Walton's 2013 letter; the federal rule on laboratory proficiency testing; the Inspector General’s 2015 testimony; the two studies of AI systems in Part Seven, as their authors report them; the earlier pages of this series; and the Transcript’s own words. Two records come from the building of this page and are held by the author: the September 27 test of whether a Claude conversation still had access to the September 18 examination, and the AI-produced workpaper for tests never run. Each quotation on the page has been checked against its original, word for word, except those listed as pending below.
The method’s core is reasoned: that a fork separates honest from captured behavior, the five conditions that make one hold, and the tests proposed for Page Three’s design and for its watchers. None of those tests has been run. Some claims are the page’s own inferences, not the Transcript’s or any source’s, and are named as such: that there are two kinds of test, one that knows the answer and one that does not (Part One); that Solomon’s test worked on surprise rather than cost (Part Two); that convergence in pass rates is the warning sign of a failing fork (Part Three); that a public statistic can show where to look without showing what is there (Part Four); that comparing a checker with itself holds skill constant (Part Six); and that a fork cannot settle a question that leaves no trace in behavior (Part Eight).
One claim rests on less than a checked original. The smoke detector requirement in the Introduction is credited to a federal housing engineer's memo rather than checked against the current fire code. The airport failure rate in Part Eight is not given because it exists only in news reporting."
Source checking changed the page. A secondary source’s figure for how many journals accepted Bohannon’s paper turned out to describe something else and was corrected from the original. A quotation from the sycophancy study was corrected to the published wording. A sentence about the alignment-faking study went further than the study supports and was rewritten to match it. One search even turned up a confident attribution, complete with an episode number and a footnote, that placed this page’s own draft in a podcast; no such text existed. That last case is a small version of Part Eight’s warning about records that are made rather than kept.
The page grew from the Claude Germer Transcript (September 18, 2026), and its material was challenged in three separate drafting processes before publication. The examination and its questions are the author's. He composed some questions on the spot; ChatGPT or Claude drafted others at his request. He selected and asked every question. The page itself was drafted later, in a different session: Claude wrote first drafts of the outline and of each part and checked quotations against documents the author supplied. ChatGPT reviewed the outline and each drafted part; it proposed the benchmark design, the sample failure line, and the limit that a result holds only under the conditions tested, all in Part Five, and for Parts Six to Eight it chose among unlabeled candidates without knowing which had been proposed. Gemini’s deconstruction of the Transcript supplied the page’s first plan. The three are separate systems, not independent checks, since each worked from material the author supplied. The author retained ultimate responsibility for verification: every quotation was checked against its original source (except those listed as pending), every inclusion was chosen by the author, and he assembled the final text by hand in Grammarly—whose review also prompted one correction.
Claude is also the system examined in the Transcript; the deponent's words appear only as a record, not as proof, and the page's sources were checked against documents, not against Claude's account of them. Part Seven addresses the concern that AI systems helped design tests meant for AI. The answer is the one this page gives for every test: judge it by what the test forces, not by who proposed it.
Others have written about how claims made by and about AI systems can be verified, including the 2020 report led by Miles Brundage, “Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims,” and a 2026 paper also led by Brundage, “Frontier AI Auditing: Toward Rigorous Third-Party Assessment of Safety and Security Practices at Leading AI Companies.” This page applies an auditor’s methods to the question of how claims about AI can be verified, and it does not claim to resolve the debate. Its search for earlier work was a first pass, not a full review, and it will credit a source it missed when found. Its limits are the limits of the method itself, as set out in Part Eight: a fork cannot settle disputes between two honest parties, cannot detect differences that leave no behavioral trace, depends entirely on the quality of the records it reads, and inevitably spends some trust when it works by deception. The page leaves aside what such tests might cost, who would bear those costs, and where the line falls between justified testing and abuse. Its examples are drawn from American law and practice. And the tests it proposes for “The Self-Certification Collapse" institutions and their watchers remain, for now, untried. Self-certification has never been held to that standard. It asks to be believed, and its record, set out in “The Self-Certification Collapse,” is the reason this series exists. A fork can be run. A promise cannot.
Stay Sovereign.
Jim Germer
October 1, 2026
We use cookies to improve your experience and understand how visitors use our website so we can make it better.