
About this series. This page is the third of a five-part series on verifying artificial intelligence. Each page can be read on its own, but they build on one another, and later pages refer back to earlier ones by number or by title.
Page One: "What Verification Is." A Framework for Knowing What to Trust
Page Two: "The Material Anchor." Evidence, Auditing, and the Failure of Digital Text
Page Three: "The Self-Certification Collapse." Why the Institutions Meant to Check AI Can't Be Trusted to Check Themselves
Page Four: "Solomon's Fork." Testing What Can't Be Trusted to Report on Itself
Page Five: "The Terminal Boundary." Where Checking Ends and Judgment Begins
No public company in the United States is allowed to audit its own books. No student grades their own final exam. No referee plays for one of the teams. Most of us understand, almost without thinking about it, that a check performed by the party being checked is not really a check at all.
And yet some of the most consequential industries of our time have come to check large parts of their own work. It rarely happens by announcement. It happens gradually, and it tends to become visible only after something goes wrong.
The first two pages in this series asked how a single claim can be trusted. “What Verification Is” argued that a reasoner cannot verify its own work simply by thinking about it harder. “The Material Anchor” argued that real checking needs something outside the claim, something that can fail on its own terms. This page asks the same question at a larger scale. What happens when the checker is an institution that is checking itself? And if that arrangement cannot be trusted, what would it take to build a check that can? The last part of this page turns those same questions back on itself and reports where each of its claims stands.
On October 29, 2018, Lion Air Flight 610 took off from Jakarta, Indonesia, and crashed into the Java Sea thirteen minutes later, killing all 189 people on board. On March 10, 2019, in strikingly similar circumstances, Ethiopian Airlines Flight 302 crashed six minutes after leaving Addis Ababa, killing all 157. Both airplanes were new Boeing 737 MAX jets. Both crashes involved MCAS, software designed to push the airplane’s nose down automatically in certain flight conditions. In all, 346 people died.
The House Committee on Transportation and Infrastructure, which investigated the crashes for eighteen months, found that Boeing made "fundamentally faulty assumptions" about MCAS and other critical technologies. The more lasting lesson, though, concerns who was responsible for checking Boeing's work. Much of the work of certifying the 737 MAX as safe was done not by the Federal Aviation Administration's own staff, but by Boeing employees authorized to act for the FAA. The committee described them as “Boeing employees who are granted special permission to represent the interests of the FAA and to act on the agency’s behalf in validating aircraft systems and designs’ compliance with FAA requirements.” In a 2016 internal Boeing survey of those employees, conducted while the 737 MAX was being certified, 39 percent of those who responded said they perceived “undue pressure.” The committee’s conclusion was direct: “Excessive FAA delegation to Boeing has eroded FAA’s oversight capabilities.”
It would be easy to read this as a story about one careless regulator. The record supports a more useful reading. Delegation began for real reasons. As the Department of Transportation’s Inspector General explained, “Recognizing that it is not possible for FAA employees to oversee every facet of such a large industry, Federal law allows the Agency to delegate certain functions to private individuals or organizations.” The FAA oversees almost 292,000 aircraft and nearly 1,600 approved manufacturers. The engineers who understood the 737 MAX best worked at Boeing. The expertise, the staff, and the speed were all on the company’s side, so the agency handed pieces of the checking to the company, one reasonable decision at a time.
Piece by piece, the checking moved inside the company until the regulator no longer understood what it was approving. The Inspector General found that the “FAA did not have a complete understanding of Boeing’s safety assessments performed on MCAS until after the Lion Air accident.” The House committee found that the FAA officials who approved Boeing’s request to remove references to MCAS from pilot training material “remained unaware of the redesign of MCAS until after the crash of the Lion Air flight.” No one ever announced that Boeing would certify its own airplane. No one had to. The agency had given them the keys.
If this sounds familiar, it should. “What Verification Is” argued that a reasoner reviewing its own work is not getting a second opinion; it is the same doctor looking at the same chart again. “The Material Anchor” argued that a check is worth something only when it adds a second, independent point of failure. Self-certification is the same problem at the scale of an institution. Auditors already have a name for it: the self-review threat, one of the threats to independence named in the AICPA Code of Professional Conduct. When the checker and the checked are the same actor, there is only one place an error can be caught, and it is the same place the error was made. This part of the argument is reasoned rather than measured: it follows from what makes a check independent, the same logic the first two pages built. None of it requires anyone to be dishonest. A sincere, competent self-check is still not an independent one.
The reasons that pulled the FAA toward delegation apply with at least as much force to artificial intelligence, because much of the expertise needed to evaluate frontier AI systems sits inside the companies that build them. OpenAI offers a clear, documented example. Its Preparedness Framework gives an internal Safety Advisory Group the job of reviewing a new model’s risks and recommending next steps to the company’s leadership. The framework states that “the members of the SAG and the SAG Chair are appointed by the OpenAI Leadership.” The group advises; it does not decide. The same document states that “OpenAI Leadership can also make decisions without the SAG’s participation.” In plain terms, OpenAI's leadership picks its own safety reviewers and can decide without them.
A second example shows how a safety commitment can rest on self-report alone. In July 2023, OpenAI announced that it was “dedicating 20% of the compute we’ve secured to date over the next four years” to the problem of aligning future superintelligent systems. It was a public and specific promise, and no one outside the company had access to confirm whether it was kept. What happened to that promise belongs to Part Three. The point here is narrower: a commitment that only the committed party can verify is a form of self-certification, however public it is.
None of this shows that every AI company works this way, and this page makes no such claim. It shows that the structure the 737 MAX investigation condemned is present, in documented form, at one of the world's most prominent AI companies. The fuller record of who holds deployment authority at OpenAI, Anthropic, and Google DeepMind, and what each has promised, is laid out on this site's earlier page, "The AI Alignment Committee Exists. Who's On It?"
The House committee had a name for where the FAA ended up: "grossly insufficient oversight by the FAA—the pernicious result of regulatory capture." That word, capture, returns later on this page. For now, it is enough to see how the road there began. Delegation is not capture. It is one of the ways capture becomes possible.

For sixteen years, until it filed for bankruptcy in December 2001, Enron retained Arthur Andersen as its auditor. On paper, Andersen looked exactly like an independent checker. It was a separate firm, with its own partners, professional licenses, and reputation, and it was one of the five largest accounting firms in the world. The federal indictment later filed against the firm describes the rest of the relationship. Andersen "earned tens of millions of dollars from Enron in annual auditing and other fees." Enron "was one of ANDERSEN's largest clients worldwide, and became ANDERSEN's largest client in ANDERSEN's Gulf Coast region." Andersen also "performed both internal and external auditing work for Enron." At a Senate hearing in December 2001, the committee's chairman put numbers to it: "Arthur Andersen claims they were paid $25 million for its accounting services and $27 million for its consulting services." A later Senate investigation found that Enron's board had allowed Andersen "to provide internal audit and consulting services while serving as Enron's outside auditor.
"What happened when Enron collapsed belongs to Part Three. The point here is the structure that existed before anything went wrong. Andersen did not need to be dishonest for that structure to matter. Dependence creates predictable pressure, not inevitable corruption. But anyone looking at the arrangement could ask a simple question: what would Andersen stand to lose by telling Enron something Enron did not want to hear?
That question is the center of this part, because separation is not independence. A checker can have its own name, its own license, and its own office, and still depend on the party it checks. Dependence takes several recognizable forms: the fees it is paid, the future business it hopes to win, the access it needs to do its work, or a relationship it would rather not damage. Part Four sets out how to measure them. For now, the test is simpler: what matters is not whether the checker is separate, but what it stands to lose by saying no. A checker has to be able to survive disagreement.
The accounting profession has official terminology for this, and it is worth borrowing exactly. The AICPA's framework for independence separates two ideas. Independence of mind, often called independence in fact, is "the state of mind that permits the performance of an attest service without being affected by influences that compromise professional judgment." Independence in appearance is avoiding circumstances that would cause "a reasonable and informed third party, having knowledge of all relevant information, including safeguards applied, to reasonably conclude that the integrity, objectivity, or professional skepticism of a firm or a member of the attest engagement team had been compromised.
"No outsider can see into a checker's mind, but anyone can see who pays the checker, who hires it, and who can walk away from it. It is tempting to treat independence in appearance as cosmetic, a matter of looking proper. Read through the logic of this series, and you'll see the opposite: it is the only part of independence anyone else can check. That reading is reasoned, not quoted; the code states the two standards, and this page draws the conclusion. Verification can examine structure even when it cannot examine motives.
The same structural question appears again in AI. Outside evaluation organizations exist, and they are genuinely more independent than a company's own safety group. They are separate organizations with their own staff and published findings. METR, one of the best-known, has evaluated models from several major labs before release. Its own report on OpenAI's o3 and o4-mini models, published in April 2025, is candid about the terms of that work. "METR received access to earlier checkpoints of o3 and o4-mini from OpenAI three weeks prior to model release." "We did not access the model's internal reasoning, which is likely to contain important information for interpreting our results." "This evaluation was conducted in a relatively short time." The evaluator was separate. Timing, scope, and the information available for review depended on what the company chose to provide.
Funding raises the same question from another direction. In July 2026, Google DeepMind's chief executive, Demis Hassabis, proposed a new body to test frontier AI models for dangerous capabilities before release. As Axios reported it, the body would be "funded by the industry," and labs would "initially share their models with the body voluntarily, up to 30 days before release." A body funded by the companies it evaluates, and dependent on their willingness to participate, can have excellent staff and still be exposed to the same pressure.
None of this is a finding that any evaluator has failed. METR's candor about its own limits is part of what makes its work useful. The Andersen question applies unchanged: what does this checker stand to lose by delivering bad news. If a checker the company pays cannot be relied on, the natural next thought is that something larger will catch the problem: the market, or the government. That is where Part Three begins.
Companies that check their own work have a ready answer to the first two parts of this page: the market would punish us. It is not a foolish argument, and it deserves to be stated at full strength. A company caught in a provable lie about safety would stand to lose its valuation, its best researchers, its enterprise customers, and its standing with the governments that regulate it. In artificial intelligence especially, such a lie would be hard to keep. Former employees leave and talk. Independent researchers publish. Rival labs have every commercial reason to expose a competitor's failure. This defense is a large part of why self-certification can feel safe to the people doing it. A page that waved it away would not be taking the problem seriously.
The defense is right that markets punish. It is wrong about when they punish, and about what they are able to see.
Start with when. The market ultimately delivered a decisive verdict on Enron. In November 2001, Enron announced that it would restate its earnings, and less than a month later it filed for bankruptcy. Arthur Andersen was convicted of obstruction of justice in 2002 and did not survive as a firm, even though the Supreme Court overturned the conviction in 2005. Congress responded with the Sarbanes-Oxley Act, which became law on July 30, 2002. Every one of those consequences was real, and every one of them came after the harm, once the hidden numbers were out. A punishment that arrives after the loss is accountability, not prevention.
The second limit concerns visibility rather than timing. A market can react only to what it can see, and a company that certifies itself controls much of what there is to see. OpenAI's 2023 pledge of 20 percent of its compute for superalignment, described in Part One, was public. Whether the pledge was being kept was not. The answer surfaced only when people inside the effort left. In May 2024, the team's co-leader, Jan Leike, resigned and wrote that "Sometimes we were struggling for compute and it was getting harder and harder to get this crucial research done." Four days later, Fortune reported that "a half-dozen sources familiar with the Superalignment team's work said that the group was never allocated this compute." The team was disbanded. The market learned what it learned because insiders chose to talk. That is the very channel the defense relies on, and here it worked late and by chance, not by design. And once the gap was visible, the response was not a penalty. In October 2024, less than five months after those reports, OpenAI announced that it had "raised $6.6B in new funding at a $157B post-money valuation." Even then, the size of the gap stayed hidden for almost two more years, until The New Yorker reported in April 2026, citing four people who had worked on or closely with the team, that it had received roughly 1 to 2 percent of the company's compute.
These are two cases, and the claim here stays at their size. To this page’s knowledge, no one—including this page—has systematically searched for cases in which a safety failure drew a swift market penalty. What the two cases show is narrower, and still important: the market reached Enron only after the damage was done, and it learned about the compute gap only because people inside chose to speak. The full record of the pledge, the reporting, and the silence that followed is laid out on this site's page "The Superalignment Record.”
If the market arrives late, government might seem the natural answer. The Claude Germer Transcript (September 18, 2026), an extended recorded examination of Claude conducted by the author, is the source this page draws on most. It runs to more than fifty numbered questions, many written to test or break an earlier answer, and Claude checked facts against primary sources as it went, correcting itself on the record when a claim did not hold up. It was taken under no oath. It records Claude's reasoning, which this page quotes and credits, and checks it against outside records wherever it can. Its first question asked Claude to name every entity in the world with the legal authority to audit a frontier AI company's safety claims before a new model is deployed: "not advise, not recommend, but actually compel a change." Later, it took each traditional mechanism of democratic accountability in turn. The point is not that any of them is failing. It is what they share. Courts need an injury or a legal claim before they can act, so they review deployments that have already happened; the State of Florida's complaint against OpenAI, discussed on the first two pages of this series, is a court case about systems already in public use. Elections come years apart and bundle thousands of issues together. Legislatures write the rules, but they do not review individual models. Inspectors general audit what an agency has already done. The press reports what has happened, or what someone chose to disc
lose, and has no power to compel a halt. Every one of these mechanisms acts after deployment.
The closest thing to a check before release is the European Union's AI Office. Since August 2, 2026, according to the European Commission, its powers over the providers of the most advanced models include "requesting information, requesting access to a model for evaluations, requiring risk mitigation measures, and issuing fines of up to 3% of global annual turnover or requesting a provider to restrict the making available on the market, withdraw or recall the model." Those are major powers, and they are new. But the Transcript described the office's authority as ongoing compliance monitoring, not a gate a model must clear before release.
Underneath it all is a simple distinction. Prevention and accountability are different institutional functions. Courts, elections, legislatures, inspectors general, the press, and markets were built for accountability. They are not failing at prevention; they were never designed for it. The Transcript put it plainly: stopping a deployment in advance is "a genuinely novel institutional problem that none of the existing tools were built to solve." Handing the job to a government agency does not solve it by itself. It "just moves the same open question to a different address," with an agency's staff making an unverified judgment that is checked only afterward. A check that acts before deployment has to be built on purpose.
If that check has to be built, the first question is how to tell a genuinely independent checker from one that only looks independent. That is where Part Four begins.


Eleven questions into the Transcript, a habit had set in. Independence kept being described as something a checker either had or lacked, in expressions such as "no external party" and "internal versus external." Then came a question the earlier answers had not faced: is independence really a yes-or-no property at all? The answer corrected the earlier language. Independence, the Transcript conceded, is a spectrum, and its own earlier line, "no independent verification exists," had been "a binary compression of a continuous variable." The Transcript named what it had done: "collapsing a graduated property into a clean yes/no for the sake of a cleaner sentence." That is the move Page One calls Complexity Damping, turning up in a new place. It is also the right place for this part to begin, because the correction is on the record rather than asserted after the fact.
What replaced the binary view was a series of questions, each focused on an aspect visible to an outsider: Who selected the checker, and who holds the authority to remove it? A checker that can be dismissed by the party it examines holds its position on that party's sufferance. Who pays for it? Does it get access by right, or only by invitation that can be withdrawn? Can the party being checked edit or suppress the checker's findings before anyone else sees them? And has any finding the checker issued ever actually cost the checked party something? A checker that has never faced real resistance may be independent on paper and nowhere else. None of these questions requires reading anyone's mind. Each one looks at structure or at a record, which is what Part Two concluded verification can do even when it cannot examine motives.
None of these dimensions is new, and the page should say so plainly. The Transcript built its list by reasoning from its own examples. Asked later how it knew the list was complete, it admitted that "no field that already studies verifier independence" had been checked against it. Both of the obvious fields had covered this ground long before. Federal law makes a public company's audit committee "directly responsible for the appointment, compensation, and oversight" of its outside auditor, which answers the questions of who selects and who pays. The same law bars an auditor from providing most consulting services to a client it audits, which is Part Two's Andersen problem written into statute. For government agencies, the legal scholars Kirti Datla and Richard Revesz have argued that agencies "fall along a continuum ranging from most insulated to least insulated from presidential control," judged by features such as protection from removal, fixed terms, and control over their own budgets. The Transcript arrived at a scale others had already built. What it added was applying that scale to the checkers of AI.
The same question about completeness turned up something the list had missed. The Transcript found the gap on the spot, and the gap was time: "A verifier who reviews a company once and moves on has different incentives than one in an ongoing, renewable relationship with the same company." A checker that works with the same company year after year slowly loses its edge, even when every other safeguard holds. "Nobody has to be captured on purpose for this to happen." Auditing already has a name for this: the familiarity threat, listed in the same AICPA code that names the self-review threat from Part One. It also has a rule for it. Under the Sarbanes-Oxley Act, a firm may not audit a public company if its lead audit partner, or the partner reviewing the audit, "has performed audit services for that issuer in each of the 5 previous fiscal years." The rule applies however well the partner has performed. It assumes nothing about anyone's character; it assumes only what time does to a working relationship. And the sixth dimension surfaced only because someone asked the question designed to break the list, the falsification discipline Page One borrows from Karl Popper, at work on this page's own material.
The dimensions do not always improve together. Sometimes strengthening one weakens another. The European Union's AI Office shows how. It was "established within the European Commission," and the AI Act gives the Commission, acting through the office, "the ability to conduct evaluations of GPAI models, request information and measures from model providers, and apply sanctions." The Transcript judged it "a government body whose staffing and mandate don't depend on the companies it regulates," and by the Transcript's measure it scores well on who selects the checker, who funds it, and whether its access is compelled. But the same standing authority means an ongoing relationship with the same small group of frontier labs, with no end date. The Transcript noted the tension in its earlier praise: the continuous authority it counted as a strength is "the same structural feature that creates a familiarity threat over a long enough horizon." That is a structural exposure, not a finding that the office has softened; the Commission's enforcement powers over general-purpose models took effect only in August 2026. Compelled access has its own limit. It gets a checker "in the door," in the Transcript's phrase. It does not, by itself, make the checker willing to "say the hard thing once inside.
"Independence is not a single property. You have to judge it dimension by dimension, never sum it into a single label. The useful question is never "Is it independent?" It is "Independent in which respects, and how much?"
Knowing how independence is measured shows what a checker has to be built to withstand. The next question is how to build one.

Part Three ended with a requirement: a check that acts before deployment has to be built on purpose. The most obvious design people reach for is a single strong regulator, one agency with the power to approve or block a new AI system. The Transcript rejected it. A single "AI czar," it warned, would be "one point of unchecked, unverifiable judgment." That is self-certification moved to the top. The same body would weigh the evidence and decide what to do about it, and nothing would check that body until after the fact.
The principle that replaces it is short: "The person deciding whether evidence justifies action should not be the same person or institution that produced the evidence." It is not a new principle, and the page credits its source. Montesquieu stated the classical form in 1748: "To prevent this abuse, it is necessary from the very nature of things that power should be a check to power." The Transcript also borrowed a working model from financial regulation, where "no single agency runs the whole system." The Transcript added one move that makes the rest of the design work. A decision made under real uncertainty cannot be verified before it is made, so the design stops trying to check the content of each decision and instead constrains the structure around it: who holds power, for how long, and what happens if the power is abused. The Transcript called this a shift "from checking content to constraining structure.
"The design that follows has four bodies, and each has one job.
The Evaluator does the technical testing. Its access to training data and internal safety communications is compelled, not granted by the company, which answers the problem Part Two found with outside evaluators whose access depends on what a lab chooses to provide. "Its only job is producing findings. It has no power to act on them.
"The Decision Body is separate from the Evaluator. It reviews the findings and decides whether to approve, delay, or block a deployment. Its members serve fixed terms and rotate out on a schedule, which answers the familiarity threat from Part Four. Every decision is published with its reasoning, so a finding cannot be quietly set aside. Separating finding from deciding — and linking them through a public record — addresses a point the Transcript made in a single sentence: "Verification without enforcement authority is diagnosis without treatment."
The Oversight Body never touches the technology. The Transcript calls this body the Overseer; this page calls it the Oversight Body, and quotations preserve the Transcript's wording. Its only job is to check whether the other bodies followed their own procedures and whether they show signs of capture. It can remove members of the other bodies, and it has its own protection from removal, "so that capturing the Decision Body doesn't also capture the body meant to catch that capture." That answers the problem Part Three left behind, where handing verification to a government agency simply moved unchecked judgment to a different address.
The Rotating Comparator is a separate team with no connection to the others. Periodically it re-evaluates the same system the Evaluator has been reviewing, to catch the kind of drift an Evaluator would not report about itself.

The design's sharpest idea concerns how people are removed. A decision made under uncertainty often cannot be proven wrong, even afterward, because no one can see what would have happened had a launch been stopped, or allowed; in the Transcript's words, "the counterfactual is unobservable." A rule that removes decision-makers only for provably bad decisions "is therefore nearly useless." What works instead is removal for failures that can be checked: "did they follow the stated procedure, did they disclose their reasoning, did they recuse when conflicted." Public disclosure is what makes all of this usable. Rotation, removal, and an Evaluator's dissent each depend on a visible record of what was decided and why.
Four bodies on an organizational chart are not automatically four independent checks. A Rotating Comparator staffed with new people, but using the Evaluator's own methods and reading the Evaluator's own inputs, will repeat the Evaluator's errors rather than catch them. The question that prompted this part of the Transcript put the distinction exactly: "methodological independence rather than numerical redundancy." Adding checkers who work the same way, the Transcript answered, "only helps against failures that are independent and random across attempts." Real independence between checkers requires a different method, a different way of failing, no shared source feeding them, and no checker seeing the others' conclusions before reaching its own. The word-count failure told on Page One is the small version of the same lesson. Redundancy is not independence. More checkers using the same methods and the same evidence can catch a random slip, but not a shared blind spot. For that, a checker has to be able to fail differently.
None of this makes the system immune to failure. What it does, in the Transcript's words, is make abuse "require multiple separate, differently-vulnerable failure points to break at once, rather than one." It remains a design grounded in reasoning rather than demonstration: no one has built or tested it, and Part Seven returns to what would prove it wrong.
Even a careful design erodes. Independence built into a system does not maintain itself; it wears away gradually, often without anyone deciding to let it go. The next question is how anyone might recognize that happening.
Site Content — Postscript, added September 26, 2026: As this page was being completed, the design this section critiques stopped being hypothetical. President Trump announced plans for an "AI Force," modeled on the Space Force, and a new AI czar to run it. The announcement offered no specifics on the office's authority, budget, or place in government. Trump's own stated standard was that the government would not "hinder or stifle" AI's growth, and would rely on the existing criminal and civil justice system to catch "bad" behavior after it happens. He has called AI safety concerns a "hoax." This paragraph is demonstrated, not speculative — the reporting is dated, public, and cited below. What it demonstrates is not this page's argument being proven right in the abstract; it is a real proposal that, point for point, matches the design the Transcript rejected: no fixed term, no compelled access to evidence, no separation between finding a problem and deciding what to do about it, and no body positioned to remove the czar if the czar is the one who fails.
Sources: Axios, "Trump wants a new AI czar and an 'AI Force' modeled on Space Force" (Sept. 19, 2026); CNN, "Trump proposes 'AI Force' with AI czar" (Sept. 19–20, 2026); NBC News, "Trump says he's creating an AI force and appointing a czar" (Sept. 19, 2026); CNBC, "Trump says Scott Bessent won't be AI czar" (Sept. 25, 2026)


Part Four showed that a checker can lose its independence without anyone deciding to let it go. That is the discouraging half of the story. The encouraging half is that the same fact also makes the loss observable. In the Transcript's words, "Capture, unlike pure judgment, leaves a trail." A single judgment call cannot be checked against anything, but a checker's behavior over years can: "A judgment call has no baseline to compare against. A verifier's track record does." Parts Two and Four looked at structure, meaning who pays, who hires, and who can walk away. This part looks at behavior. Between them, they cover the two things an outsider can observe. Motive is the one thing no one can.
One set of signs appears in the checker's own output over time. Adverse findings against one company decline year after year, while findings against comparable companies do not, and nothing about the first company's actual risk has changed. Problems are still found at the same rate, but they are described more gently: "an area for continued attention" where the checker once wrote "a material deficiency." Recommended fixes drift toward what the company had already planned to do. People and money move between the checker and the company it checks, which can be counted as hard numbers. Federal audit rules already address one version of that movement. Under the Sarbanes-Oxley Act, a firm may not audit a public company if the company's chief executive, chief financial officer, controller, or chief accounting officer worked on that company's audit for the firm "during the 1-year period preceding the date of the initiation of the audit." The strongest sign of all comes from comparison. A fresh checker, with no history with the company, reviews the same target, and the two sets of findings are laid side by side. The Transcript called this "the strongest test" and "close to a direct measurement of drift." It is where Part Five's Rotating Comparator earns its place in the design.
None of these signs proves anything on its own. The Transcript was careful to say so: each is "not proof of capture." Each is a specific, measurable change that justifies a closer look, and none of them depends on the checker's account of itself.
Output-based signs share one weakness: they arrive late. By the time a finding softens, the softened finding has already been issued. The Transcript therefore looked further upstream for "the leading indicator, not the trailing one." Findings get released on a schedule that tracks the political or news calendar rather than the date the technical work was finished. Internal drafts are consistently sharper than the version that gets published. Technical staff lose resources relative to communications staff before any finding has changed. An assessment stated plainly inside the institution becomes hedged by the time it reaches a hearing or a press release. These signs can warn before the damage is done, but they come with a cost. Each one "requires access to material that isn't normally public": drafts, staffing records, and internal assessments that no institution volunteers.
One objection has been waiting since Part One, which noted that the expertise needed to evaluate AI sits largely inside the companies that build it. If the only people qualified to check the labs have worked for the labs, why not simply exclude them? The Transcript's answer was that excluding them would not produce independent evaluators; it "produces incompetent ones." An evaluator who does not understand what it is evaluating is not independent in any meaningful sense, only unqualified. Independence pushed far enough costs competence. So the answer cannot be exclusion. It has to be measurement, aimed at the one kind of case where genuine expertise and capture behave differently. Does a former industry evaluator issue adverse findings against its own former employer at the same rate as against companies it has no connection to? Does it argue for weaker standards specifically in the specialty it came from? Where does it go to work when its public service ends? The Transcript called expertise "necessary and dangerous at the same time." The general method behind the first test, building a situation in which an honest checker and a captured one are forced to act differently, is the subject of Page Four.
Every sign so far assumes someone is on the other end: a company whose findings soften, a former employer, a job waiting after government service. But a checker can drift with no one capturing it at all, by quietly coming to serve itself: its budget, its reach, its survival. The Transcript called this "in some ways worse," because there is no one to trace the money or the people to, and the revolving-door tests find nothing. The idea comes from economics. The economist William Niskanen built a model of government bureaus, first published in 1971, that "assumes that bureaucrats act to maximize the budget of their bureau." Even this form of drift leaves a trail, just a different one. Real risk goes up and down, because the world does. A checker serving its own budget reports risk that always justifies the same funding or more, year after year, whatever is actually happening, and that one-way relationship between stated risk and requested money can be tracked across many budget cycles. The measures of success change too. A checker that began by asking whether it prevented real harm starts reporting how many reviews it completed, because "activity is easier to demonstrate to a legislature or the public than outcomes."
This last form of drift points somewhere uncomfortable. An institution protecting itself has every reason to starve the body that watches it. The Transcript said that self-serving drift and the Oversight Body’s vulnerability "aren't separate problems. They're the same problem, reinforcing each other."
You can measure each of these signs, and none of them requires trusting the checker. But each one requires someone outside the relationship to track it for years, with access to records "neither party has an incentive to volunteer." In the design proposed in Part Five, that role belongs to the Oversight Body.
Which raises the question the whole design rests on: what happens when the Oversight Body itself fails?
Part Six ended with a question: what happens when the Oversight Body itself fails? The Transcript's answer was unusually precise. It did not argue that the design was generally strong or weak. It counted.
A single failure remains contained. If the Evaluator issues a flawed finding, two safeguards intervene: the Oversight Body, which checks whether procedures were followed, and the Rotating Comparator, which independently reviews the same system and may reach a different result. If the Decision Body errs, the Oversight Body can flag any departure from process, and because the Decision Body must publish its reasoning, outsiders can respond even if the Oversight Body does not. As the Transcript puts it, “One captured Evaluator, with the other three bodies functioning, is a contained failure.” This is the purpose of building four bodies rather than one: no single failure undermines the entire system.
Two failures are different, and the Transcript was specific about which two: "The framework becomes unreliable at exactly two simultaneous failures, and it matters which two." If the Evaluator and the Oversight Body fail together, findings go soft and nothing is watching for exactly that. The Decision Body sees only what the Evaluator reports, and the Rotating Comparator reviews the system only on a schedule, one the design never specified precisely. Between comparisons, two failed bodies can operate unseen. If the Decision Body and the Oversight Body fail together, the mirror image happens. Accurate findings are ignored or distorted, and no one is left to flag that the Decision Body abandoned its own process.
Three or four failures are not a meaningfully higher bar than two, because once the Oversight Body is gone, the system has already lost its only means of noticing the others failing. In the Transcript's words, "a captured Overseer doesn't just remove one of four protections, it removes the mechanism that would have revealed the other three failing too." The Oversight Body is "the check on the check." The design's protection is not evenly spread across its four bodies; "the system's redundancy isn't symmetric." This is the oversight seam. One case the Transcript did not address is the Oversight Body failing alone. Following its logic, which is this page's inference and not the Transcript's statement, nothing breaks at once in that case, but the next failure anywhere else goes unseen.
A failing Oversight Body would not announce itself. It would leave a measurable sign. The Transcript named it: the Oversight Body’s interventions against the other bodies "trend toward zero over time, independent of whether those bodies' actual performance changed." That can be tracked like any approval rate, with the same timing, rate, and comparison tools Part Six described. The likely cause is ordinary rather than dramatic. Part Six ended on it: an institution protecting itself has every reason to starve the body that watches it. The Transcript's remedy was structural. The Oversight Body’s funding and standing have to be made "structurally independent enough that the institutions it watches can't starve it the same way they'd starve any other rival claim on a shared budget." The Transcript did not say how that could be done, and this page does not know of a design that has settled it.
The practical lesson for anyone building such a system is to "put the Overseer function first, not last," and to "resource the Overseer as if it's the most important body, not the cheapest one to add later." The watcher is the easiest role to underfund, because its work is invisible when it is going well. The accounting profession has already lived one version of this. After Enron, Congress created the Public Company Accounting Oversight Board, which describes itself as "a nonprofit corporation established by Congress to oversee the audits of public companies." It is the clearest real case of a society adding a watcher over the checkers. A group of AI governance researchers proposed the same kind of body for frontier AI in January 2026, a PCAOB-style "auditor of auditors" for the firms that audit AI companies. It is not a pure Oversight Body in this page's sense, because it also writes auditing standards and brings enforcement actions, jobs the four-body design keeps apart. And it has its own watcher: "The SEC has oversight authority over the PCAOB, including the approval of the Board's rules, standards, and budget." Who watches that watcher is a question this page names and leaves for Page Five. Funding the Oversight Body first does not make the design safe. It protects the design's most exposed point; it does not remove the exposure.
The Transcript supported its advice with a sweeping claim this page will not repeat as written: "Every real institutional collapse — Enron's auditors, the FAA's self-certification before the 737 MAX, the SEC before 2008 — traces back to the watcher role being underfunded, understaffed, or structurally dependent on the thing it was watching." That is a universal statement built on three examples, the same kind of overreach the Transcript caught in itself elsewhere. The record supports a narrower conclusion. Each of the three involved a watcher that failed. Arthur Andersen depended on the client it audited, as Part Two showed. The House committee found that "Excessive FAA delegation to Boeing has eroded FAA's oversight capabilities." And in 2008, the SEC's own Inspector General found that the program overseeing Bear Stearns "failed to carry out its mission." Three cases cannot establish "every," and in the SEC case the record points to missed warnings rather than to a lack of money.
The claims at the center of this part may be mistaken, and this page indicates where. If, in practice, systems with a dedicated oversight body fail about as often as those without, the seam claim would not hold. If a single unified body with comparable funding and staff does as well as the four-body design at resisting capture and drift and at avoiding error, then the claim that separating the jobs is what does the work is wrong. Neither test has been run. The Transcript was direct about what that means: the design "has only ever been challenged by reasoning about its structure," and "Its survival so far is not yet meaningful confirmation."
Two questions lie beyond this page. Page Four asks how to test an oversight body without trusting its own report. Page Five asks who checks the Oversight Body, and why that question never fully ends.
What this page can leave with the reader is a question that travels. The Transcript found that this design's protection is uneven, that one body's failure matters more than the others' because it hides everyone else's. This page cannot claim that every design has such a point; that would be a generalization from a single case, and no one has tested it. But the question works on any system of checks, whether an audit regime, a safety agency, or a lab's internal review: whose failure would stop us from seeing everyone else's? If there is a clear answer, that is where the design is most exposed, and where money and scrutiny belong. If there is not, the protection may be genuinely even, and that is worth knowing too. Find the seam before it finds you.
In the end, the seam "can't be closed by better institutional design alone," because "something always has to watch the watcher."
This page has asked every institution to report honestly where its claims stand. Part Eight does the same for this page.


This page has asked every institution in it the same questions: what it has shown, what it has argued, and whose work it builds on. The page owes the same answers about itself. The page sorts its claims into three tiers and two statuses.
A claim is demonstrated when it reports something that happened and a reader can check it against a record, either a primary document or the Transcript itself. A claim is reasoned when it follows from argument rather than from a test or a record. A claim is credited when the idea began with someone else and the page names its source. A claim is pending when it rests on an outside source that has not yet been checked against the original. A claim is dropped when it failed that check and was removed. Claims that are removed are noted here for transparency.
A single rule sorts every claim on this page. Quoting the Transcript shows only what the Transcript said, not that it was correct. Its statements are Claude’s reasoning, and the page treats them as such. When Part Seven quotes the Transcript saying the design “becomes unreliable at exactly two simultaneous failures,” the quotation is demonstrated, but the underlying point remains an argument. Four claims are specific to this page: that delegation is not capture but one way capture becomes possible (Part One); that independence in appearance is the only part outsiders can check (Part Two); what follows when the Oversight Body fails by itself (Part Seven); and the question Part Seven asks the reader to carry to other systems. These are the page’s own inferences—reasoned claims based on its argument, not on the Transcript.
Its proposed design rests on argument, on foundations laid in other fields, and on one recorded examination. Start with what happened. The historical record on this page is demonstrated, and each piece of it points to a document a reader can open. The 737 MAX crashes, the FAA’s delegation to Boeing employees, and the committee’s conclusion that delegation “eroded FAA’s oversight capabilities” come from the House committee’s final report and the Department of Transportation’s Inspector General. OpenAI’s Safety Advisory Group, and its leadership’s freedom to decide without it, come from OpenAI’s own Preparedness Framework, and the compute pledge comes from OpenAI’s announcement. Andersen’s fees and its dual role at Enron come from the federal indictment, a Senate hearing, and a Senate investigation; the fee figures are Andersen’s own, as the committee chairman stated them. METR’s terms of access come from METR’s own report, and the industry-funded testing proposal comes from Axios’s account. Jan Leike's resignation statement, Fortune's report on the compute, and OpenAI's funding announcement are each quoted from their source, and the 1 to 2 percent figure comes from The New Yorker's April 2026 report. So are the Sarbanes-Oxley rules on audit committees, consulting, partner rotation, and cooling-off periods; the PCAOB’s description of itself and the SEC’s authority over it; and the SEC Inspector General’s finding on Bear Stearns. The European Commission’s description of the AI Office’s powers is also demonstrated, but it is the newest item in the record and the most likely to change.
The Transcript’s own corrections belong in this group too. It called its earlier yes-or-no language “a binary compression of a continuous variable,” and it found the missing dimension of time only when asked to break its own list. Those claims are about what the Transcript did, and the record shows it.
The record also has edges, and the page should mark them. What it shows is sometimes narrower than it might sound. Part Three rests on two cases, and to this page's knowledge, no one, including this page, has systematically searched for others in which a safety failure drew a swift market penalty. Part One shows the delegation structure at one AI company, not at all of them. Part Seven declined to repeat the Transcript’s claim that “every” institutional collapse traces back to a weak watcher; three cases cannot support “every,” and in one of them the record points to missed warnings rather than a shortage of money. Some statements that read like part of the record are the Transcript's interpretation. That the AI Office's staffing and mandate do not depend on the companies it regulates, and that its authority amounts to ongoing monitoring rather than a gate a model must clear before release, are the Transcript's descriptions, not findings this page checked; so is the comparison to financial regulation's division of labor. An AI system's knowledge of recent events is only as current as its training or its searches that day, which is one more reason the page quotes the Commission, not the Transcript, for what the Office can do. No claim is pending. One claim was dropped. Goodhart’s Law, the idea that a measure stops measuring once it becomes a target, fits Part Six’s point about checkers that report activity instead of outcomes, but this page could not check its wording against an original source, so Part Six makes the point without it.
This page draws on multiple fields and credits each source where its ideas appear, wherever possible. The principle of separation of powers is from Montesquieu (1748). The concepts of self-review and familiarity threats, and the distinction between independence of mind and independence in appearance, are from the AICPA's code of professional conduct. The rules governing who selects and pays an auditor, what services an auditor may provide, how long a lead partner may serve, and restrictions on employment after an audit come from the Sarbanes-Oxley Act, which has governed public-company audits since 2002. The idea that agencies exist along a continuum of independence is from legal scholars Kirti Datla and Richard Revesz. The model of a bureau seeking to maximize its own budget is from economist William Niskanen. The requirement that a real check must be able to fail on its own terms is drawn from the first two pages of this series. Each source addressed the problem of checking power long before the arrival of AI.
Now to what is proposed. The page’s core is reasoned: the rule that whoever produces the evidence should not decide what it justifies, the four bodies, removal for failures of process rather than of outcome, the difference between more checkers and independent ones, the signs of capture and drift, and the oversight seam. None of it has been built, and none of it has been tested. Part Seven named the two tests that could show it wrong. One asks whether systems with a dedicated oversight body fail less often than systems without one. The other asks whether a single body with comparable funding and staff does as well as four separate ones. Neither test has been run. Part Four’s list of dimensions is reasoned as well. The Transcript built it by reasoning from its own examples, and this page compared it with auditing and administrative law only afterward. What the page adds is bringing these structures to the checkers of AI, as others have also begun to do, and counting where such a design breaks. The Transcript said the same of its own work: “a translation, not a discovery.”
The page leaves several questions open, and naming them tells the next reader, and the next page, where to start. Some concern the page's own checking. The Transcript's list of dimensions was not checked against the full research literature on verifier independence; this page compared it with two legal traditions, not with every relevant field. The search for earlier proposals of this kind, which found published work applying audit structures to AI, was a first pass, not a full review. The test for a missing dimension is the one that surfaced time: ask what a checker could depend on that none of the six questions would catch.
Others concern the design itself. A blind spot shared by all four bodies cannot be caught by any of them, and that problem belongs to Page Five. Neither this page nor the Transcript has addressed how to fund an oversight body that cannot be starved by the institutions it watches. The design assumes the evidence it examines is genuine; the problem of fabricated evidence raised in "The Material Anchor" applies to the Evaluator as much as to anyone. A check that acts before deployment in one jurisdiction does not constrain a lab in another; the page identifies this collective-action problem but does not solve it. Full checking is never free, so any real system must decide which elements warrant the most scrutiny, a choice that, as the Transcript notes, is itself an unchecked judgment. What a preventive gate would cost to run, and how much it would slow deployment, are implementation questions beyond this page's scope.
That is where the page stands. Its account of what happened rests on records a reader can open. Its proposed design rests on argument, borrowed foundations, and one recorded examination, and it has not yet met the evidence that could break it. Until someone runs those tests, the design stays in the reasoned tier.
Stay Sovereign.
Jim Germer
September 26, 2026
Quotations from the Claude Germer Transcript, September 18, 2026. © 2026 Jim Germer. All rights reserved.
We use cookies to improve your experience and understand how visitors use our website so we can make it better.