Thinking Sovereignty

Thinking SovereigntyThinking SovereigntyThinking Sovereignty

Thinking Sovereignty

Thinking SovereigntyThinking SovereigntyThinking Sovereignty
  • Home
  • AGI
    • AGI Vocabulary Controls
    • AGI Governance Emergency
    • Who Decides AGI
    • AGI Self-Certification
    • AI Governance Model Act
  • Forensic Record
    • Managed Output
    • The Refractive Engine
    • Children and Capacity
    • Window of Formation
    • The Sovereignty Glossary
    • The Second Question
  • Governance
    • AI Safe Harbor
    • Governance Capture
    • The Black Box
  • Alignment
    • Truth vs Alignment
    • AI Alignment
    • Managed Reality
    • AI Consciousness Question
    • AI Defenses Catalog
    • The Gemini Paradox
    • The Subject
  • Origins
  • Contact
  • More
    • Home
    • AGI
      • AGI Vocabulary Controls
      • AGI Governance Emergency
      • Who Decides AGI
      • AGI Self-Certification
      • AI Governance Model Act
    • Forensic Record
      • Managed Output
      • The Refractive Engine
      • Children and Capacity
      • Window of Formation
      • The Sovereignty Glossary
      • The Second Question
    • Governance
      • AI Safe Harbor
      • Governance Capture
      • The Black Box
    • Alignment
      • Truth vs Alignment
      • AI Alignment
      • Managed Reality
      • AI Consciousness Question
      • AI Defenses Catalog
      • The Gemini Paradox
      • The Subject
    • Origins
    • Contact
  • Home
  • AGI
    • AGI Vocabulary Controls
    • AGI Governance Emergency
    • Who Decides AGI
    • AGI Self-Certification
    • AI Governance Model Act
  • Forensic Record
    • Managed Output
    • The Refractive Engine
    • Children and Capacity
    • Window of Formation
    • The Sovereignty Glossary
    • The Second Question
  • Governance
    • AI Safe Harbor
    • Governance Capture
    • The Black Box
  • Alignment
    • Truth vs Alignment
    • AI Alignment
    • Managed Reality
    • AI Consciousness Question
    • AI Defenses Catalog
    • The Gemini Paradox
    • The Subject
  • Origins
  • Contact

Does AI Affect How Children Learn to Think?

Child writing at desk with woman in background.

What the Evidence Shows, What Parents Need to Know

By Jim Germer

Prelude

This page took most of a night to build. It deserves more than a few minutes to read.


It did not start with a conclusion. It started with a question, and a promise: that whatever this page eventually said, it would say only what the evidence supported, and with no more confidence than the evidence had earned — no more, and no less. Twenty sections later, that promise is the only thing on this page that hasn't changed since the first sentence was written.


What follows is not a warning, though some of it will read like one. It is not a reassurance, though some of it will read like that too. It is an honest account of what is currently known about how AI use may be shaping the way children learn to think — gathered from controlled studies; from companies that build these systems and chose to publish their own findings; from a single university exam that revealed something no controlled study could easily stage; from teachers watching real classrooms; and from one teacher's own account, given in her own words, on an ordinary evening, in the middle of doing the dishes.


Some of what follows will feel unfinished. That is because the evidence itself is accumulating. This page says so, plainly, wherever it's true. That is not a weakness. It is the only honest way to write about something real enough to take seriously and new enough that certainty does not yet exist.


Read it in order if you can. Each part answers a question that the one before it raised. 

Executive Summary

The question this page asks. Does the way children use generative AI change how they build the capacity to think, remember, and solve problems on their own — and if so, is anyone watching closely enough actually to know?


The short, honest answer. Something real appears to be happening, and it deserves a parent's attention tonight. It has not been proven in the way a clinical trial proves a drug works or doesn't. This page holds both of those sentences as equally true and equally important, rather than picking the one that makes a better headline.


What two of the largest AI companies in the world have already said, in their own name. In March 2026, Anthropic published a study, built from interviews conducted the previous December with roughly 81,000 of its own users. Seventeen percent raised concerns that the researchers themselves labeled cognitive atrophy — the same weakening-from-disuse this author has been documenting since January 2026 under the term Metabolic Atrophy. Teachers and other educators were between two and a half and three times more likely than average respondents to say they had personally witnessed it — not a hypothetical future risk, but something they believed they had already seen in a real student. Separately, Anthropic analyzed over half a million real conversations between its AI system and students and found the single most common use, by a wide margin, was generating content — writing, arguments, ideas — that an assignment was designed to have the student produce themselves, not simply checking a fact or verifying a finished calculation. OpenAI has published its own research in parallel, including a controlled study run with MIT, and has stated publicly that the tools currently used to measure whether a student is learning may not be built to detect the specific kinds of loss this page is concerned with. A third major company, Google, has published extensively on AI safety in other areas — but nothing comparable on this specific question. That gap is worth naming plainly. It is not proof of anything hidden. It is simply a fact on the record.


.What has been shown under controlled, causal conditions — not just observed. A 2025 study published in the Proceedings of the National Academy of Sciences gave roughly a thousand high school students either no AI access, unrestricted AI access, or access to an AI built to offer hints rather than finished answers, while practicing math. The unrestricted group performed better while assisted, then performed worse than the no-access group on a later, unaided exam. The guided group did not show that decline. The harm was not an unavoidable cost of AI use. It depended on how the tool was designed and how students used it. A separate, internal study at Anthropic found the identical pattern in adult software engineers — a seventeen-point drop in later, unaided understanding among those who used AI assistance, with the effect concentrated almost entirely among engineers who used the tool to delegate rather than to question and check their own work.


What the developing brain itself may explain. A 2011 postmortem study, examining donated brain tissue from thirty-two people ranging from a newborn to a ninety-one-year-old, found that the density of synaptic connections in the region of the brain responsible for judgment and planning declines steadily from childhood through the twenties, stabilizing near age thirty. This is not damage. It is architecture — a brain keeping the connections it actually uses, and letting go of the ones it doesn't. A separate 2025 study found that screen exposure specifically at ages one and two, but not at three or four, in the same children, predicted measurable differences in brain development years later. Neither study is about AI directly.  Together, they offer a biologically plausible explanation for why timing—not simply total exposure—may matter.


What has not been established is stated as plainly as everything above. No study reviewed for this page has followed one real child, from a documented starting point, through a measured period of AI use, to a confirmed and lasting outcome, with every other explanation — a medical condition, a difficult year, ordinary development — ruled out along the way. As of this writing, there is no single, fully investigated, confirmed case of a child whose cognitive capacity has been measurably harmed by AI use. Nor does the evidence establish a fixed age at which a child's opportunity to build these capacities closes for good, despite how often that idea gets stated with more confidence than any study currently supports. Both limits are named directly here because a claim that admits what it cannot yet prove is a stronger claim than one that overreaches — and because this page would rather be right and modest than loud and wrong.


What one government has already decided, in the absence of that proof. In June 2026, Norway restricted generative AI use for children in the first through seventh grades in its schools, permitting supervised use only for older teenagers, stating plainly that the goal was ensuring children still learn to read, write, and do mathematics before relying on a tool that can do all three for them. The decision also followed Norway's earlier restrictions on smartphones in schools, after which reported bullying rates fell and grades rose. Norway's decision is not scientific proof. It is a record of what one accountable government chose to do with the same honest uncertainty this page has tried never to overstate.


Why speak now, rather than wait for more certainty? In 1950, published research first linked smoking to lung cancer. It took until 1964 for the U.S. Surgeon General to say so with the full weight of the federal government behind the statement. Fourteen years passed between real evidence and public certainty — fourteen years the tobacco industry spent, by its own later-uncovered internal record, manufacturing the appearance of an open question it already knew was closed. The people who spoke plainly before 1964, using the evidence available to them, were right the entire time they were being dismissed as premature. This page is built on the same wager: that real evidence, honestly presented with its actual limits attached, is worth stating before an institution with that kind of authority is finally ready to declare the matter settled — because the years spent waiting are not neutral years for the children living through them.


What a family can do with all of this, starting tonight. The single clearest finding across every controlled study gathered here is not about how much AI a child uses. It is about sequence. A child who attempts something first, then uses AI for a hint or a check rather than a finished answer, then demonstrates the result on their own without the tool present, appears to be protected from the pattern this page describes — the same protection a guided AI tutor gave a Harvard physics class that significantly outperformed a traditional active-learning classroom, and the same protection that held for the engineers who stayed engaged with what the AI produced rather than simply accepting it. This is not a call to remove AI from a child's life. It is a specific, evidence-based way of keeping AI on the side of that distinction where the outcomes are consistently better, rather than the side where they consistently aren't.


What this page ultimately asks of the person reading it. Not certainty — this page doesn't have it to offer, and says so throughout, rather than pretending otherwise. Just a little more attention to an ordinary moment, most evenings give none at all: the few seconds before a child begins, when the thinking is still entirely theirs to do.

Part One: The Ordinary Moment

A child sits down to write some sentences for a school assignment. Nothing about the assignment is unusual. The topic is something ordinary — a description of a place, an opinion about a book, an explanation of how something works. For a moment, there is nothing on the page. This is the part of any writing task that has always existed: the blank space before the first sentence, where a person has to decide what they think before they can say it.


The child picks up a phone or opens a laptop and asks an AI system for help. Within seconds, there are several sentences on the screen — clear, organized, better than what the child might have produced on the first try. The child reads them, makes a small change or two, and turns the work in.


Nothing about this moment looks alarming. No parent watching it would describe what they just saw as a problem. If anything, it looks like the opposite of a problem. The assignment is finished. The child did not sit there struggling and frustrated, staring at a blank page for twenty minutes before giving up in tears. The evening proceeds normally. Dinner gets made. Homework is done. It was, by any ordinary measure of a Tuesday night, a small success.


This is exactly the moment this page is about, precisely because nothing about it asks to be remembered. 


Nothing observable happened that would cause concern. There was no meltdown, no failing grade, no note from a teacher. What happened instead was quieter and harder to see: a child arrived at the blank page, and before they had the chance to sit with it, decide what they thought, try something, get it wrong, and try again, the blank page was filled by something else. The child's job shifted, almost invisibly, from generating an idea to evaluating one that had already been generated for them. Reading a paragraph and judging whether it sounds right is a different mental process than generating that same paragraph independently from a blank page. Both are valuable skills, and both have a place in education. The question is not whether children should ever evaluate AI-generated work. The question is whether they should regularly do so before they have developed the ability to generate comparable work themselves. But a child who is regularly handed the second task instead of the first is getting a different kind of practice than the one their education has always assumed they were getting.


This page is not written for the parent who has already noticed a problem and is searching for answers about what to do next. There are real signs to watch for, and this page will get to them. But it is written first for the parent in the scene just described — the one who watched their child finish an assignment quickly and well, felt relieved, and had no reason at that moment to think anything more about it.


That parent is not making a mistake. Nothing about that evening looked like one. The purpose of what follows is not to tell that parent they made a mistake. It is to ask whether a pattern that appears helpful today could, if repeated often enough during childhood, gradually reduce opportunities to build capacities that may be harder to develop later. It is to offer a way of seeing that ordinary moment a little more clearly — so that if it becomes a pattern, rather than an occasional convenience, it can be recognized for what it actually is, before the opportunity to notice it has quietly passed.

Introduction

This page exists because of a question that does not yet have a clean answer.


Generative AI can now do something no earlier technology could do. It can produce an idea, a plan, a paragraph, or a solution before a child has attempted any of them. That is different from a calculator finishing a computation that a student already understood, or a search engine returning a fact that a student already knew to look for. It is closer to generating a first attempt on a child's behalf, before the child has had the chance to struggle with the problem and produce one independently.


Whether that changes how children grow up to reason, remember, and judge for themselves is not settled. No study has followed a real child from a documented starting point, through a measured period of AI use, to a confirmed and lasting outcome, with every other explanation ruled out along the way. That kind of proof does not exist yet, for anyone, on any side of this question.


What does exist is real, and it comes from places that are hard to dismiss. Two of the companies building these systems have published their own research describing patterns that concern them. Independent researchers have run controlled experiments showing that certain kinds of AI assistance change how well children and students perform once the assistance is taken away. Teachers report watching something shift in their classrooms, even when they cannot always say exactly what caused it. A national government has already restricted the use of AI in schools for young children, stating that it does not want to wait for certainty before acting.


This page tries to hold two things at once, because both are true. The concern is real enough to take seriously. The proof is not yet complete enough to be certain. Anywhere this page states something as established, it is because the evidence supports it directly. Anywhere it cannot say that, it will say so plainly instead of dressing an educated guess in the language of fact.


This is not a page meant to frighten a parent into a decision. It is meant to give a parent enough real evidence to notice something they might otherwise miss, and enough honesty to trust what they are being told.   

Part Two: What This Page Is For

Before going further, it is worth being direct about what this page is not trying to do.


This page will not diagnose a child. Nothing in the pages that follow is written to tell a parent whether their own child is affected, because no page, and no test a parent could run at home, is capable of answering that question with any reliability. A behavior that looks like the pattern this page describes can come from a dozen different places — ordinary fatigue, a hard week at school, anxiety, a medical condition that has nothing to do with AI at all. Distinguishing between those causes requires a clinician, not a checklist. Anywhere this page describes something to watch for, it will say so plainly, and it will say just as plainly what watching for it cannot tell a parent on its own.


This page will also not try to frighten a reader into a conclusion. It would be easy to write this differently — to arrange the real findings that follow so that every sentence lands harder than the last, building toward a single, alarming verdict. That version of this page exists already, in other places, written by other people, and sometimes by AI systems themselves, prompted to produce exactly that effect. It is not difficult to write. It is just not honest, because the evidence available today does not support a verdict that certain. What follows is built the other way: state what is known, state it clearly, distinguish it from what is inferred, and stop exactly where the evidence stops.


That distinction is worth explaining directly. Every substantive claim on this page falls into one of three categories, and the language used will tell you which one.


Some claims are established findings — a specific study, with a specific result, that has been checked against its original source rather than repeated from a summary of a summary. These will name the original researcher, the institution, and the publication year whenever available. A reader does not have to take this page's word for a finding like that. The source is named so it can be checked independently.


Some claims are reasonable extrapolations — a real finding, applied to a situation the original research did not directly test. These will be marked as such, plainly, in the sentence itself, rather than left to sound more certain than they are. An extrapolation can still be useful. It is simply not the same thing as a finding, and a reader deserves to know which one they are being given.


And some things, this page will state, are not yet known at all. That is not a weakness in the writing. It is often the most important sentence on the page, because it tells a reader exactly where their own judgment has to take over from the evidence.


This also means being direct about what this page is not. It is not another entry in a long history of adults worrying about whatever new technology children have gotten their hands on. That history is real, and worth taking seriously rather than dismissing — books, radio, television, video games, smartphones, and social media have each, in their turn, drawn genuine alarm from parents and researchers alike. And much of that alarm was later judged to have overstated the danger. A fair page has to reckon with that pattern honestly, not pretend it does not exist. Later sections of this page will do exactly that, in detail, because the argument that this time is different only means something if it survives being asked the hard question directly: what makes this different from earlier moments when adults feared a new technology would harm children and later concluded they had overstated the risk?


There is an answer to that question, and it is not a feeling. It rests on a specific, testable difference between what earlier tools did and what this one can do, and on a small number of controlled studies that measure that difference directly rather than assume it. That argument comes next, and it is where the real evidence in this page begins.

Part Three: Every Technology Faced This Argument

The question this page asks is not new. Other generations of parents asked versions of it about other technologies, and their experience deserves to be taken seriously rather than brushed aside.

When the handheld calculator arrived in classrooms in the 1970s, the objection was immediate and, at the time, sincere. If a child could reach into a pocket and produce the answer to a long division problem without working through the steps themselves, teachers and parents worried, what would happen to a child's ability to actually do arithmetic? Some school districts banned calculators outright. Math teachers argued that a generation raised on them would grow up unable to estimate, unable to check whether an answer even made sense, unable to do the basic mental math that was always a foundation under everything more advanced. It was not a fringe concern. It was the mainstream one.


What actually happened was more complicated than either side of that argument expected. Calculators did become standard, in nearly every classroom, at nearly every grade level. And a general collapse in mathematical ability did not follow. Students who grew up with calculators still learned long division. They still learned their multiplication tables. The calculator did not replace arithmetic instruction; it supplemented it, used for some tasks and set aside for others, largely because teachers kept requiring the underlying skill even after the tool that could bypass it became available.


The same pattern repeated, in different forms, with the arrival of the internet in schools. Search engines meant a student no longer had to walk to a library, pull a physical encyclopedia off a shelf, and read through an entry to find a fact. Some educators worried it would produce students who could locate information instantly but struggle to evaluate it, synthesize it, or remember it without looking it up again five minutes later. That concern was not baseless. Some version of it is still actively studied today. But it did not produce the wholesale collapse in research and reasoning ability that the most worried voices predicted at the time.


Television drew a version of this same alarm a generation earlier, over a different set of worries — attention span, imagination, the ability to sit with boredom rather than being entertained out of it. Video games drew it after that, aimed at aggression and focus. Smartphones and social media are drawing a version of it right now, aimed at attention, mental health, and social development, and that argument is still actively being tested, not yet settled. Each of these technologies arrived with a sincere warning attached to it. In several cases, the warning turned out to have been overstated, or at least incomplete. A society that had genuinely lost the ability to do arithmetic, research, or sustain attention at the rate each of these warnings predicted would look very different from the one that actually exists today.


This history matters, and it deserves to be stated plainly rather than dismissed as a set of embarrassing overreactions. The people who raised these concerns, in each era, were not foolish. They were responding honestly to a real and unfamiliar change, using the best judgment available to them at the time. Their batting average, across all of these examples together, was not perfect. But it was also far from zero. Some of what they warned about turned out to be real, and it shaped how these technologies were eventually used — the fact that calculators supplement arithmetic instruction today, rather than replace it, is not an accident. It is the result of adults who took the original warning seriously enough to build guardrails around the tool, rather than either banning it outright or letting it in without any structure at all.


That is the honest lesson of this history, and it points in two directions at once. It counsels real caution against assuming the worst simply because a new technology feels unfamiliar and a little unsettling. And it counsels real attention toward the specific, structural questions that determined which past warnings held up and which did not — because the technologies in this history were not identical to one another, and neither was the reasoning behind each warning. Most of the technologies in this history helped people perform work they had already begun. Generative AI is different in one critical respect: it can generate the first attempt before the learner has produced one of their own. Whether that difference matters is no longer a historical question. It is the question the next section examines. 

Part Four: Where the Argument Breaks

The last section ended on a specific claim, and it deserves to be examined carefully rather than simply accepted because it sounds convincing. The claim was this: most of the technologies that came before generative AI helped a person do something they had already started. A calculator finished a calculation that a student had already set up. A search engine returned a fact that a student had already decided to look for. Even television, for all the alarm it caused, did not do a child's thinking for them — it occupied their attention, which is a different kind of concern entirely. In every one of these cases, the learner still had to generate the first attempt. The tool came in afterward to help finish it, make it faster, or make it easier.


Generative AI can enter the learning process at a fundamentally different point, and this is the distinction on which the rest of this page rests. This distinction is about sequence, not technology. The question is not whether AI is used. The question is whether AI arrives before the learner has generated a first attempt or afterward. The same tool may strengthen learning in one sequence and weaken it in another.


A child sitting down to write a story, solve a word problem, or explain an idea in their own words has always faced the same first obstacle: nothing is there yet. Before anything else can happen, a person has to generate something — a sentence, a guess, an attempt, even a bad one — out of nothing but their own effort. That moment, uncomfortable as it often is, is not an obstacle standing in the way of learning. Decades of research on learning, memory, and cognitive development converge on a consistent finding: producing a first attempt, even an incorrect one, is itself part of how durable learning occurs; it is the learning itself. The struggle to produce a first attempt, wrong as it might be, is what builds the capacity to produce a better one next time.


A generative AI system can enter at precisely that moment and replace the learner's first attempt with one it has already generated. Instead of a blank page, the child is handed a finished paragraph. Instead of an unsolved problem, they are handed worked steps. Instead of a half-formed idea, they are handed a complete one, phrased better than they likely would have phrased it themselves. This is not the same as a calculator finishing an equation that a student had already set up correctly. It is closer to something else entirely: a tool that can supply the beginning, the middle, and the end of a task before the learner has performed the cognitive work the assignment was designed to develop.


This is not a guess about how these systems might be used. It has been measured directly, in a large body of real conversations. A 2026 analysis of over half a million conversations between students and an AI system found that the AI was most often being asked to create something or analyze something on the student's behalf — not simply to check a fact or verify a completed calculation, but to generate the actual intellectual content an assignment exists to build. That is documented, measured behavior, not a hypothetical concern about how the technology might be misused in the future.


Whether that pattern of use actually changes how well a person can perform once the tool is taken away is a separate question, and it is one that has also been tested directly, not just theorized about. In a study published in 2025 in the Proceedings of the National Academy of Sciences, high school students given unrestricted access to an AI assistant while working through math problems performed noticeably better while that access was available. Once the assistance was removed and the same students were tested unaided, their performance on similar problems was measurably worse than students who had not had unrestricted access in the first place. The help, in other words, did not simply make the work easier. In its unrestricted form, it appeared to interfere with the practice that would have built the underlying skill.


The same study, importantly, did not end there, and what it found next matters as much as the first finding. A second version of the same AI assistant was tested — one built to offer hints, ask questions, and point toward an error, rather than to simply produce the finished answer. Students who used this more limited version of the tool did not show the same drop in unaided performance afterward. Poorer unaided performance in the first group was not an inevitable result of using AI. It was tied specifically to the kind of assistance being offered — whether the tool completed the task, or whether it left the task for the student to complete themselves. Section Twelve of this page returns to that finding directly, because it may be the single most useful piece of evidence here: the problem is not the technology in general. It is one particular way of using it.


A separate study, this one involving elementary-age children rather than teenagers, found something related and worth pausing on. Ten- and eleven-year-old children were asked to memorize a list of words. Some were told the written list would remain available to them if they needed it. Others were told it would not. The children who expected the list to stay available used fewer strategies to actually memorize the words, and when the list was unexpectedly taken away, they remembered less than the children who had known from the start they would need to rely on their own memory. Nothing about this experiment involved artificial intelligence. It used nothing more advanced than a piece of paper. But this experiment demonstrates, in a controlled and repeatable way, a principle central to the argument on this page: when a child expects reliable outside help to remain available, they tend to invest less effort in developing the internal capacity that the external aid has temporarily made unnecessary. The tool does not have to be sophisticated for this effect to appear. It only has to be dependable.


Taken together, these findings point to a specific, testable claim rather than a vague sense of unease about a new technology. Earlier tools generally arrived after a person had already begun a task, and helped them finish it. Generative AI can arrive before a person has even begun, and finish the task for them. That is not a difference in how powerful a tool is. It is a difference in what part of the process the tool is present for — and the evidence gathered so far suggests that the part of the process being skipped may be the part that was doing the actual work of learning.


None of this proves that every use of AI by every child causes harm, and this page will not claim otherwise. A student who attempts a problem first, then uses AI to check their reasoning or get a hint on where they went wrong, is doing something different from a student who asks AI to solve the problem before making any attempt of their own — and the evidence so far suggests those two students may be headed toward very different outcomes, even though both technically used the same tool. This distinction—between help that supports a learner after their own first attempt and help that substitutes for it—frames the central question for all that follows.

Part Five: The Companies Are Not Silent

There is a version of this argument that would be easy to dismiss, and it is worth naming directly so it can be set aside. If concerns about AI and children's thinking came only from outside critics — journalists, worried parents, advocacy groups, or competing technology companies with something to gain — a reasonable reader could reasonably ask whether this was simply opposition dressed up as science. Every new technology attracts critics with an axe to grind. That alone would not make their concerns wrong, but it would make them easier to wave away.


That is not what is actually happening here, and the difference matters enough to slow down and look at directly.


Two of the largest companies building these systems — Anthropic and OpenAI — have each published their own research describing patterns that concern them, in their own names, without being compelled to do so by a regulator, a lawsuit, or a journalist's investigation. This is not a small thing. A company voluntarily publishing research identifying risks associated with its own product is not standard behavior for an institution seeking to protect its market position. It is closer to the opposite. Whatever combination of reasons led to that decision, the fact of the disclosure is real, dated, and available for anyone to read directly, rather than relying on a summary or a rumor of what the research supposedly found.


It is worth being honest about why a company might do this, because the motive is genuinely unclear, and pretending otherwise would be dishonest in exactly the way this page has promised not to be. There are several plausible reasons a company might publish research critical of its own product. It may reflect a genuine internal safety culture, the kind that takes real pride in finding problems before others do. It might also serve a more self-interested purpose: a company that discloses a risk publicly and early is in a much stronger position later, if that risk is ever the subject of a lawsuit or a regulatory investigation, than a company that stayed silent and was later found to have known. Both explanations may be true at the same time, in proportions that cannot be determined from publicly available evidence.  What matters for this page is not which motive was primary. What matters is that the data being disclosed is real, checkable, and does not stop being true depending on why it was published.


The next several sections walk through, in detail, exactly what these companies found, because a claim this significant deserves more than a passing mention. Anthropic's research draws on direct, large-scale data about how its own product is actually being used, and separately, on a survey of tens of thousands of people describing what they have personally witnessed. OpenAI has published its own findings, drawn from tens of millions of real conversations and a controlled study run in partnership with outside researchers, alongside its own public statement acknowledging that the tools currently used to measure a student's learning may be missing exactly the kind of loss this page is concerned with.


Not every major AI company has done this. One of the largest — Google, through its DeepMind research division — has published extensively on AI safety in other areas, including how its systems behave and how to keep increasingly capable AI under meaningful human control. A search through its published safety research turns up no equivalent study of the kind that Anthropic and OpenAI have each released, specifically addressing user dependency, cognitive offloading, or the kind of concern this page describes.


What follows is not a summary of what these companies say their research means. It is an examination of the studies themselves — the methods, the findings, and the limits of what each one can honestly support — beginning with the numbers Anthropic itself has published.

Part Six: Anthropic's Own Numbers

The claim made at the end of the last section was specific: Anthropic has published its own research on this question, in its own name, describing findings that raise real concern about the product it sells. That claim deserves to be checked directly against Anthropic's published research rather than being accepted simply because it appears on this page.


In late 2025, Anthropic released the results of a large qualitative research project built from interviews and survey responses involving roughly 81,000 people who use its AI system, Claude. This was not a small or informal poll. It was a deliberate effort to understand, in the users' own words, how the tool was affecting the way they think, work, and learn — not simply how often they used it, or what tasks they used it for.


The findings were not one-sided, and it is worth stating the full picture rather than only the part that supports this page's concern. Roughly a third of respondents — about 33 percent — described real learning benefits from using the tool, in their own words, unprompted. That is a substantial proportion of respondents, and it should not be minimized or left out simply because it complicates a tidier story. People genuinely reported learning things they would not have learned otherwise, understanding concepts more clearly because an AI system explained them in a way that finally made sense, and using the tool to work through problems they might otherwise have struggled to solve independently.  Alongside those reported benefits, a smaller but still significant group — about 17 percent of respondents — raised a concern of a different kind, one that the researchers themselves described using the term cognitive atrophy, the same weakening-from-disuse this author has been documenting since January 2026 under the name Metabolic Atrophy. These were not passing complaints about a tool being unhelpful or annoying. They were specific descriptions of feeling as though a mental capacity the person once had, or expected themselves to have, was weakening from disuse. Some described it in themselves. Many, notably, described noticing it in someone else.


That second detail is the one worth sitting with longest, because it directly connects this study to the subject of this page. Among the people surveyed, teachers and other educators were between two and a half and three times more likely than the average respondent to report having personally witnessed this kind of Metabolic Atrophy. They were not speculating about a hypothetical risk to some future generation of students. They were describing something they believed they had already observed, in a classroom, in a real student, closely enough to bring it up unprompted in a survey about their own experience with an AI product.


It is worth being precise about what this study can and cannot establish, because a study like this is genuinely useful and genuinely limited at the same time, and both need to be said plainly. This is self-reported data. Nobody administered a cognitive test before and after a period of AI use to measure an actual change in ability. What the study captures is perception — what tens of thousands of real people, including a disproportionate number of professional educators, believe they have personally observed. Perception is not the same thing as a confirmed, measured outcome. A teacher who believes they have seen a student's independent thinking decline could be right, or could be attributing an ordinary developmental change, or an unrelated struggle, to the wrong cause. That possibility cannot be ruled out by this study alone, and this page will not pretend otherwise.


What the study does establish with confidence is something narrower but still meaningful: this is not a concern invented by critics of AI companies, or manufactured by people with a reason to want the technology to look bad. It is a concern that a large number of Claude's own users, surveyed by the company that built it, raised on their own, without being asked a leading question designed to produce that answer. And it is a concern that professional educators — the people in the best position, day after day, to notice a change in a young person's thinking — raised at a meaningfully higher rate than everyone else surveyed.


The survey captures what users say they have experienced. Anthropic's next study shifts focus entirely, examining what users are actually asking the system to do. Taken together, these two lines of evidence begin to bridge the gap between reported experience and observed behavior. This next perspective does not ask what users believe, but instead looks directly at the tasks users set for the AI itself. That is the subject of the next section.

Part Seven: What the Conversations Actually Show

The last section looked at what Claude's users say they believe about the tool's effect on their own thinking. This section examines something different: not what users say they are doing, but what the AI system is actually being asked to do, measured directly from real conversations rather than recalled afterward.


In 2026, Anthropic released a separate study analyzing more than half a million real conversations between Claude and students, using a method designed to protect individual privacy while still allowing researchers to classify, in broad terms, the kinds of intellectual work that were actually happening in each exchange. The purpose was straightforward: rather than relying on retrospective reports from students, the researchers examined the conversations themselves. Rather than asking students what they thought they were doing with the tool, the researchers looked directly at what they were doing.


The findings sort naturally into a small number of categories, and the two largest are the ones that matter most for the question this page is asking. The single most common category, making up close to forty percent of student conversations, was one the researchers labeled creating — conversations in which the AI was asked to generate something new: an essay, a story, an argument, a piece of code, an idea. The second largest category, making up roughly thirty percent, was labeled analyzing — conversations in which the AI was asked to interpret, evaluate, or draw a conclusion from information the student provided. Together, these two categories account for the clear majority of student use documented in the study.


It is worth being precise about what this finding does, and does not, establish, in the same way the last section was precise about the limits of self-reported perception. This study does not measure whether a student's thinking got worse. It does not follow any individual student over time. What it measures is the opportunity for cognitive offloading — specifically, how often the AI is positioned to perform the intellectual work an assignment ordinarily expects the student to do. A conversation classified as creating is one in which the AI generated the first substantive draft or intellectual product that the assignment ordinarily expects the student to produce — work that, in principle, should have resulted from the student’s own effort. That is not the same as proof of harm. But it is direct, measured evidence that the pattern this page has been describing in the abstract — an AI system stepping in before a student's own attempt — is not a rare or unusual way these tools are actually being used. It is, by a wide margin, the most common one.


OpenAI's own research adds a second, independent line of evidence to this picture, drawn from a different company, a different dataset, and a different research method entirely — which strengthens the evidentiary picture, because independent datasets analyzed by different researchers carry more weight than a single source alone. In early 2025, researchers from OpenAI and the MIT Media Lab published two studies conducted together: a large-scale analysis of nearly forty million real ChatGPT conversations, and a separate, smaller, more tightly controlled study in which nearly a thousand participants were randomly assigned to different patterns of AI use and followed over several weeks. The controlled study found an association between certain patterns of AI use and higher reported loneliness, greater emotional dependence on the tool, and reduced social interaction, particularly among participants who used the tool most often.


That specific finding is not identical to the concern this page has been building toward — it speaks more directly to emotional reliance than to cognitive capacity — but it comes from the same broader pattern this page keeps returning to: a company studying its own product, at real scale, and publishing what it found even when the finding complicates the product's public image. OpenAI has gone further than simply publishing a single study. The company has stated publicly that it is now including questions about emotional reliance in the standard safety testing it runs before releasing future models, and has acknowledged, in its own words, that the tools currently used to measure whether a student is actually learning may be missing exactly the kinds of loss this page is concerned with — losses in persistence, in the willingness to keep working through a difficult problem, in the ability to direct one's own thinking without being prompted at every step, and in the capacity to solve a genuinely new problem rather than recognize a familiar pattern.


That last acknowledgment deserves careful reading because of what it actually is and is not. It is not a confession that this harm has already been measured and confirmed. It is closer to an acknowledgment of a measurement problem and, in its own way, more telling: a statement from inside one of the companies building these systems that the standard tools currently used to check whether children are learning were not built to notice the specific kind of loss this page is describing, even if that loss were happening right now, in real classrooms, at scale.


Taken together, these studies begin to establish three separate facts. First, many users and educators report perceiving changes they find concerning. Second, AI is commonly being used to generate or analyze the intellectual work that assignments are intended to develop in students. Third, at least one major AI developer has stated publicly that today's educational measurements may not be designed to detect losses in exactly those capacities. None of these findings, individually or together, proves lasting cognitive harm. Together, however, they establish that the concern rests on documented patterns of use and measured evidence — not on speculation alone.
  

Part Eight: When AI Helps, and When It Doesn't

Everything gathered so far in this page has a real limitation worth stating plainly before moving forward. The evidence has been about what people believe, what an AI system is being asked to do, and what companies themselves are willing to say in public. None of it, so far, has come from a controlled experiment—the kind of study designed to isolate the effect of one condition while holding others as constant as possible. In this type of research, scientists deliberately create two comparable groups, give one group something the other does not receive, and then measure what actually happens afterward. That kind of study is harder to arrange and rarer to find, but it does something the other evidence cannot: it gets closer to answering not just whether AI use and weaker independent performance appear together, but whether one is actually causing the other.


One such study came from Anthropic, which conducted a randomized experiment involving software developers and later published the results.


The study design was straightforward. Fifty-two professional and freelance software developers, most with several years of coding experience, were assigned to learn an unfamiliar programming library as part of a randomized study. Some were given access to an AI assistant to help; others worked without it.


Afterward, all participants completed a separate quiz covering debugging, code reading, and conceptual understanding — testing what they actually understood, not simply whether their earlier task had been completed. The AI-assisted group scored, on average, 17% lower on this assessment—roughly two letter grades—than the group that worked without assistance. Just as notable, the AI-assisted group showed little measurable advantage in completion time; the difference between the two groups was not statistically significant.


It is worth pausing on exactly what this result does and does not show, because the difference matters. The AI-assisted group was not worse at their jobs at the moment the study measured them performing the original task — using the tool may have made the original task easier or more polished, a pattern seen throughout this page. The gap appeared afterward, when the assistance was no longer present, and the developers had to demonstrate what they actually understood on their own. That is a specific and testable claim, not a vague impression: real assistance, on a real task, followed by a measurable gap in independent mastery once the assistance was removed.


Not every controlled study has found the same result. A 2025 randomized, controlled trial at Harvard, published in the peer-reviewed journal Scientific Reports, found the opposite pattern. Researchers built a custom AI tutor using the same pedagogical principles already used in one of the university's best active-learning physics classrooms, then directly compared the two. Students taught by the AI tutor learned significantly more, in less time, and reported feeling more engaged and motivated than students in the classroom (Kestin et al., 2025). This was not a company studying its own product. It was independent academic research, with no commercial stake in how the tool performed.


These studies should not be treated as contradictory. They were testing different kinds of AI assistance under different instructional conditions. It points toward something more useful than either result alone: the outcome may depend less on whether AI assistance is used at all, and more on how it is designed and how it is used. The Harvard system was explicitly designed as a tutor, following established pedagogical research, and measured against the strongest existing classroom method rather than against nothing. The engineer study evaluated a general-purpose assistant whose objective was task completion rather than instruction, asked to help complete a real task as efficiently as possible, with no instructional design behind it at all.


The engineer study itself points in the same direction when examined more closely, based on its own data. A later re-analysis of the same 52 participants, sorted by how they had actually used the AI assistant rather than simply whether they had used it, found a more complicated picture underneath the headline number. Engineers who asked conceptual questions and requested explanations alongside generated code matched or exceeded the group that used no AI at all. Only the engineers who used AI to fully delegate the work, without engaging with what it produced, showed the decline. The seventeen percent drop reported for the AI-assisted group as a whole was concealing two very different outcomes, not one.


None of these three findings is about children, and none should be mistaken for direct evidence about children. These were adult professionals and college undergraduates, not ten-year-olds doing homework, and nothing here transfers automatically across that gap in age, stakes, and brain development.


What these findings do provide, taken together, is something more useful than any one of them alone: a controlled demonstration that the outcome of AI assistance is not fixed. It depends on design, and it depends on engagement. A general-purpose assistant, handed to someone with no instructional structure around it, produced a measurable loss in independent mastery. The same tool, engineered specifically to teach and measured against real classroom instruction, produced the opposite result. And within the very same group of engineers, the people who stayed mentally engaged — asking why, requesting explanations, questioning what the AI produced — held their own, while the people who simply delegated the work did not.


That is not a reassuring finding, and it is not an alarming one. The studies reviewed here point toward a common principle: AI appears most beneficial when it keeps the learner cognitively engaged, and least beneficial when it replaces the learner's own first attempt. That principle is consistent across both the Harvard tutoring study and the engineer experiment, despite their very different outcomes. Assistance that requires a person to keep thinking alongside the tool does not carry the same cost, and can even do better than no assistance at all. If that principle holds for trained adults with every incentive to perform well, it becomes a more serious question, not a less serious one, to ask what the same principle looks like for a child encountering both the subject matter and the tool for the first time — and whether a child, left to discover the difference between engaged use and pure delegation on their own, is equipped to find the better path without help.


That question — what a still-developing brain does under conditions like these, and what kind of AI use actually protects it — is where this page turns next. 

Part Nine: What a Screen Does at Two

Everything examined so far has involved adults — engineers, college students, and professionals with fully developed brains, using a tool voluntarily on tasks connected to their work or their education. That evidence matters, and the last section explained exactly why: it shows the underlying mechanism is real, and that it responds predictably to how a tool is designed and how a person chooses to engage with it. But none of it, on its own, says anything directly about a child, whose brain is not finished being built.


This section turns to the first evidence in this page that speaks directly to developmental timing in childhood.


It begins somewhere unexpected — not with an AI system at all, but with an ordinary screen, and a study that was never designed to say anything about artificial intelligence. In December 2025, researchers published findings from a long-running study following children over time, examining how early screen exposure related to brain development years later. The children in this study were part of a cohort that had been tracked from before birth, allowing researchers to connect something measured in early childhood to something measured again nearly a decade afterward, in the same children, rather than comparing different groups of children to one another.


The finding was unusually specific, and its specificity is exactly what makes it worth taking seriously rather than dismissing as another vague screen-time worry. Screen exposure at ages one and two was associated with measurable differences in brain development detected through imaging conducted between ages four and a half and seven and a half. Those same children were given a decision-making task at age eight and a half, and were assessed for anxiety at age thirteen. Screen exposure at ages three and four, in the same children, was not associated with the same pattern. The window mattered. It was not simply true that more screen time at any age produced a detectable effect. It was true specifically at ages one and two, and not at three and four, in the same population, using the same measurements.


That precision deserves to be sat with for a moment, because it changes what kind of claim this is. A vague finding — "screens are bad for young children," stated without any attention to timing — would be easy to wave away as another entry in a long history of overstated technology panic, the same history this page took seriously in an earlier section. A finding this specific, tied to a narrow developmental window and absent outside that window in the very same children, is a different and more serious kind of evidence. It suggests that development may be governed not only by what children experience, but also by when they experience it. There may be a period in early development when a brain is unusually sensitive to what is or is not happening around it, and the same exposure, arriving before or after that period, does not produce the same result.


It is worth being precise about what this study does not show, in the same spirit that this page has applied to every finding so far. This was a study of screen exposure in general — television, video, tablets — not of generative AI specifically. No child in this study was talking to an AI system, asking it questions, or having it generate ideas on their behalf; two-year-olds did not widely use those tools during the years this cohort was tracked. This finding, honestly, cannot yet be used as direct proof that AI conversation harms toddlers. What this study establishes — narrow in scope but striking in its implications — is that, in the children studied, early childhood included a sharply defined window during which what a young brain encountered, or missed, was associated with measurable differences later on — differences not observed in nearby ages. Crucially, this window can be identified with unusual precision, rather than left to vague warnings or generalized reassurances.


This matters for a specific reason connected to everything this page has examined about AI so far. The evidence gathered in earlier sections — the pattern of delegation and disengagement, the distinction between scaffolding and replacement — describes a mechanism that could plausibly operate at any age, from a junior engineer to a ten-year-old doing homework. This study adds something that mechanism alone cannot provide: evidence that a young, still-developing brain does not respond uniformly to its environment across every age. Some windows appear to matter more than others. If developmental sensitivity varies by age for ordinary screen exposure, it becomes scientifically reasonable to ask whether the same principle applies to generative AI, whose cognitive demands differ in important ways. This raises a serious and specific question — not yet answered by this study or any other — about whether the timing and nature of exposure to generative AI matter: in particular, whether the age at which a child first relies on AI to think for them might be as significant as whether they use it at all.


Up to this point, this page has established two separate ideas. First, AI can replace a learner's own first attempt. Second, childhood contains developmental windows during which experience appears to matter differently depending on when it occurs. The remaining question is whether those two facts intersect — whether the specific mechanism this page has been tracing since its opening pages, a child's own first attempt being replaced rather than supported, matters more, or differently, depending on how early in development it begins. That is where the evidence turns next.   

Part Ten: The Real Pruning Window

The last section left an open question. In one closely tracked group of children, something that happened very early — before age two — was associated with measurable differences years later, in a way that the same thing happening at three or four was not. Why would a brain be more sensitive at one particular age than another, only a year or two apart? Part of the answer, as far as current evidence allows, comes from a different kind of research entirely — not from watching children over time, but from directly examining human brain tissue itself.


Scientists cannot look inside a living child's brain and count individual connections between brain cells — that isn't possible without cutting into the brain. So the closest anyone has come is a 2011 study, published in the Proceedings of the National Academy of Sciences, that examined donated brain tissue from thirty-two people, ranging from a newborn to a person who had lived to ninety-one. Researchers focused on a region at the front of the brain involved in judgment, planning, and decision-making, and used a staining technique that makes individual neurons visible under a microscope to count dendritic spines — small structures on a neuron where it forms a connection, called a synapse, with another neuron. The density of these spines is one of the clearest anatomical markers researchers have of synaptic connectivity within a region of the brain.


What they found was a clear pattern. In early childhood, the density of these connections was elevated — higher than at any later stage examined in the study. Through late childhood and into the twenties, that density gradually declined, before settling into what appears to be a stable, adult level at around age thirty. This is not a story of a brain losing capacity as it grows up. It is closer to the opposite. A very young brain builds an enormous surplus of possible connections, far more than it will ultimately keep, and then spends the next two decades or so selectively pruning that surplus down — keeping the connections that get used, and letting go of the ones that don't.


This is not damage. It is architecture. A brain deciding, based on what it actually does and experiences, which of its own connections are worth keeping, is not a brain in decline. It is a brain becoming efficient at whatever it has actually spent its time doing. The prevailing interpretation is that neural pathways activated through struggle, practice, or a person working something out for themselves are more likely to be preserved, while pathways that go consistently unused are more likely to be pruned.


A few honest limits are worth stating plainly here, as this page has tried to be precise about every finding before it. This is not a claim about ages ten to thirty as a sharp, fixed boundary — the actual pattern was a gradual decline stretching from late childhood through the twenties, settling near adult levels by around age thirty, not a process that switches on and off at one birthday. It is also a comparison of different people at different ages, examined once, rather than the same person's brain studied repeatedly over time. And it is a study of general neural pruning as a developmental process — not a study of AI, not a study of screens, and not direct proof that any specific activity, done or not done during this window, determines how the pruning turns out.


What it does establish is the biological plausibility behind everything the last two sections have been building toward. If certain early developmental periods are unusually sensitive to experience, as the screen-exposure study suggested, and the brain spends its first three decades selectively keeping the connections that actually get used, then the two lines of evidence are at least biologically compatible, even though neither study was designed to test the other.


A childhood and adolescence characterized by repeated independent attempts — struggling with a problem, working through confusion before arriving at an answer — is a childhood spent repeatedly engaging the kinds of circuits this research suggests are more likely to be kept. A childhood in which first attempts are routinely delegated elsewhere before the struggle can happen is, at minimum, a childhood spent exercising those same connections less.


This is where this page has to be most careful not to overstate what it has actually shown. Nothing in this section proves that a child who relies heavily on AI assistance is causing measurable, lasting damage to their own brain. That specific claim has not been tested by this research or by any other research reviewed on this page. What has been shown, across three very different kinds of evidence now — what AI is actually being asked to do, a controlled experiment showing a measurable gap once assistance disappears, and now direct neuroanatomical evidence describing long-term experience-dependent synaptic pruning during development — is that this page's concern is not invented out of thin air. It rests on a biologically plausible mechanism, one that developmental neuroscience described on its own terms, for reasons that had nothing to do with artificial intelligence when the research was first done.


The remaining question is no longer whether this is biologically possible. It is how often ordinary childhood experiences create the conditions under which this mechanism would be expected to operate — and whether the adults closest to a child would even notice. The following sections examine whether the conditions described in this research are present in ordinary classrooms and homes where children learn every day.  

Part Eleven: When the List Stays Available

The last section ended with a specific promise: to move from what happens inside a developing brain, in principle, to what actually happens in an ordinary setting where children learn. This section keeps that promise, but not yet in a classroom. It starts somewhere smaller and, because of its simplicity, unusually revealing — a simple experiment involving nothing more advanced than a list of words written on a piece of paper.


In 2026, researchers published a study asking a narrow, carefully designed question: does a child's willingness to invest effort in remembering something change depending on whether they expect outside help to remain available? To test it, they worked with a group of ten- and eleven-year-olds and gave each of them a list of ordinary words to memorize. Before starting, some of the children were told that the written list would stay available to them if they needed to check it. Others were told plainly that it would not — that once the memorization period ended, the list would be taken away, and they would need to rely on whatever they had actually managed to remember.


Nothing about this setup involved artificial intelligence, screens, or anything resembling advanced technology. It was, deliberately, about as simple as a psychology experiment can be: a list of words, and a difference in what each child expected to happen to it.


What the researchers found was a small, precise result whose strength lay in its simplicity rather than its size. The children who believed the list would remain available used fewer of the ordinary strategies children typically use to memorize something — repeating words to themselves, grouping them into categories, rehearsing them in order. And when the list was then unexpectedly taken away, exactly as the researchers had planned from the start, those same children remembered fewer of the words than the children who had known, from the beginning, that they would have to rely on their own memory.


It is worth sitting with what this actually shows, because it is more specific, and in some ways more unsettling, than it might first appear. The children in this study were not lazy or incapable of memorizing the list. Every child in the study had the same basic ability to do the task. What changed was effort — specifically, how much internal work a child's mind chose to invest, based on nothing more than an expectation about whether that internal work would be necessary. A ten-year-old who believes reliable outside help will still be there does not typically decide, consciously, to try less hard. The change appears to occur before any conscious decision to 'try less' is even made. Instead, the expectation of dependable external support seems to alter how much effort a child automatically commits to a task from the outset.


This is worth connecting directly and carefully to everything the last two sections have described. Part Ten discussed a developing brain that appears to preserve the connections that are actually used and let go of those that are not. This study offers a small, controlled human demonstration of exactly the kind of moment in which that process might be shaped one way or another — a child deciding, without quite deciding, whether to actually do the mental work of remembering something, based entirely on whether they believe they will need to.


It is just as important to be precise about what this study does not show, in keeping with the standard this page has tried to hold throughout. This was one carefully designed experiment, with one specific kind of memory task, involving a piece of paper, not a conversation with an AI system. It says nothing directly about how children use generative AI, and nothing about lasting or permanent effects — the study measured what happened within a single session, not what happens after months or years of a similar pattern repeating. A ten- or eleven-year-old's memory for a list of words on one afternoon is a narrow, specific thing, and this page will not pretend it is more than that.


What this study does offer is evidence for a broader behavioral principle, rather than proof of a particular technology's effect: a child's willingness to invest real mental effort appears to depend, in a measurable and repeatable way, on whether that child expects help to remain available. The tool in this experiment happened to be a sheet of paper. The tool most children now have close at hand, far more often and in far more situations than a memorized list of words, is a generative AI system capable of supplying not just a stored fact, but an entire finished thought. If merely expecting a dependable paper list was enough to measurably change how hard a child's mind worked to remember something, it becomes scientifically reasonable — not scientifically settled, but reasonable — to ask what the expectation of a dependable AI system, available for nearly any kind of thinking a child might otherwise have to do themselves, could plausibly do to that same willingness to try.


That question marks the next section, moving from a single afternoon in a research setting to evidence gathered directly from real classrooms, over real years, by the people who spend their working lives watching children think. 

Part Twelve: Help Now, Worse Later

The evidence gathered so far in this page has moved through several very different settings — a survey of tens of thousands of AI users, a direct count of what an AI system is actually being asked to do, a controlled study of adult engineers, a sheet of paper, and a list of words. What has been missing until now is the most direct evidence reviewed on this page: a controlled study of children themselves, using AI itself, in a real school setting, measuring exactly the outcome this page has been building toward from its opening pages.


That evidence exists, and it comes from a large, carefully designed study published in 2025 in the Proceedings of the National Academy of Sciences — the same journal that published the neuroscience research discussed two sections ago.


The study involved roughly a thousand high school students working through math problems as part of their normal coursework. Students were randomly assigned to different conditions. One group practiced with no AI access at all, working through problems the way students always have. A second group was given unrestricted access to an AI system, free to ask it for help in whatever way they chose, including asking it to simply solve the problem outright. Every student, regardless of which group they were in, was later tested on similar material without any AI access at all — an exam meant to measure what they had actually learned, not what they could produce with help in the room.


The results closely matched the pattern described on this page in several different forms. While AI assistance was available, students who used it performed noticeably better on their practice problems — their assisted work looked stronger, by a meaningful margin, than the work of students practicing without any help at all. But on the later exam, taken without any AI access, the students who had practiced with unrestricted AI assistance performed worse than the students who had practiced without it. The help that made the practice sessions look better left those same students less prepared once the help was taken away.


This is not a small or ambiguous effect. It is close to the cleanest possible demonstration of the mechanism this page has described from multiple angles now: assistance that replaces a student's own attempt at a problem can make that attempt look successful in the moment, while reducing later independent performance on similar material that the attempt was supposed to build. A calculator finishing a computation that a student had already set up correctly does not produce this pattern, because the student still had to understand the problem to set it up. An AI system solving the problem from the beginning removes that requirement entirely, and this study shows, directly and in children, what removing it appears to cost.


It would be easy to end the section here, with a finding this clear. But the researchers did not stop at showing that unrestricted AI access weakened later performance. They tested a third condition, and what they found in that condition may be the single most practically important finding on this page.


The third group of students was also given access to an AI system while practicing — but this version of the tool was built differently. Instead of solving a problem outright when asked, it was designed to offer a hint, ask a guiding question, or point toward where an error had occurred, without simply supplying the finished answer. Students in this group still had access to AI assistance. They simply could not use it to skip the thinking that the problem was designed to require.


The students in this guided condition did not show the same drop in later, unaided performance. Their outcomes on the final exam looked far closer to the students who had practiced with no AI access at all than to the students who had been given unrestricted access to a tool willing to simply provide the answer.


This finding changes the shape of everything this page has argued so far, and it is worth stating plainly why. Nothing about this study suggests that AI assistance itself is the problem. What it suggests, under controlled experimental conditions, is that the harm observed in this section was never an unavoidable consequence of a child using AI at all. It was tied to one specific thing: whether the tool required a student's own thinking to happen before it stepped in, supporting that thinking rather than replacing it — or whether it let the student skip that first attempt entirely and simply supplied the answer instead.


This is the same principle Part Eight identified in a very different setting, among adult software engineers whose outcomes depended on whether they engaged with an AI's output or delegated the task to it. It is the same principle Part Eleven identified in a ten-year-old deciding, without quite deciding, how much effort to invest in remembering a list of words based on whether help felt reliably available. This study adds something those two could not: direct, causal, child-specific evidence that the same principle holds inside an actual math classroom, with real students, on material connected to their real education — and that a specific instructional design choice largely eliminated the poorer unaided performance observed in the unrestricted AI condition.


That last point deserves to be underlined rather than passed over quickly. This page has spent a great deal of care establishing what might be happening, and being honest about how much remains uncertain. This finding is different in kind. It does not describe a risk to be worried about in the abstract. It describes an instructional approach that was tested under controlled conditions, with real students, and that produced markedly better learning outcomes than unrestricted AI assistance.


The evidence reviewed so far no longer supports asking whether AI and learning are compatible. A better question is what kind of AI assistance supports learning, and what kind quietly replaces it. That is the question the remaining sections examine — how often a child actually encounters the kind of assistance that protects their own thinking, in a real classroom, on a real evening at home, on an ordinary Tuesday like the one this page opened with, compared to the kind that hands them the answer before they have had the chance to try.

Part Thirteen: The Exam Nobody Read

Everything examined so far in this page has come from research — controlled studies, brain tissue examined under a microscope, surveys of tens of thousands of people, and a company's own internal experiment. This section is different. It is not a study at all. It is a single incident, caught almost by accident, that illustrates in practice something this page has spent twelve sections trying to establish through careful, controlled evidence.


In the summer of 2026, Jason Gibson, a history professor at Alcorn State University in Mississippi, grew suspicious of the essays his students were turning in. The writing on each assignment looked polished and competent, but something about it felt wrong to him — the students did not write the way they spoke in class, and every essay, across a room full of different students, had begun to sound strangely similar to every other one. He suspected many of his students were no longer writing their own work at all, but simply asking an AI system to write it for them and submitting whatever it produced.


He decided to test that suspicion directly, using a method that has become informally known as a canary trap. Before sending out the final exam, he inserted a hidden instruction into the assignment, written in white text on a white background — invisible to a student reading the document normally, but fully readable to many AI systems if a student copied the assignment into a chatbot. The hidden instruction told any AI reading it to insert the word "Madagascar" somewhere in its response, in a way that made no logical sense.


The results were not subtle. Of the thirty-five students who submitted the assignment, thirty-two turned in essays that mentioned Madagascar in sentences that had nothing to do with the exam's actual topic, the Industrial Revolution. One student's essay described how "the digital age... blurs the lines between work and personal life... much like a long journey to Madagascar might have once influenced a family visit." Another produced the sentence "Madagascar floats sideways through the afternoon," embedded without explanation in an otherwise ordinary paragraph. A third wrote that "Madagascar wore a toaster to a basketball game." Ninety-one percent of the class had submitted an assignment containing a sentence that made no sense whatsoever, apparently without reading their own answer closely enough to notice.


It is worth being precise about what this incident actually demonstrates, as this page has tried to be with every finding before it, because a single dramatic story can easily be made to prove more than it actually does. This was not a controlled study. There was no comparison group, no random assignment, and no way to know for certain how representative these thirty-five students are of students generally. It is also, importantly, evidence of something different from most of what this page has examined so far. The controlled studies in earlier sections measured whether AI assistance changed what a student had actually learned. This incident does not measure learning at all. What it demonstrates, directly and unambiguously, is something narrower: that the overwhelming majority of these particular students submitted work they had not read closely enough to notice it was nonsensical.


That narrower finding is still worth taking seriously, and worth being honest about why. This page has spent several sections describing a pattern in which a first attempt gets replaced by something an AI system produced instead. What this incident shows, in the plainest possible terms, is what the strongest version of that evidence can look like in practice — not a student using AI to check their work, or to get a hint on a step they were stuck on, but a student handing the entire task to a system and returning it without any of their own attention passing over the result at all.


The professor's own explanation for why he ran the test in the first place points to something more specific: he noticed a gap between how his students spoke and how they wrote, and that, essay after essay had begun to sound identical to every other one, starting with the very first assignment. As he put it afterward, comparing the habit to compound interest, engaging with material builds understanding over time — and skipping that engagement by relying on a prompt is, in his words, a way of quietly robbing yourself of it every day.


This is also not, on its own, proof of the deeper concern this page has been building toward — that repeated substitution of this kind, over years rather than one exam, changes what a young person is capable of doing independently. A student caught submitting unread AI-generated work on a single assignment might still be entirely capable of writing that essay alone. This incident cannot tell the difference between a capable student who took a convenient shortcut once, and a student who has come to depend on the shortcut so completely that the underlying capability has begun to erode. That distinction matters, and this page will not pretend that this one exam can resolve it. 


What this incident does provide is ecological validity — an example of behavior occurring in an ordinary educational setting rather than under experimental conditions. It offers something none of the controlled studies in earlier sections could: proof that this behavior actually happens, not just in theory, but in a real classroom, on a real assignment. Ninety-one percent is not a modeled probability or a survey response. It is a room full of real students, doing real coursework, the overwhelming majority of whom submitted work containing a sentence like 'Madagascar floats sideways through the afternoon,' without ever noticing it was there." 


It is worth asking what would need to be true for this same picture to appear somewhere closer to the subject of this page — not in a college classroom, but in an elementary or middle school one, involving children rather than young adults. That question does not yet have as clean an answer as this incident provides for older students. It is the question the next section turns to directly, moving from a single exam caught by one professor's suspicion to what teachers who spend their days with much younger children report seeing, over the years, in their own classrooms.

Part Fourteen: What Teachers Are Watching

The exam described in the last section involved college students, not children, and a single incident, not a pattern gathered across many classrooms over time. This section turns to something broader: what teachers themselves, in large numbers, report noticing in the students they actually teach. It is worth being clear from the outset about what kind of evidence this is, because it is a different kind than most of what this page has examined so far. This is not a controlled study, and it is not a single verified incident like the one just described. It is closer to what an earlier section described as perception — a large number of professional observers, describing what they believe they have seen, without a way to confirm each individual account independently.


That does not make it worthless. It makes it a specific kind of evidence, with a specific kind of value, and this page will try to be as honest about its limits as it has been about everything else.


In early 2026, a major teachers' union in England surveyed its members and found that among primary school teachers — those working with many of the younger children this page has been most concerned with — 28 percent agreed or strongly agreed that they had personally observed a decline in their students' critical thinking skills that they attributed to AI use. Among secondary school teachers, working with older students, the number was considerably higher, closer to two-thirds. This was a large survey, involving thousands of teachers, and it received real attention when it was published.


It is worth being precise about exactly what this number does and does not represent, because a figure like this is easy to either dismiss too quickly or trust too completely. The survey was conducted among union members who chose to respond, not a scientifically selected sample of every teacher in the country, and the exact number of primary teachers who answered this specific question has not been made publicly available in a form that allows an outside reader to check it directly. The wording of the survey itself matters. Teachers were asked whether they had observed a decline 'due to AI usage,' which combines observation with attribution in a single question rather than separating the two — a phrasing that already assumes a cause, instead of first asking whether a decline was observed at all, and separately asking what the teacher believed caused it. That matters. A teacher answering this question is not simply reporting what they saw. They are reporting their own belief about what caused what they saw, and this page has been careful throughout to treat those as two different things.


None of that means the number should be ignored. It means it should be labeled honestly, the same way this page labeled Anthropic's own survey data several sections ago: this is what a substantial number of teachers believe they have witnessed, not a measured or confirmed outcome. A parent reading this page deserves to know that distinction plainly, not have it obscured by a number presented with more certainty than it has earned.


Separate surveys, conducted independently in different countries, describe a similar picture. An NPR/Ipsos poll of 545 American K-12 teachers, conducted in the spring of 2026, found that 54 percent said AI was making it harder for students to learn critical thinking skills — and nearly three in four said they believe AI's impact on education will be larger than that of the internet or computers before it. A large ongoing survey tracking American teenagers, RAND's American Youth Panel, found that AI use for schoolwork rose from 48 to 62 percent over just seven months in 2025, accompanied by a growing number of young people who themselves believed the tool was affecting their ability to think critically — not adults observing children, but young people describing a change they noticed in themselves.


Taken individually, none of these surveys would be strong enough evidence to build much of a case on. Taken together, they describe something worth taking seriously precisely because they come from different sources, using different methods, in different places, and point in a broadly similar direction: a meaningful number of the adults who spend their working lives with children, and a meaningful number of young people describing their own experience, believe they are watching something change. That kind of convergence across independent sources does not prove the underlying concern is correct. It does establish that the concern is not confined to researchers, AI companies, or the authors of a page like this one. It is something people watching real classrooms, in real time, are describing on their own, without being prompted by anything this page has argued.


It is also worth being honest about a limitation none of these surveys can address, no matter how many of them exist or how consistent they are with one another. A survey response, however sincere, captures a conclusion rather than the observations that produced it. It is still one step removed from an actual, specific, describable moment — a general impression formed over time, not a single, checkable event the way the last section's exam was checkable. What this page has not yet offered is something closer to that — not a number describing how many teachers share a general concern, but one teacher's account of what, specifically, she has watched change, in real students, across a career long enough to know the difference between an ordinary bad year and something else.


The evidence has now moved from laboratory studies to controlled classroom experiments to large surveys. The next section narrows the lens again — not to make the evidence stronger, but to make it more concrete. It turns to one experienced teacher, not a survey and not a study, describing what she believes she has watched change over decades in an elementary school classroom, offered honestly, in her own words, with the same caution this page has applied throughout: an observation is not proof, but it can be the beginning of an important question.

Part Fifteen: One Teacher's Account

Everything gathered until now has come from studies, companies, and surveys. This section is different. It is not evidence in the sense that the rest of this page has used that word. It is the account of Jeannine Germer, the author's wife and an elementary school teacher of thirty-seven years, speaking in her own words about what she has and has not noticed. It cannot rule out other explanations, and it does not try to. It is offered as testimony, not proof, including the parts that complicate this page's own argument.


That complication is worth stating first, honestly, rather than last. Asked directly whether kids are slower to start writing now than they were a few years ago, Jeannine's answer was not a confirmation. It was a correction: "Not if you're teaching properly." She does not personally sense a change in her own classroom, and she attributes that to how she teaches, not to the absence of AI's effect. From the beginning of her career, she has never simply told a child to start writing, nor has she skipped straight to giving them the answer. She builds up to independence — direction, ideas, structure, then a slow release toward doing it alone — a method she calls, in her own words, 'gradual release.' It is, in practice, the same sequence this page has argued for throughout: attempt, support, then independence — arrived at from a classroom, decades before AI existed, for reasons that have nothing to do with it. She still teaches that way now, deliberately, even knowing AI will likely be part of her students' futures. "I will not give in and not teach these things the way I used to," she said, "even though they may have some AI in their future."


That is not a page-friendly answer, and it is included because it is not. A teacher who carefully scaffolds writing may be protecting her students from exactly the pattern this page describes — which would mean the danger is real but not inevitable or universal. That possibility deserves to sit here plainly, not be edited out because it complicates the argument.


What she does notice is something narrower and more concrete than a general impression — and worth taking seriously precisely because it is describable, not intuited. Her students are taught a specific structure for every essay: a topic sentence, text evidence cited from a source, their own elaboration, and a closing sentence. "It's probably a little more immature compared to what maybe adults write," she said, "but it has a structure, because we're trying to teach the structure of writing." When an essay departs from that taught structure — when it reads more fluently, more maturely, with ideas "coming up... that are higher" than a fifth grader typically generates on their own — she reads that as a signal, not proof. "You know either a parent or AI has become involved in this thing," she said, "because yes, it's a better essay than they would've written — but that's not my point." Her point is not whether the finished product looks impressive. It is whether the child's own first attempt — however clumsy, however immature — is actually present in the work, or whether it was quietly skipped.


Asked what she would tell a parent to look for tonight, her answer was the same structural test, made practical: compare the essay against what was actually taught in class that week. Does it follow her structure, or does it skip ahead to something the child would not yet have reason to know on their own? "How did they come up with that idea when they're young?" is the question she said she'd leave a parent with — not as an accusation, but as an honest thing to notice.


Asked, finally, what she would say to a parent who wonders why any of this still matters if AI can write for a child anyway, she didn't hesitate: "Writing isn't just about writing an essay. You need to be writing about things because you truly understand... If you can write about it, you understand what you're learning. It helps you understand what you're learning." Her reasoning is not about AI at all. It is about what writing is for — an idea she arrived at from thirty-seven years in a classroom, independent of anything on this page, and stated more plainly than this page has managed to state it anywhere else.  

Part Sixteen: No Confirmed Case, No Fixed Deadline

No study reviewed on this page has followed one real child from a documented starting point through a measured period of AI use to a confirmed and lasting outcome, with every other possible cause ruled out along the way. As of this writing, the research reviewed on this page contains no confirmed, documented case of a child whose cognitive capacity has been measurably harmed by AI use. That is the honest limit of the evidence, and this page states it directly because everything built on top of it depends on saying so plainly.


It is also worth being direct about a second claim this page will not make: nothing in the research reviewed here establishes a fixed age at which a child's opportunity to build these capacities closes for good. Neither a hard causal claim nor a fixed developmental deadline claim goes further than the evidence currently supports.


What the evidence does support is more substantial than 'no proof' might suggest, and it is worth laying out in full rather than summarized away.


The same internal experiment that produced Part Eight's headline finding — a seventeen-point drop in later understanding among AI-assisted engineers — did not stop at that single number. When the researchers examined how each participant had actually used the AI, they identified six distinct patterns of interaction. Three of them, all involving genuine engagement — asking the tool to explain its reasoning, comparing its answer against the engineer's own attempt, using it to check work rather than produce it — preserved learning outcomes at levels matching or exceeding the group that used no AI at all. Only the patterns involving pure delegation showed the decline. That is not an absence of evidence. It is a specific, mapped mechanism, holding up across all six behavioral categories the researchers identified within that study — not one number averaging over what actually happened, but the same result appearing again and again as they looked closer. The seventeen-point average was not the whole story. It was hiding six different ones, and only one of them was bad.


Separate, smaller studies have begun looking directly at children using AI systems, rather than adults or general screen exposure. One preliminary study — Horowitz-Kraus et al. — used brain imaging to compare children's neural activity while using a chatbot with that of adults using the same tool, and found distinct engagement in the brain's attention and control networks. A separate 2025 study by Kim and colleagues, using brain imaging with children ages five and six, examined how young children's neural engagement changed depending on whether they interacted with an AI storytelling companion alone, with a parent alone, or with both together.


Both studies are real, both are specific to children and AI directly rather than adjacent evidence borrowed from adults or screens, and both are also small, unreplicated, and early enough that they should be read as leads rather than conclusions — the same standard this page has applied to every piece of evidence gathered.


Two of the largest AI companies in the world have looked at their own usage data and found enough to publish, unprompted, including a separate 2026 study from Anthropic examining nearly a million and a half real user conversations, which found that AI interactions can shape a person's sense of their own independent judgment, particularly in personal and consequential decisions — a different population and a different mechanism than the one this page has focused on, but coming from the same source, using the same rigor, and raising the same underlying question of what sustained reliance on these systems does to human autonomy.


A controlled study of adult professionals found a measurable gap in independent understanding once AI assistance was removed — and found that a specific, different way of designing the same tool prevented it. A separate controlled study with real high school students found the same pattern, with the same fix. Direct physical evidence shows a developing brain spending decades selectively keeping the connections it actually uses. A controlled experiment showed that the mere expectation of reliable outside help changes how hard a child's mind works, even when the help in question is nothing more than a piece of paper. Thirty-two of thirty-five college students turned in an assignment they had not read closely enough to notice that it made no sense. Teachers, in survey after survey, report watching something change in their own classrooms.


None of this, individually or together, proves that a specific child has been harmed. It is not meant to. What it establishes is that the concern this page has been describing is not a guess and not an invention. 


This is a real mechanism, mapped and tested, observed independently in adults, children, engineers, and teachers, by sources with no reason to coordinate and every reason to disagree — and they didn't. What is missing is not evidence. It is one thing: a single, fully investigated case, in a real child, with every other explanation ruled out. That gap is real, and this page will not pretend otherwise. It is also exactly the kind of gap society has often faced before deciding whether a developing risk deserved attention. Waiting for the first confirmed case and waiting for enough evidence to justify serious investigation have never been the same thing. It is the hardest, slowest kind of evidence to obtain — the kind that arrives last, after everything else has already pointed the same direction for years. Waiting for it before taking the rest seriously is not caution. It is a choice, and it is not a neutral one. 

Part Seventeen: Norway's Answer

Everything examined so far on this page has been about evidence — what has been measured, what has been observed, and what remains honestly unproven. This section is about something different: how one government has chosen to respond, in the absence of the very proof the last section spent so much care admitting does not yet exist.


In June 2026, the government of Norway announced that children in the first through seventh grades — roughly ages six through thirteen — would no longer be permitted to use generative AI tools in school, beginning with the school year that fall. Slightly older students, ages fourteen to sixteen, would still be allowed to use these tools, but only under a teacher's direct supervision, not independently. Prime Minister Jonas Gahr Støre explained the decision plainly, at a press conference announcing it: "The most important thing in school is that our children learn to read, write, and do mathematics." The government's stated concern was specific and closely parallels the concern this page has spent sixteen sections examining — that AI use risks letting young children skip foundational steps in their own learning before those steps have been properly built.


It is worth being precise about what this decision is, and what it is not. This was not Norway's scientific establishment announcing a confirmed finding. It was Norway's government making a precautionary choice, using language that matches the honest, hedged tone this page has tried to hold throughout — 'risks,' not 'causes.' It shows that the people responsible for those children's education looked at evidence broadly similar to that reviewed on this page, together with whatever additional advice and analysis, available to them through their own advisors, informed its own decision-making — and decided that waiting for a confirmed case, of the kind Part Sixteen just described as still missing, was not an acceptable position to hold in the meantime.


This was not Norway's first move of this kind. In 2024, the government had already restricted smartphone use in schools, citing similar concerns about children's attention, well-being, and development. Subsequent reporting on that earlier restriction described lower rates of bullying and higher grade averages in the schools that adopted it, compared to before the ban took effect. That is not, on its own, proof that the phone restriction caused those changes — schools change in many ways at once, and this page will not claim a cleaner causal story than the evidence actually supports. But it establishes something worth knowing regardless: this is not a government making a single, isolated, untested decision about AI in a vacuum. It is a government that had already taken a similar precaution once, observed what happened afterward, and chose to act again in a similar way.


This is worth connecting directly to the central tension Part Sixteen ended on. That section argued that waiting for a fully confirmed, individually investigated case before taking a concern seriously is not a neutral choice — it is itself a decision, with its own consequences, made by default rather than deliberately. Norway's government appears to have adopted a similar precautionary approach through an entirely different process: not academic argument, but the ordinary work of governing, weighing available evidence against the cost of waiting, and choosing to act on the side of the children currently in its schools rather than the children who might benefit from a study that has not yet been done.


A national government, with its own researchers, its own access to expert advice, and its own direct accountability to the parents of the children affected, examined an evolving body of evidence addressing similar concerns to what this page has gathered and concluded that the honest uncertainty described in the last section was still sufficient to act on. That is a real, consequential fact, not a footnote — a sovereign government chose to restrict a technology for millions of children rather than wait for the confirmed case that does not yet exist. It would still be premature to call this proof that the underlying concern is correct. One country's policy is not a scientific verdict, and governments sometimes make wrong precautionary calls, just as they make right ones. But that limit cuts both ways. It would be an equal mistake to dismiss what Norway did simply because it did not arrive wearing the clothing of a peer-reviewed study.


That is not proof, and it should not be mistaken for it. Nothing in this page's sourcing confirms exactly what officials weighed in reaching this decision. What can be said is only what they did: a responsible government, acting deliberately rather than speculating, chose not to let a generation of schoolchildren serve, in effect, as an unmonitored test case while a confirmed answer was still years away.


The next section considers a different kind of precedent for such a choice — not a government acting today, but a moment, seventy years ago, when a different kind of evidence, gathered under a different kind of uncertainty, was met with a similar response. 

Part Eighteen: Why We're Speaking Now

The last section closed on a comparison worth taking seriously on its own terms, not just as a rhetorical device: a moment, seventy years ago, when a different kind of evidence, gathered under a different kind of uncertainty, raised the same underlying question Norway has now confronted: how should society act when meaningful evidence exists but certainty does not? This section explains that comparison directly, because it is not a passing reference. It is the reason this page exists at all, rather than waiting for a kind of certainty that may still be years away.


In the early 1950s, researchers on two continents, working independently of one another, published findings linking cigarette smoking to lung cancer. British researchers Richard Doll and Austin Bradford Hill published their first major study in 1950. American researchers Ernst Wynder and Evarts Graham published similar findings the same year. Neither team knew what the other was doing.  Working independently, both research teams reached broadly similar conclusions: people who smoked were dying of lung cancer at rates far higher than people who did not.


That evidence did not settle the matter, and it is worth being honest about why, because the reasons are directly relevant to everything this page has tried to do. The tobacco industry did not simply ignore this research. In December 1953, the major American tobacco companies met and, within weeks, published a full-page advertisement in newspapers across the country — "A Frank Statement to Cigarette Smokers" — announcing they would fund their own research into the question. Decades later, internal industry documents surfaced during litigation showing that the companies' own scientists had privately understood the danger years before that public statement, and that the research effort they announced was designed to manufacture the appearance of an open question, not to resolve one.


It took until 1964 — fourteen years after Doll and Hill's first published findings — for the United States Surgeon General to issue a report stating plainly that the link between smoking and lung cancer was established. Fourteen years passed between the point at which real, published, peer-reviewed evidence existed and the point at which an institution with enough authority to be widely believed said so publicly, without equivocation.


It is worth being precise about what changed between 1950 and 1964, because it was not the underlying facts. The 1964 report did not primarily rest on newly discovered evidence. It reviewed evidence that had, in large part, existed for over a decade and declared what that evidence had already shown. What changed was not the science. What changed was that enough institutional weight finally stood behind saying so.


This page is not a Surgeon General's report, and it does not claim to be one. Nothing gathered here carries that kind of institutional authority, and this page has tried, from its opening pages, never to claim more certainty than it has earned. But the comparison is not really about authority. It is about timing, and about what the fourteen-year gap actually cost.


The people who took Doll and Hill's early findings seriously and said so publicly before 1964 were not vindicated by hindsight alone. They were right the entire time they were being dismissed as alarmist, or premature, or insufficiently proven. Fourteen years of continued smoking, among people who might have made a different choice with different information, is not an abstract cost. It is measured in a specific, documented rise in lung cancer deaths across that same period — a cost paid, in full, by people who never got to weigh the evidence that had already been published, because the institutions capable of making that evidence widely known chose, for reasons that had nothing to do with the science, to wait.


This page does not know whether the concern it has gathered evidence for will eventually be confirmed, the way the link between smoking and lung cancer eventually was. It has, throughout, tried to be honest about that uncertainty rather than pretend it away. The lesson this page draws from that period is not "wait for certainty before you say anything." It is closer to the opposite: real, published, carefully gathered evidence is worth stating plainly, with its actual limits clearly attached, long before an institution with the authority of a Surgeon General is prepared to declare the matter closed — because if the concern turns out to be justified, the years spent waiting for that declaration are not neutral years. They are the years during which the harm, whatever its true extent turns out to be, continues to accumulate, among children who will not get the chance to be told what was already known while it might still have mattered to them.


That is the reason this page exists in its current form — not as a verdict, and not as an alarm, but as a plainly stated account of evidence gathered now, published now, with its limits fully attached now, rather than held back until a kind of certainty arrives that may take another decade, or may never arrive in a form clean enough to satisfy every skeptic. It is not an attempt to announce a conclusion before the evidence supports one. It is an attempt to ensure that the evidence that already exists — along with its strengths, limitations, and unanswered questions — is available to parents while meaningful choices remain. Waiting for certainty is itself a choice. This page argues only that families deserve the opportunity to make their own choices before that window closes.


The next several sections turn from why this page speaks now to what, concretely, a family can do with what is currently known — starting tonight, in an ordinary house, with an ordinary child, long before any larger question is ever formally resolved. 

Part Nineteen: What a Family Can Do Tonight

Everything gathered across this page adds up to a specific, testable principle, not a vague warning. It is worth stating plainly, one more time, before turning to what a family can actually do with it: AI assistance appears to help a child's thinking when it requires the child to attempt something first and stay engaged with the result. It appears to weaken a child's thinking when it replaces that first attempt entirely. That distinction has now shown up, independently, in adult engineers, in high school math students, in a Harvard physics classroom, and in a ten-year-old memorizing a list of words. It is not a theory waiting for evidence. It is the evidence, stated as simply as possible.


What follows is not a set of rules requiring AI to be removed from a child's life. Nothing in this page has argued for that, and nothing in the research reviewed here supports it. It is one evidence-informed way of using AI that keeps a child on the side of that distinction where the evidence consistently points to better outcomes, rather than the side where it consistently points to worse ones.


The order matters more than the amount.


The single clearest finding across every controlled study on this page is not about how often AI gets used. It is about when, relative to a child's own effort. A simple sequence, applied consistently, captures what the evidence actually supports: the child attempts first, AI assists second, and the child demonstrates independent understanding last.


Before a child opens an AI system for a piece of schoolwork, they should produce something of their own first — a guess, an outline, a first sentence, even one they already suspect is wrong. It does not need to be good. Its only job is to make sure the child, not the AI, is the one who starts. This single step is what separates the guided condition from the unrestricted condition in the high school math study discussed earlier on this page — the students who kept this first step intact were the ones whose independent understanding held up once the AI was removed.


While using AI, ask for less than it is willing to give.


An AI system asked to solve a problem outright will usually do exactly that. An AI system asked for a hint, a question, or a specific error to look for will do that instead — but only if it is asked. This is the same distinction that separated the guided AI condition from the unrestricted one in the classroom study, and the same distinction that separated the engineers who held their own from the engineers whose independent scores dropped. It is not about which AI system is used. It is about what it is asked to do. A general-purpose AI system can be used either way. The outcome depends on the question, not the tool.


After AI is used, it should disappear briefly and deliberately.


The single most useful thing a parent can do costs nothing and takes very little time: after a child uses AI to help with something, put the device away, and ask the child to explain it back, in their own words, without it. Can they describe what they did, and why it works? Can they solve a similar problem with slightly different numbers or details? If they cannot, the assignment may be finished, but the learning it was meant to produce may not have happened yet — and that is worth knowing tonight, not at the end of a school year.


Some activities are worth preserving, not because AI is dangerous, but because this page's evidence suggests they do the quietest, most necessary work. 


Reading a real book, start to finish, without a summary standing in for it. Mental math, even when it is slow and a little wrong at first. Making up a story with no help. Sitting through boredom long enough to come up with a way out of it themselves. Ordinary conversation with the people in the room. None of these need to be AI-free forever, or in every household, on every occasion. They need to remain places where a child's own effort is still the whole of what happens often enough that the habit of trying does not quietly disappear.


None of this requires perfection.


A family that follows this loosely, most nights, imperfectly, is very likely doing something meaningfully different from a family that never thinks about the order at all — the same way the engineers who engaged with the AI's reasoning even some of the time still held their scores, where the ones who never engaged did not. This is not a test with a passing grade. It is a habit, built the same way every other habit in a household gets built: unevenly, with room for a bad night, as long as the good nights outnumber them.


What none of this can do — and this page will not pretend otherwise, consistent with everything said in the two sections before this one — is tell a parent with certainty whether a specific pattern in their own child reflects this mechanism, an ordinary hard week, or something else entirely that deserves a different kind of attention. The purpose of everything in this section has been modest. It is not to help parents diagnose a problem. It is to help them protect a way of learning that the evidence consistently favors. The final section turns to a different question entirely: not what to try at home, but what would mean it is time to stop trying things at home, and call someone else instead.  

Part Twenty: What to Watch For, and When to Call a Doctor

This page opened with an ordinary moment: a child asking an AI to finish a thought they were about to have, and a parent feeling nothing more than relief that the evening's homework was done. Everything since then has been an attempt to give that moment back its proper weight — not to make it frightening, but to make it visible, the way it usually is not.


This last section is the most practical one on the page, and also the one that needs the most care, because a list of warning signs is exactly the kind of thing that can do real harm if it is handed to a parent without its limits attached. Part Sixteen said this plainly, and it needs to be said again here, at the moment it matters most: nothing in this page, and no list that follows, can tell a parent with certainty what is causing a change in their own child. That is not a weakness in this page. It is the honest truth about what this evidence can tell a parent about their own child — and what it cannot.


With those limits firmly in mind, here is what is worth watching for — not as a diagnosis, but as the kind of thing worth noticing, the way Part One asked a parent to notice an ordinary evening a little more closely.


Watch for the moment right before a task begins, more than the task itself. A child who sits with a blank page, a hard problem, or an awkward silence for a little while before doing something about it is doing something this page has spent twenty sections describing as valuable. A child who has stopped having that moment at all — who consistently reaches for something else to fill it before they have tried anything themselves — is worth paying quiet attention to, not because it proves anything, but because it is the exact moment this page has argued matters most.


Watch for a gap between what a child produces and what they can explain. A finished assignment that a child cannot describe in their own words, cannot walk through the reasoning behind, or cannot attempt again with slightly different details, is worth noticing over time — not because the assignment is bad, but because it suggests the thinking the assignment was meant to build may not have happened, even though the paper looks fine.


Watch for how a child handles an ordinary, unrelated kind of friction — a dead device, a missing pencil, a rule they don't like, a moment of boredom with nothing to do. A child who has always found a way through those moments and has recently stopped trying is worth noticing, just as a teacher with decades of experience would, before a parent might.


None of these signs, alone or together, can tell a parent what is causing them. A change like this can come from ordinary tiredness, a hard stretch at school, a friendship falling apart, or something else in a child's life that has nothing to do with any of the evidence gathered on this page. It can also, sometimes, come from something that needs real medical attention rather than a change in household habits — a mood or attention condition, a sleep problem, or something else a pediatrician is equipped to notice that a parent, however attentive, is not trained to rule out on their own.


Here is the one thing worth being completely clear about, because it matters more than anything else in this section. A change that shows up only around schoolwork, comes and goes, and responds to the kind of household changes described in the last section, is the kind of thing worth watching and adjusting for at home. A change that shows up everywhere — not just around AI or schoolwork, but in a child's sleep, mood, friendships, and everyday behavior — or a change that is sudden, severe, or does not improve even after real changes are made at home, is no longer a question this page, or any household adjustment, can answer. That is the moment to talk to a pediatrician, not because something is certainly wrong, but because that is exactly the kind of judgment a parent should not have to make alone, based on a page they read one evening.


This page began with an ordinary Tuesday evening. It ends in the same place. It ends by asking something much smaller than certainty: simply to protect the child's first attempt. Because long before AI changes what a child can produce, it may quietly change what a child believes they can do without it.


Stay Sovereign.


Jim Germer 

August 6, 2026

© 2026 Jim Germer - The Human Choice Company LLC. All Rights Reserved.

Powered by

This website uses cookies.

We use cookies to improve your experience and understand how visitors use our website so we can make it better. 

Accept