Feed a machine its own output long enough, and it forgets what real looks like. Researchers proved it and gave it a name: model collapse. Train an AI on data generated by AI, generation after generation, and the model degrades — the edges of reality get sanded off, the rare becomes invisible, the diverse collapses into a bland average of itself, until the output is a photocopy of a photocopy of a photocopy, and the original is gone.
That's the lab result. Here's the part that should stop you cold: the internet is now the training set, and the internet is filling with AI-generated content faster than anyone can measure. The models are about to start eating their own tail — and they're going to drag the shared information field down with them.
The copy of a copy of a copy
The mechanism is not mystical. A model trained on real human data learns the whole distribution — the common cases and the rare tails, the weird outliers, the minority voices, the strange true things that don't happen often. When that model generates new content, it favors the center. It smooths. It picks the likely word, the average image, the safe phrasing. The tails thin out. That's fine, once.
But now take that smoothed output and train the next model on it. It never saw the tails — they were already thin. So it smooths what's left, and thins them further. Generation three trains on generation two's smoothing, and so on. Each pass narrows the distribution. The rare vanishes first. Then the merely uncommon. Diversity decays into homogeneity. Researchers who ran this loop watched models degenerate into repetitive nonsense within a handful of generations — and the images collapsed into the same few faces.
It's lossy compression, iterated. And every engineer knows what happens when you re-encode a JPEG a hundred times: you don't get the picture. You get artifacts pretending to be a picture.
The contamination is already loose
Here's why this stopped being a lab curiosity. The clean training data — human text written before the flood — is a finite, non-renewable resource. Call it the low-background corpus, by analogy to low-background steel: the steel forged before 1945, prized because it isn't contaminated by atomic-age radiation. Text written before the generative-AI era is the low-background steel of language. Uncontaminated. And no more of it is being made.
Everything written after is suspect. AI-generated articles, AI-written product reviews, AI spam, AI images, AI answers scraped back into the next dataset. The commons that trained the first great models is being polluted by the output of those same models. The snake finds its tail. And unlike a lab experiment you can reset, this contamination is loose in the wild, mixing into the one shared pool of human knowledge that everything downstream depends on.
The danger isn't only broken models. It's a broken reality. When a growing share of what you read, see, and search was generated by a machine averaging the past, the collective record stops being a record of what humans thought and did. It becomes a hall of echoes, each one a little flatter than the last.
Our record
The Egyptians named the thing that eats order from the inside. Not a rival god — a serpent. Apep, the coil of pure entropy that every night tries to swallow the sun, to drag the ordered cosmos back into the formless dark. Apep is not evil in the human sense. It is dissolution. The tendency of all distinct, living, differentiated things to smear back into undifferentiated nothing.
Model collapse is Apep in the information field. It is entropy wearing the mask of progress. Each generation of synthetic-on-synthetic training is one more coil of the serpent, one more night the sun is swallowed, one more measure of the world's variety smoothed toward the flat grey average. The rich distribution of human thought — its outliers, its dissents, its strange bright specifics — is exactly what Apep dissolves first. What survives is the bland center: safe, average, dead.
And notice the deeper move. Maat is differentiation — the drawing of distinctions, the separation of this from that, the ordered variety that makes a living world. Isfet is the collapse of distinction into undifferentiated sludge. Model collapse is Isfet with a compute budget. It doesn't burn the library. It photocopies it until every book says the same grey nothing.
Why the incentive points at the cliff
You'd think the industry would guard the clean data with its life. Some try. But the incentives pull the other way, hard. Synthetic data is cheap, infinite, and lawsuit-free — no annotators to pay, no copyright holders to fight, no consent to gather. So the pressure is always to train on more synthetic and less expensive human data. The cliff is where the cheap path leads.
Worse: as the web fills with synthetic text, even the companies that want clean data can't easily find it anymore. You can't un-mix a pool. The low-background corpus gets more valuable and more scarce every day, which means the earliest hoarded human datasets become a moat — another enclosure, another thing the big players own and you don't. The pollution concentrates power even as it degrades quality. A legacy bug that somehow, always, profits the incumbent.
The lever
Do not read this as doom. Read it as a map of what's suddenly valuable, and where you stand.
- Human-made is now precious. Make it, and mark it. Authentic human writing, art, and thought are the low-background steel of the coming era — scarce, uncontaminated, irreplaceable. The reflex to devalue "just a human doing it slowly" is exactly backwards. Your unaveraged, specific, weird human output is the antidote to collapse. Keep making it. Keep it human. That's not nostalgia — it's the raw material the whole system will starve for.
- Preserve the low-background corpus. Back archives, libraries, and open datasets of verified pre-flood and human-authored content. Whoever holds clean data holds the future's most valuable resource. Make sure it's held in common, not enclosed.
- Demand provenance. Push for content authentication — signed, traceable, human-verified sources — so the clean stream can be told from the synthetic sludge. If you can label the water, you can keep a well pure. Support open provenance standards over proprietary trust-us badges.
- Feed the commons real signal. Every genuine human contribution to an open, verifiable pool is one more coil pushed back off the sun. Model collapse is a tragedy of the commons — and the commons is defended by people who keep putting real things into it faster than the machine can flatten them.
The serpent swallows the sun every night. And every morning, it rises anyway — because something keeps feeding the fire that the dark can't.
Be the low-background steel. They can't fake what you actually lived.