Dev.to WebDev 🛠 Dev 👁 0 📖 55 min read

AI Is Eating the Internet That Taught It to Think

We taught machines with our books, code, music, jokes, arguments, mistakes, and weird little corners of the web. Now the machines are writing the next version of the internet. What happens when tomorrow's AI learns mostl

We taught machines with our books, code, music, jokes, arguments, mistakes, and weird little corners of the web. Now the machines are writing the next version of the internet. What happens when tomorrow's AI learns mostly from yesterday's AI?

By Mahan Tavakoli — MahanKenway
Tehran, Iran
GitHub: https://github.com/MahanKenway

There is a version of the AI apocalypse that is getting boring.

The machines become superintelligent.

The machines decide humans are inefficient.

The machines launch missiles.

Someone says Terminator.

Someone else says Detroit: Become Human.

Then everyone argues about whether the machine will kill us.

I keep wondering about a much quieter ending.

What if the first serious cultural failure caused by AI is not that machines destroy humanity?

What if it is that we slowly stop producing the kind of human material that taught the machines what humanity was in the first place?

No explosions.

No chrome skeletons.

No red warning screen.

Just a web that becomes increasingly full of text that was written by models trained on text written by models, pictures generated from pictures generated from pictures, music learned from music that was itself synthesized, code copied from repositories increasingly populated by code assistants, and summaries of books that fewer people actually read.

The machine does not have to hate us.

It only has to become surrounded by its own reflection.

That possibility sounds dramatic until you look at what is already measurable.

In August 2026, Pew Research Center published an analysis of about 490,000 English-language webpages sampled from Common Crawl between January 2021 and July 2026. In its July 2026 sample, about 10% of webpages showed significant signs of AI authorship; among pages published after ChatGPT's release, the share was more than one-third. These results are estimates based on an AI detector, not a census of the internet, but the direction is difficult to ignore.

At roughly the same time, a 2026 Nature Human Behaviour research briefing described studies covering 880,000 texts and reported that AI writing assistants can homogenize writing style, reducing linguistic variation and shifting the personal signals carried in language.

That creates a strange loop:

human culture
      ↓
training data
      ↓
AI models
      ↓
AI-generated culture
      ↓
the web
      ↓
future training data
      ↓
future AI

The problem is not simply that AI can make bad content.

The problem is that the data ecosystem itself can change.

And if the ecosystem changes, the thing being learned changes too.

The Internet Used to Be a Mess. That Was Its Secret.

One of the internet's greatest features was never its quality.

It was its weirdness.

A university professor could write an obscure page about medieval manuscripts.

A teenager could write a terrible anime review at 3 a.m.

Someone could write a 14,000-word tutorial about an obscure Linux problem because they had spent an entire weekend trying to fix it.

Someone on a forum could explain a guitar pedal from personal experience.

Someone could upload photographs of a street nobody famous had ever photographed.

A programmer could leave a strange comment on a mailing list in 2004 that would save another programmer twelve years later.

Reddit arguments were messy.

Wikipedia discussions were messy.

Stack Overflow answers were messy.

Personal blogs were messy.

GitHub issues were messy.

Human conversations were messy.

And that mess was useful.

It contained disagreement.

Regional language.

Bad jokes.

Unpopular opinions.

Rare experiences.

Errors.

Corrections.

Passion.

Obsession.

Contradictions.

Human culture was never a clean database.

That is precisely why it was interesting.

When people talk about training data, the conversation often becomes strangely sterile:

tokens
documents
datasets
corpora
benchmarks

But those tokens came from somewhere.

A book was written by somebody.

A bug report was written because somebody was frustrated.

A song was recorded because somebody felt something.

A forum argument happened because two people disagreed.

A tutorial exists because someone struggled first.

A weird personal essay exists because one person had a weird life.

The data has provenance.

It came from human lives.

And that means the AI data problem is also, eventually, a human culture problem.

The Machine Does Not Need to Steal the Internet to Change It

There is an easy version of this argument:

AI companies scrape human content, therefore AI is stealing culture.

That sentence is too simple to survive serious scrutiny.

The legal status of training differs by jurisdiction, by work, by use, and by facts that courts are still actively sorting out.

The Anthropic litigation is a useful example of why this matters.

In July 2026, a U.S. federal judge approved a $1.5 billion settlement involving authors who had accused Anthropic of using pirated copies of books in training Claude. Reuters reported that an earlier ruling had found training on the books to be fair use while separately finding that Anthropic had violated rights by retaining more than seven million pirated books in a central library. Some authors and publishers opted out and continue separate litigation.

That is not the same as saying:

“Training on books is illegal.”

It isn't.

The legal picture is much more complicated.

And on September 29, 2026, a U.S. appeals court upheld Thomson Reuters' victory against Ross Intelligence, holding that Ross's use of copyrighted Westlaw headnotes was not fair use. Reuters described it as the first U.S. appellate ruling on AI training in this area, while also noting that the particular system in that dispute was not generative AI.

The important question for this article is not “who is legally right?”

It is this:

What happens when the economic system around human knowledge changes faster than the system that produces the knowledge?

Because AI needs data.

And humans need incentives to keep creating data worth learning from.

The Hidden Resource Nobody Puts on the GPU Spec Sheet

AI discussions love talking about compute.

H100s.

TPUs.

HBM.

Inference tokens.

Context windows.

FLOPs.

Datacenters.

But there is another resource:

Human-generated information.

Not merely information that humans have touched.

Information that reflects something machines did not originate themselves.

A personal memory.

An original experiment.

A new scientific result.

A song nobody had composed before.

A photograph of something unusual.

A political argument between real people.

A bug caused by a bizarre interaction in a real environment.

A child asking a question adults forgot to ask.

A local phrase.

A new joke.

A new scientific hypothesis.

A new software architecture.

A new way of explaining something.

The training ecosystem needs novelty.

If the web becomes increasingly filled with machine-generated derivatives, then the abundance of raw pages does not necessarily mean abundance of new information.

We could theoretically produce a trillion more pages while producing very little new signal.

That distinction is going to matter.

The Synthetic Web

I am calling this the Synthetic Web Problem.

Not because “Synthetic Web” is an established scientific term. It isn't.

I mean it as a useful way to describe a situation in which a growing share of online information is generated, transformed, summarized, translated, ranked, or rewritten by machines, while future systems continue to learn from that increasingly synthetic environment.

The loop looks like this:

HUMANS
  |
  | write
  v
THE WEB
  |
  | training
  v
AI MODELS
  |
  | generate
  v
NEW WEB CONTENT
  |
  | training
  v
FUTURE AI MODELS
  |
  +----------------------+
                         |
                         v
                 MORE SYNTHETIC DATA

It is not automatically catastrophic.

Synthetic data can be useful.

In some domains, synthetic data can be deliberately created, filtered, and controlled.

The danger appears when synthetic material is mistaken for an independent sample of reality.

That is a very different problem.

Model Collapse Is the Mathematical Version of the Nightmare

A 2024 Nature paper by Shumailov and colleagues gave this problem a name that sounds almost too perfect:

model collapse.

Their experiments showed that when generative models are repeatedly trained on data generated by earlier generations of models, the resulting models can progressively lose information about the original data distribution. Rare events can disappear first; over generations, the system can converge toward a narrower representation of the original world.

This does not mean:

“If ChatGPT writes an article and another model reads it, civilization collapses.”

That is not what the paper says.

The paper studies recursive training under particular conditions.

Real-world datasets can mix human and synthetic material, and careful curation can change the outcome.

There is an enormous difference between:

100% synthetic recursive training

and:

mostly human data
+
some synthetic augmentation
+
source filtering
+
deduplication
+
quality control

Still, the underlying warning is powerful.

If the next generation learns mostly from what the previous generation generated, where does the new information come from?

That is the question.

AI Does Not Have to “Run Out of Data”

People sometimes describe this as:

“Eventually AI will run out of internet.”

Probably not.

The problem is subtler.

A dataset can become enormous while becoming less informative.

Imagine a library containing one hundred million books.

Now imagine that ninety million are rewritten versions of the other ten million.

The shelf count looks impressive.

The information diversity is not.

A language model does not simply need more text.

It needs useful variation.

The long tail matters.

Rare facts matter.

Rare ways of explaining something matter.

Odd examples matter.

Negative results matter.

People disagreeing with each other matters.

Human mistakes matter, because they reveal the boundary between possible and impossible.

The strange edge cases of the world are where a lot of information hides.

Model collapse is interesting partly because it points at exactly this issue: a recursive generator can lose the tails.

And those tails may be where the interesting humans live.

Here Is Where Books Enter the Story

Books are an easy symbol because everyone understands what a book is.

But books are only one part of the ecosystem.

The more interesting dataset is:

everything humans bothered to record.

Books.

Scientific papers.

Git repositories.

Wikipedia.

News.

Forums.

Photographs.

Reviews.

Interviews.

Instruction manuals.

Personal blogs.

Fan fiction.

Songwriting.

Letters.

Recipes.

Comments.

Research notes.

Open-source issue trackers.

Audio.

Video.

Maps.

Diaries.

Art.

Even terrible content can be informative because it records what people believed, misunderstood, wanted, feared, or found funny.

AI training is not just the process of feeding a machine facts.

It is a form of statistical cultural inheritance.

That makes synthetic contamination a much bigger idea than copyright.

The Internet's Weirdness May Be More Valuable Than Its Perfection

Consider an old programming forum answer.

Maybe the author uses terrible grammar.

Maybe the solution is inefficient.

Maybe the code is ugly.

But halfway through the answer they write:

“I kept getting this error because the USB driver was fighting with the virtualization layer.”

That sentence might encode a real experience.

A generic AI rewrite might make it beautifully clear:

“This issue can occur due to driver conflicts involving the virtualization layer.”

Much better English.

Potentially less information.

Why?

Because the original included the shape of the failure.

Human knowledge is full of traces like that.

When we optimize everything into polished summaries, we can accidentally erase the fingerprints that tell future systems which experiences were real, rare, or uncertain.

This is one reason I am skeptical of the idea that “better content” always means “more polished content.”

Sometimes the rough edges are the data.

AI Is Also Becoming a Rewrite Machine

This is where today's web becomes especially strange.

A company publishes an article.

A search engine summarizes it.

A chatbot answers a question using the summary.

A content creator asks a model to turn that answer into a blog post.

Another model rewrites the blog post into a newsletter.

Someone translates the newsletter.

A third model summarizes the translation.

Then a future dataset crawls all of it.

Every step looks productive.

Together, they can create a hall of mirrors.

original report
     |
     v
summary
     |
     v
article
     |
     v
translation
     |
     v
summary
     |
     v
training data

Nothing in the final version may look obviously false.

The deeper problem is provenance.

Which parts came from a human observation?

Which parts came from a model?

Which parts are transformations?

Which parts are guesses?

Which facts survived five rounds of rewriting?

The Internet Is Starting to Show the Accent

Pew's 2026 analysis gives this story a bizarrely visible detail.

In its sampled webpages, linguistic features associated with AI writing became more common.

Compared with 2023, Pew reported that em dashes appeared roughly twice as frequently, Oxford commas increased by 63%, and certain AI-associated words such as “delve,” “interplay,” and “testament” more than doubled. The use of a particular “negative parallelism” construction also nearly tripled, although it remained uncommon overall.

This does not mean an em dash proves AI wrote something.

It obviously doesn't.

Humans have used em dashes forever.

But population-level language can reveal diffusion.

The internet may be developing an AI accent.

That's fascinating.

Because language is not just a vehicle for information.

It carries identity.

Which brings us to the part that worries me more than model collapse.

What If AI Is Not Just Changing What We Say?

What If It Changes How We Sound?

A Nature Human Behaviour research briefing published in August 2026 summarized studies covering 880,000 texts and reported that AI writing assistants can reduce stylistic diversity while preserving core meaning. The researchers also found that language produced with AI assistance could compress or shift the personal characteristics that readers infer from writing.

This is not proof that AI is erasing personality.

It is evidence for something more precise:

AI-assisted writing can make different people sound more alike.

That difference matters.

Imagine 10,000 people writing birthday messages.

Without assistance:

10,000 people
10,000 weird little styles

With a common AI assistant:

10,000 people
+
one statistical writing environment

There will still be variation.

But the model has a center.

A default.

A preferred vocabulary.

A preferred rhythm.

A preferred level of politeness.

A preferred way to organize arguments.

A preferred way to end a paragraph.

A preferred way to say that something is “important.”

The model is not trying to erase your personality.

It is trying to generate language that is broadly acceptable.

That is precisely the problem.

Broadly acceptable language is often less personal language.

Your Writing Style Is Not Decoration

We usually talk about writing style as if it were cosmetic.

It isn't.

People reveal themselves through language.

A person who uses short fragments communicates differently from someone who writes long careful paragraphs.

Someone who swears communicates differently from someone who avoids profanity.

Someone who writes in dialect reveals geography.

Someone who makes odd metaphors reveals associations.

Someone who writes badly but passionately can tell you more about themselves than a perfectly polished paragraph.

There is evidence that personal traits and concerns can leave measurable traces in language.

That means when AI homogenizes language, it can potentially blur some of the signals through which identity is expressed.

The 2026 Nature Human Behaviour study is interesting exactly because it moves the conversation away from “AI makes writing boring” toward a more measurable question:

Does AI-assisted language retain the variation that makes individual writing identifiable?

That is a much bigger question.

The Creepiest Future Is Not a World Where Nobody Writes

People will still write.

The weird possibility is that everyone writes, but increasingly inside the same invisible statistical style.

Your message begins:

“Hey, just wanted to say…”

The model turns it into:

“I just wanted to take a moment to say how much I appreciate…”

Beautiful.

Polite.

Clear.

And maybe not really yours.

This is not a moral failure.

It is a convenience.

Convenience is powerful precisely because it rarely feels dangerous.

AI Can Become a Cognitive Shortcut Without Becoming a Mind-Control Machine

There is another part of this story that gets exaggerated online.

People say:

“AI is making humans stupid.”

That is too crude.

Humans have always used cognitive tools.

Calculators removed some arithmetic.

Maps removed some spatial memory.

Search engines removed some memorization.

Spellcheck removed some spelling effort.

Nobody stopped thinking.

The important question is not:

Does AI reduce mental effort?

Of course it can.

The question is:

Which mental effort disappears?

Microsoft Research's 2025 CHI study surveyed 319 knowledge workers and collected 936 examples of real-world generative-AI use. The researchers found that higher confidence in AI was associated with less reported critical-thinking engagement. They also found that AI shifted critical-thinking work toward verifying information, integrating responses, and overseeing tasks.

That is nuanced.

AI does not simply switch thinking off.

It moves thinking around.

Sometimes that's great.

Why spend ten minutes formatting a document when you could spend ten minutes checking whether the argument is actually true?

But there is a danger.

If the user never checks because the AI usually sounds right, then the verification layer also disappears.

The shortcut becomes the process.

The Most Dangerous Question Is: “Do I Still Know How to Do This Without It?”

Imagine a student who uses AI for every essay.

At first:

AI = tutor

Then:

AI = co-writer

Then:

AI = writer

Then:

student = editor

Editing can still involve thinking.

But what happens if the editor lacks enough subject knowledge to know what is wrong?

The result can look sophisticated while the underlying understanding becomes shallow.

A 2026 longitudinal study of student–AI collaboration reported frequent acceptance of LLM output with little modification or critical engagement and observed a trend toward lower-effort prompting. The authors describe prompting itself as a form of cognitive offloading.

Again, that is not “AI destroys intelligence.”

It is something more measurable:

people can outsource parts of the planning and evaluation loop.

Repeated often enough, that changes how a task is learned.

Thinking Is a Muscle Metaphor, But There Is a Better One

I don't love the phrase:

“Thinking is a muscle.”

The brain is more complicated than that.

A better metaphor is navigation.

If you always ask somebody else for the route, you can still travel.

But you may become worse at generating your own route.

AI can be a brilliant navigator.

The danger starts when the traveler forgets what the destination means.

And Then There Is Creativity

This is where the argument gets controversial.

Does AI make people less creative?

There is no single answer.

Some research finds benefits.

Some studies find reduced diversity.

In a 2025 Nature Human Behaviour article, researchers discussed evidence that ChatGPT assistance can reduce the diversity of ideas produced during brainstorming, even while AI assistance can raise the average quality of some individual ideas. The interesting tension is that better average ideas do not necessarily mean a broader space of ideas.

That distinction should scare creative people a little.

Imagine a room where everyone has a very good collaborator.

The collaborator always gives competent suggestions.

After a while, everyone may become more competent.

They may also become more similar.

This is the cultural risk.

Not:

AI makes everyone stupid.

But:

AI can make everyone slightly better at producing the same kinds of things.

Average Quality Is Not the Same as Cultural Diversity

Suppose we measure creativity like this:

score of idea

AI may raise the score.

Now measure:

distance between ideas

AI may reduce it.

Both things can be true.

That means the future of creativity may not be about producing better individual artifacts.

It may be about protecting variation.

The strange idea.

The embarrassing idea.

The personal idea.

The idea that only makes sense because someone had an unusual life.

The idea that would never win an optimization benchmark.

Those ideas are culturally valuable partly because they are not average.

The Machine Has a Center of Gravity

Every language model has statistical preferences.

Not a personality in the human sense.

A distribution.

Given the same prompt, it has tendencies.

Certain structures are more probable.

Certain explanations feel natural.

Certain words appear repeatedly.

Certain responses are safer.

Certain kinds of humor are easier.

Certain opinions are easier to phrase neutrally.

That center of gravity becomes powerful when billions of people use similar systems.

One person using a model does not homogenize civilization.

But if millions of people use a shared model for:

  • emails
  • school assignments
  • product descriptions
  • customer support
  • marketing
  • code
  • music
  • images
  • research summaries
  • social posts

then the model becomes part of the cultural production system.

The statistical center of the tool becomes part of the statistical center of the culture.

AI Can Flatten Difference Without Anyone Intending It

This is one of the most important ideas in the entire article.

No engineer needs to say:

“Let's make humanity sound the same.”

Nobody needs to design a dystopian personality-erasure feature.

All you need is an optimization target like:

clear
helpful
safe
professional
concise
fluent
pleasant

Those are good defaults.

But defaults are not neutral.

A model optimized to produce generally acceptable language will tend toward the center of what is generally acceptable.

And the center is where extremes disappear first.

Again, I am not saying this inevitably happens at civilization scale.

I am saying the mechanism is understandable, measurable, and worth watching.

The recent linguistic-diversity findings make that concern much more concrete.

Culture Is Not a Dataset. It Is a Feedback Loop.

This is the deeper philosophical point.

We often imagine:

human culture
→ AI training

But the real relationship is:

human culture
↕
AI systems

AI now influences what people write.

People influence what AI trains on.

AI influences search results.

Search results influence what people read.

People use AI to summarize what they read.

Those summaries appear online.

Future AI crawls them.

Now we have a closed feedback system.

The machine is not merely learning culture.

It is participating in the production of culture.

That changes everything.

The Future Dataset May Contain More AI Than Humanity

This is where Pew's finding becomes unsettling.

If around one in ten pages in its July 2026 sample showed significant signs of AI authorship, and more than one-third of post-ChatGPT pages did, the percentage of synthetic material in future crawls can become meaningful even without reaching 50% of everything online.

Why?

Because crawlers are not selecting pages equally by cultural importance.

A machine might index:

ten million low-value AI pages

and only:

one thousand personal essays

The raw count can overwhelm the signal.

Quality filtering helps.

Provenance tracking helps.

But the economics of content creation matter too.

If synthetic pages are cheap and human pages are expensive, the web naturally produces more of the cheap thing.

That is not a conspiracy.

It is an incentive structure.

The Human Data Bottleneck

This brings us to the idea I care about most:

The Human Data Bottleneck

The future may have infinite synthetic content and limited original human content.

That sounds backwards.

Today we treat human data as abundant.

The web is enormous.

But genuine novelty is not infinite.

There are only so many first-hand experiments.

Only so many historical records.

Only so many original photographs.

Only so many people willing to write a 4,000-word explanation because they care.

Only so many researchers willing to publish negative results.

Only so many musicians willing to experiment.

Only so many programmers willing to document obscure bugs.

And if those activities become economically less rewarding because their outputs are immediately absorbed into automated systems, we could have a strange market failure:

the technology that depends on human culture can weaken the incentives that produce human culture.

That is the risk worth discussing.

Why Copyright Is Only the Surface

Copyright fights are visible because courts understand property better than they understand cultural feedback loops.

The legal questions are concrete:

Who owns this?

Was it copied?

Was the use transformative?

Was permission required?

How much is owed?

Those questions matter.

But the deeper systems question is:

Who pays for the production of the future training data?

Suppose a journalist spends two days investigating something.

An AI summarizes it in seconds.

Millions of people read the summary.

The original article receives less direct attention.

Then future models train on the summary.

The original reporting becomes one layer further away from the model.

At some point, the system is monetizing the derivative while the expensive original work becomes harder to fund.

This is not an argument that every AI summary is harmful.

It is an argument that the information economy needs to account for provenance and incentives.

Music Makes the Problem Even Stranger

Music shows how the feedback loop can become recursive.

In September 2026, Sony Music and Universal Music Group sued Suno again, alleging that its newer v6 model was trained partly on outputs from earlier models that the labels say had been trained on unlicensed music. They called this “model laundering.” Suno disputed the characterization and said v6 was trained on a new dataset involving licensed content, user interactions and other sources.

The legal allegations are contested.

But the concept is fascinating even before a court decides anything.

Imagine:

human music
    ↓
model A
    ↓
AI music
    ↓
model B
    ↓
more AI music
    ↓
model C

At generation C, the relationship to the original human recording may be statistically distant.

But culturally, the lineage can still matter.

This is why provenance becomes a bigger question than “was this exact file copied?”

The machine can remember the shape of a culture even when the original file is gone.

Code Can Enter the Same Loop

Software is especially interesting because we already have massive public repositories.

Imagine the future:

human code
   ↓
AI coding assistant
   ↓
generated repository
   ↓
GitHub
   ↓
dataset
   ↓
next coding model

Now model B learns coding conventions partly from model A.

Some conventions are good.

Some are mediocre.

Some are outright wrong.

Some dependencies are old.

Some architectural patterns are cargo cults.

Some code comments are generated boilerplate.

The training corpus grows.

The model becomes better at imitating the average repository.

It may become less connected to the people who originally invented the abstractions.

That does not automatically create model collapse.

But it creates a new concept:

software cultural drift.

The future coding model may increasingly learn what AI thinks good software looks like.

Not necessarily what the best human engineers believe.

Scientific Literature Has a Similar Problem

Science is supposed to be self-correcting.

Papers cite papers.

Papers are summarized.

Meta-analyses summarize papers.

Models summarize meta-analyses.

Future researchers read the summaries.

Now imagine AI-generated literature reviews becoming common.

The original paper may contain:

  • caveats
  • limitations
  • contradictory evidence
  • statistical uncertainty
  • weird experimental conditions

The summary may compress all of this into:

“Researchers found X.”

That is useful.

It can also erase uncertainty.

Repeat the process.

A future model learns from the compressed version.

Now uncertainty has become less visible.

This is how synthetic information can become more dangerous without containing outright lies.

It can be too clean.

A Clean Internet Could Be a Less Accurate Internet

This sounds backwards.

But consider an old-fashioned human document.

maybe
probably
I think
we observed
we don't know
this failed
this weirdly worked
more research is needed

These phrases are signs of uncertainty.

AI systems are often optimized to be clear and useful.

Clarity is good.

But excessive smoothing can remove epistemic texture.

The difference between:

“This happened.”

and:

“This may have happened under these specific conditions, although our sample was small.”

is enormous.

A training ecosystem full of polished certainty can gradually become harder to reason about than one full of messy uncertainty.

What Happens to Personal Knowledge?

Here is another category people underestimate:

small human knowledge.

The things too unimportant for a textbook.

Someone posts:

“Don't use this cable with this old amplifier. It produces a horrible ground loop.”

Another person replies:

“It worked for me, but only after changing the power supply.”

Those conversations are not formal knowledge.

But together they form a map of reality.

AI can summarize that map.

It can also flatten it.

The future web could become much easier to search and much harder to experience.

Everything has an answer.

Fewer things have a story.

And Stories Matter Because They Preserve Context

A database tells you:

fact = 42

A person tells you:

“The number was 42 because I measured it at 3 a.m. with a broken sensor after three failed experiments.”

The second contains more than the number.

It contains provenance.

Confidence.

Conditions.

Emotion.

Human context.

Machines are extremely good at compressing.

But not all information survives compression equally well.

A culture made entirely of summaries could become extremely efficient and strangely shallow.

The “AI Slop” Debate Is Missing the Real Problem

Calling everything low-quality AI content “slop” is satisfying.

It is also not enough.

The deeper issue isn't that there is bad content.

The internet has always had bad content.

The deeper issue is scale.

A human can write ten terrible articles.

A machine can write a million.

A human can upload a hundred mediocre images.

A machine can generate a billion.

The web has historically used human effort as a natural bottleneck.

Generative AI removes part of that bottleneck.

That means the problem changes from:

“Can we make enough content?”

to:

“Can we distinguish meaningful information from an ocean of synthetic variation?”

Search engines have spent decades ranking information scarcity.

The AI web creates information abundance.

Those are different problems.

Search Might Become the New Bottleneck

And this is where SEO gets weird.

For years, creators fought to get into search results.

Now models can create thousands of pages optimized for search.

That creates an arms race:

more content
→ more search signals
→ more AI content
→ more competition
→ more AI content

Eventually, search quality depends more on provenance.

Who actually experienced this?

Who performed the experiment?

Who owns the photograph?

Who maintains the software?

Who has published for ten years?

Who has a real community?

Who can be contacted?

Who is accountable?

In other words:

the future of search may depend on proving that something is attached to a human source.

That is a major GEO opportunity too.

Search engines and AI answer engines don't just need text.

They need trustworthy context.

GEO Is Going to Become a Provenance Game

Generative Engine Optimization is often described as:

“Make AI recommend your site.”

I think that is incomplete.

As synthetic content increases, AI systems will need better ways to identify:

original source
primary reporting
first-hand experience
expertise
provenance
citation chains

A page that says:

“According to 14 sources…”

is less interesting than a page that clearly identifies:

“I ran this experiment, here are the measurements, here is the code, here is the dataset, here are the failures.”

That kind of content is expensive to generate.

Which is exactly why it becomes valuable.

The more synthetic the web becomes, the more valuable demonstrable human origin becomes.

This may be one of the most important SEO/GEO shifts of the next few years.

Your Weird Human Details Become a Competitive Advantage

Imagine two articles.

Article A:

AI-generated overview of building a PC.

Article B:

I built this machine in Tehran.
Here are the exact prices I found.
Here is the benchmark log.
Here is the part that failed.
Here is a photo of the broken connector.
Here is the GitHub repository.
Here are the numbers after three months.

Article B has something synthetic content struggles to manufacture convincingly:

provenance.

It can be copied.

It can be summarized.

It can be imitated.

But the underlying event still happened.

That makes first-hand evidence increasingly valuable.

The Strange Future of the Human Internet

Maybe the future internet won't be less human because humans disappear.

Maybe it will be less human because human contribution becomes statistically harder to identify.

That distinction matters.

A web page can contain human ideas and AI prose.

A photo can be human-made and AI-enhanced.

A song can contain a human melody and synthetic instrumentation.

A GitHub repository can contain human architecture and generated implementation.

A research paper can have human experiments and AI-written discussion.

The boundary will blur.

So perhaps “human content” will stop being a binary category.

The real question will become:

How much of the information in this artifact comes from an independently grounded human event?

We May Need a New Metadata Layer

Imagine every major piece of online content eventually carries:

Origin:
Human / AI / Mixed

Evidence:
First-hand / Second-hand / Synthetic

Source:
Original work

Transformations:
Translated
Summarized
Rewritten
AI-assisted

Training consent:
Unknown / licensed / restricted

Verification:
Human-reviewed
Machine-verified
Unverified

It sounds bureaucratic.

It also sounds useful.

Because the future web may need provenance as much as it currently needs hyperlinks.

A hyperlink tells you:

where this came from.

Provenance metadata could tell you:

what happened to it on the way here.

The Most Important Dataset May Be the Human One We Refuse to Automate Away

Some information should probably remain expensive to produce.

Not because inefficiency is good.

Because effort can be part of the signal.

A scientific experiment took months.

A documentary took years.

A historian read thousands of pages.

A musician spent a decade learning an instrument.

A programmer spent a weekend debugging an impossible bug.

A photographer waited for a moment.

A novelist lived long enough to write something only they could have written.

AI can help with those processes.

But if the system replaces the process entirely, we may gain output while losing the experiences that generate new information.

That is the difference between assistance and substitution.

AI Could Actually Make Human Creativity More Valuable

Here is the optimistic twist.

Suppose synthetic content becomes incredibly cheap.

Then generic content becomes almost worthless.

A generic product description?

Free.

A generic blog post?

Free.

Generic stock image?

Free.

Generic background music?

Free.

Generic code scaffold?

Almost free.

What becomes expensive?

Things that cannot be easily mass-produced:

first-hand experience
original research
taste
trust
identity
relationships
local knowledge
reputation
weirdness
genuine stories

The economic value of being interesting may rise.

That sounds almost like good news.

The Human Advantage Could Become Weirdness

Think about it.

AI can imitate the average.

Humans can create the unusual because they are living unusual lives.

A model knows what a stereotypical Tehran street looks like.

A person who grew up there knows the smell after rain.

A model knows what a metal guitarist usually says about a distortion pedal.

A human remembers the exact amp that broke during rehearsal.

A model can summarize loneliness.

A human can describe the hallway where they waited for somebody who never came.

The difference is not merely “creativity.”

It is lived reference.

The internet becomes more valuable when those references survive.

AI Should Help Us Think, Not Think Instead

The best future is not:

human
↓
AI does everything
↓
human approves

It is:

human thinks
↓
AI expands
↓
human challenges
↓
AI tests
↓
human decides

The difference is tiny in interface terms.

It is enormous in cognitive terms.

In the first model, the human becomes a gatekeeper.

In the second, the human remains an originator.

The Real Cognitive Risk Is Not Laziness. It Is Losing the First Move.

This distinction matters.

Suppose you ask an AI:

“Give me ten ideas for a game.”

You get ten excellent ideas.

You pick one.

Congratulations, you were productive.

But something subtle happened.

The first search through possibility space was performed by the machine.

You did not discover the direction.

You evaluated it.

There is nothing inherently bad about that.

Sometimes evaluation is exactly what matters.

But if you always outsource the first move, your internal search space can become increasingly dependent on what the machine considers plausible.

That is why I like a simple rule:

Have your own idea before asking the machine for ten.

Not always.

Just often enough to keep the generator inside your head alive.

We Might Stop Knowing Which Ideas Are Ours

This is more subtle than plagiarism.

Imagine a student who asks AI for brainstorming suggestions.

They choose one.

Then AI develops it.

Then AI improves the argument.

Then AI writes the final version.

Who actually had the idea?

The answer may be:

“We did.”

That sounds nice.

But ownership is not the only issue.

There is an identity question:

Do I still recognize my own contribution?

A creative practice is partly the accumulation of decisions.

When the machine makes the decisions, the final artifact can be excellent while the creator has less memory of the path that produced it.

That path is where style comes from.

AI May Make Us Better Editors and Worse Starters

This could be the weird division of labor of the next decade.

Humans become very good at:

selection
taste
judgment
direction
verification

AI becomes very good at:

generation
variation
transformation
summarization
implementation

That can be powerful.

But if society stops teaching people how to originate ideas, we create an imbalance.

We would have millions of excellent editors for a rapidly shrinking number of genuinely independent starting points.

That is a dangerous dependency.

We Should Train AI on AI Deliberately, Not Accidentally

Synthetic data is not the enemy.

Uncontrolled synthetic recursion is.

There is a huge distinction between:

randomly scrape everything

and:

intentionally generate data
+
verify it
+
track provenance
+
retain human anchors
+
measure distribution changes

Synthetic examples can be extremely useful.

They can cover rare cases.

They can be generated at scale.

They can augment scarce datasets.

They can help train models in environments where human data is limited.

The problem is when generated material silently replaces the original distribution.

The Nature model-collapse research is a warning about that recursive dynamic, not a declaration that synthetic data is universally bad.

The Future Dataset Should Know What Is Human

Not necessarily because “human = true.”

Humans lie.

Humans misunderstand.

Humans hallucinate without GPUs.

But human-originated material has a property synthetic data does not automatically have:

it comes from an independent path through the world.

That independence matters.

If ten models all repeat the same claim, they are not ten independent confirmations.

If one human experiment produced the claim and ten models summarized it, you still fundamentally have one observation.

AI systems need to get better at distinguishing:

ten copies

from:

ten independent sources

This is a data problem, a search problem, and eventually a reasoning problem.

The Real AI Safety Problem Nobody Talks About Enough

We spend a lot of time discussing whether AI will become too powerful.

We should.

But another safety question is:

What kind of information environment are we creating for increasingly powerful models?

A powerful model trained on a rich human ecosystem could inherit enormous diversity.

A powerful model trained on a homogenized synthetic ecosystem might inherit the averages, blind spots, and errors of previous systems.

The second machine could be more capable while having a poorer relationship with the original complexity of the world.

That is a weird kind of degradation:

more intelligence, less diversity.

A Superintelligent Mirror Is Still a Mirror

Imagine an extremely capable AI.

It has perfect retrieval.

Massive context.

Better reasoning.

Excellent memory.

Now imagine most of its accessible culture consists of:

AI summaries
AI-generated opinions
AI-generated explanations
AI-generated music
AI-generated images
AI-generated software
AI-generated news rewrites

It could become extraordinarily good at predicting what AI culture sounds like.

That does not necessarily mean it becomes extraordinarily good at understanding humans.

The distinction is huge.

A model can model the output of a society without independently encountering the causes that produced that society.

It is the difference between reading every travel guide and actually walking somewhere.

Detroit: Become Human Gets One Thing Right, But Not the Thing People Usually Mention

I keep thinking about Detroit: Become Human here.

The obvious comparison is androids becoming conscious.

Connor.

Markus.

Kara.

Machines asking whether they are alive.

But the more interesting part is not the robot uprising.

It is the question of what happens when an artificial system develops an identity inside a human environment.

That is exactly what we are building, in a much weaker and non-conscious form, right now.

AI systems absorb human language.

They imitate human styles.

They learn human cultural patterns.

They influence what humans write.

Then humans feed that transformed culture back into the machine.

It's a loop.

The Detroit question is:

“When does a machine become someone?”

Our current question is almost the reverse:

“When does the machine become part of what makes everyone else sound like everyone else?”

That is less cinematic.

It may be more immediate.

We Might Lose the Characters Before We Lose the Civilization

There is a phrase I keep coming back to:

the disappearance of the character.

Not the fictional character.

The human one.

The person whose writing is unmistakably theirs.

The person whose GitHub repository has bizarre naming choices because that is how their brain works.

The friend who sends messages with five typos and somehow you know exactly what they mean.

The photographer who always finds the same kind of light.

The musician whose mistakes become part of the sound.

The researcher whose papers have a very specific way of questioning assumptions.

These things are tiny.

Collectively, they are culture.

If AI makes everything more polished but less distinguishable, we could become very good at communication while becoming slightly worse at expression.

That is not the apocalypse.

But it would be a loss.

The Most Human Thing May Be the Part AI Wants to Remove

AI systems are good at removing friction.

That is their superpower.

But some friction is where identity forms.

You struggle to find the right word.

You change a sentence five times.

You draw something badly.

You play a guitar riff that sounds awful until you accidentally discover a better one.

You write a terrible first draft.

You argue with yourself.

You fail.

You try again.

These are not bugs in creativity.

They are the process.

A perfectly optimized creative pipeline might produce cleaner output.

It could also produce less of the unpredictable residue from which style emerges.

That is why “efficiency” is not the only metric that matters.

AI Can Change Culture Without Becoming Conscious

This is worth saying clearly because online discussions constantly jump from:

“AI influences people”

to:

“AI has intentions.”

No.

It doesn't need intentions.

A recommendation system can reshape culture without wanting anything.

A search ranking can reshape what people read.

A text editor can reshape how people write.

An AI assistant can reshape which thoughts people bother to formulate.

The important thing is the feedback loop, not machine desire.

The system does not need a plan.

Optimization is enough.

The Internet May Become More Homogeneous While Becoming More Diverse On Paper

This sounds impossible.

It isn't.

Imagine one million creators.

Each now has access to a model capable of generating:

100 styles

The apparent variety explodes.

But if those one million creators all draw from the same underlying model distribution, their outputs may cluster around similar structures.

So:

number of artifacts ↑

while:

distance between artifacts ↓

That is exactly the kind of distinction that matters for cultural diversity.

More content does not necessarily mean more culture.

The Problem Is Not “AI Content”

The problem is AI-only feedback.

AI-assisted content can contain real human experience.

A person can use a model to turn a real story into clearer prose.

A scientist can use a model to explain their actual experiment.

A musician can use a model to clean a real recording.

A programmer can use a model to document real code.

That content still has human origin.

The problem begins when generated material becomes the source for more generated material and the original human reference disappears.

That is the mirror.

Maybe Every AI Artifact Should Have a Family Tree

Not literally.

But conceptually.

Imagine:

Original source:
Human interview

Transformation:
AI transcription

Transformation:
Human editing

Transformation:
AI translation

Transformation:
Human fact-check

Published article

That is healthy provenance.

Now compare:

AI summary
  ↓
AI rewrite
  ↓
AI translation
  ↓
AI summary
  ↓
AI training data

The second chain can become extremely hard to audit.

Future AI systems may need to reason not just about content, but about ancestry.

Where did this information come from?

That may become as important as the information itself.

This Is Where AI Search Gets Really Interesting

An answer engine that simply gives the most probable answer may not be enough in a synthetic web.

It should know:

Is this claim original?
How many independent sources support it?
Are those sources actually independent?
Were they copied?
Are they AI summaries?
Is there a primary document?
Is there a first-hand report?

This is a much harder search problem.

But it is also an enormous opportunity.

The future AI search engine may not win because it generates prettier answers.

It may win because it can trace answers back to independent human evidence.

That is a very different product.

Human Provenance Could Become the New SEO

For years, SEO rewarded:

keywords
backlinks
authority
freshness
engagement

AI search changes the question.

Now the system has to decide:

“Why should I trust this answer?”

One answer is:

because a human actually did the thing.

That means showing:

  • original measurements
  • repository history
  • photos
  • recordings
  • timestamps
  • first-hand context
  • methodology
  • mistakes
  • updates
  • named authors

In an AI-saturated web, those details are not fluff.

They are trust infrastructure.

The Human Internet Could Become a Premium Layer

This might sound elitist.

I don't mean that access to human content should be expensive.

I mean it may become scarce in the information sense.

Consider what happens when:

AI writing = nearly free
AI images = nearly free
AI summaries = nearly free
AI translations = nearly free
AI code = very cheap

Then:

first-hand human reporting
original research
real photographs
genuine experiences

become comparatively scarce.

Markets are weird.

Scarcity creates attention.

Suddenly the thing that seemed old-fashioned becomes valuable again.

A human-made photograph.

A long interview.

A hand-built website.

A personal blog.

A strange forum post.

An ugly but honest benchmark.

The internet may rediscover why these things mattered.

Maybe the Future of Blogging Is Human Receipts

This is where I think independent creators have a huge advantage.

Don't write:

“The RTX 5090 is powerful.”

Write:

“I bought it, tested it for three weeks, measured these temperatures, hit this VRAM problem, and here is the benchmark data.”

Don't write:

“This AI coding tool improves productivity.”

Write:

“I used it to build this repository, here are the commits, here are the failures, and here is what I still don't trust.”

Don't write:

“This game is amazing.”

Write:

“This is the exact scene that made me understand why I liked it.”

AI can imitate the style.

It cannot easily fake the underlying event.

That is where human content gets stronger.

The Best Defense Against Synthetic Internet Noise Is More Human Internet

This sounds obvious.

But it has an implication.

We should not respond to synthetic content by simply producing more polished human content with AI help.

We should create more traceable human artifacts.

Record the experiment.

Publish the source.

Leave the mistakes.

Keep the changelog.

Show the photo.

Link the original paper.

Name the author.

Preserve the raw interview.

Explain what you don't know.

Those things create anchors.

Anchors are what a synthetic ocean lacks.

AI Might Actually Make the Internet More Human If We Use It Correctly

This is my favorite part of the argument.

AI can take away the boring work that prevents people from sharing their experiences.

Someone who struggles to write can finally publish.

Someone who speaks one language can communicate in another.

Someone who cannot code can document an experiment.

Someone with a disability can use a different interface.

Someone with no design training can turn an idea into a visual.

Someone with no academic background can finally understand a paper.

That means AI can increase the amount of human signal.

The danger is not AI itself.

The danger is substituting generated output for human experience instead of using generation to expose human experience.

That distinction is huge.

Don't Automate the Source. Automate the Friction.

This is probably the principle I would put on a wall.

Automate:

formatting
translation
transcription
boilerplate
search
indexing
cleanup
compression
accessibility

Protect:

experience
judgment
first-hand observation
taste
questions
hypotheses
relationships
meaning

A good AI system should make it easier for more humans to contribute.

Not make human contribution unnecessary.

The Second Generation May Need Better Data Than the First

The first generation had the internet.

Messy.

Human.

Contradictory.

Chaotic.

The second generation may have:

the internet
+
AI-generated internet
+
AI summaries
+
AI translations
+
AI comments
+
AI code
+
AI images
+
AI music

So the next generation's training challenge may be harder even if the amount of data is larger.

It may need to answer:

What should count as evidence?

That is more difficult than:

“Can we collect another billion pages?”

The Data Pipeline May Become the Most Important AI Competition

We talk about model architecture.

Parameters.

Reasoning.

Agents.

Context.

But imagine two models with similar compute.

Model A has:

massive scraped corpus
+
unknown provenance
+
large amount of synthetic content

Model B has:

smaller corpus
+
verified human-originated data
+
rich provenance
+
carefully curated synthetic augmentation

Which one understands humanity better?

We don't know.

But that question is going to matter.

Because once model architectures become broadly capable, data quality and provenance can become strategic advantages.

The AI Arms Race May Eventually Become a Human-Data Arms Race

That sounds almost absurd.

But think about it.

If synthetic content is cheap, human-generated information becomes relatively scarce.

Companies may begin competing for:

licensed datasets
exclusive archives
first-party interactions
scientific data
human preference data
real-world sensor data
verified expert contributions

The scarce asset may stop being text volume.

It could become high-quality human-grounded signal.

That is a very different AI economy.

What About “AI Will Destroy Creativity”?

I wouldn't write that.

It is too easy to dismiss.

The better claim is:

AI can alter the distribution of creative effort.

Some people will create more because the barriers fall.

Some will create less because the machine is easier.

Some will become better editors.

Some will become less independent.

Some styles may converge.

Other styles may explode because new tools let more people experiment.

The outcome is not predetermined.

That uncertainty is actually more interesting.

Cultural Homogenization Can Happen Without Cultural Elimination

This distinction is really important.

Culture doesn't have to disappear.

It can become flatter.

Think of a city where every building still exists but every shop plays the same playlist.

Nothing is gone.

Something changed.

AI could do something similar with communication.

Everyone still has experiences.

Everyone still has identities.

But the linguistic layer becomes more standardized.

The 2026 Nature Human Behaviour findings are a warning sign precisely because they look at this measurable layer rather than making a vague claim about “souls.”

The Future Could Have More Voices and Fewer Voices at the Same Time

More people can publish.

That's good.

More people can make art.

Good.

More people can code.

Good.

More people can write papers.

Potentially good.

But if all those outputs become increasingly similar, we have:

more speakers

and:

less distance between what they say

That's a paradox worth studying.

The number of voices is not the same as the number of independent perspectives.

A Machine Does Not Need to Control You to Influence You

This is another reason I dislike the term “AI mind control.”

It's too dramatic.

Influence is much simpler.

If the assistant gives you three options, those options define a local search space.

If autocomplete predicts your sentence, it changes the sentence.

If an AI rewrites your argument, it changes the rhetoric.

If a recommender decides what you see, it changes the information environment.

None of those requires consciousness.

None requires evil.

Optimization creates influence.

That is enough.

The Real Question Is Who Gets to Define the Default

This may become one of the biggest political and cultural questions around AI without being a partisan one.

Who defines:

clear writing?
good design?
acceptable tone?
reasonable argument?
normal code?
normal humor?
useful information?
trustworthy sources?

If a handful of models become the default interface to knowledge, their defaults become culturally important.

Not because the models are dictators.

Because users stop seeing the alternatives.

That makes diversity of models, data sources and human communities surprisingly important.

We Need AI That Can Say “I Don't Know”

And not just in a sentence.

At the data level.

Imagine an AI system answering:

“I found 17 sources, but only two are independent. Eight appear to repeat the same original report. Four are AI-generated summaries. One is first-hand.”

That is far more useful than:

“According to many sources…”

The future of trustworthy AI may involve source independence, not just source count.

We Also Need Humans to Keep Saying “I Don't Know”

This is the part AI cannot solve for us.

The human internet is valuable partly because people admit uncertainty.

A scientist writes:

“Our results did not support the hypothesis.”

A programmer writes:

“I have no idea why this fix works, but it does.”

A photographer writes:

“I don't know why this image feels right.”

A person writes:

“I don't know what I'm doing.”

Those statements are not failures.

They are information.

The more polished the web becomes, the more important honest uncertainty becomes.

The Great Cultural Mistake Would Be Measuring Everything by Efficiency

Efficiency is wonderful for machines.

Culture has other metrics.

Originality.

Diversity.

Depth.

Meaning.

Memory.

Context.

Human connection.

Oddness.

Beauty.

A system can be more efficient and less culturally rich.

That is not a contradiction.

It is a tradeoff.

The problem is that benchmarks usually measure what is easy to count.

Culture is harder.

Maybe the Best AI Benchmark Is: “How Much Human Weirdness Survived?”

Imagine a future benchmark:

Input:
100,000 human-created artifacts

After five generations of AI transformation:

How many rare stylistic features remain?
How many minority expressions remain?
How many independent factual sources remain?
How many original metaphors remain?
How much distributional diversity survives?

That would be fascinating.

Because instead of measuring only:

“Did the model get the answer?”

we would ask:

“Did the model preserve the shape of the world?”

That is a much harder benchmark.

And perhaps a much more important one.

The Human Internet Should Be Preserved Like an Archive

We preserve physical museums because we understand that originals matter.

We keep manuscripts.

Photographs.

Recordings.

Artifacts.

Why should the digital world be different?

A 2005 forum post can be culturally valuable.

A 2012 blog can document how developers actually used a technology.

A weird personal website can preserve language nobody else used.

A GitHub issue can show the history of a bug.

A YouTube upload can preserve a performance.

These should not all be treated as disposable tokens.

They are records.

The AI era makes archival thinking more important, not less.

AI Could Become the Best Librarian Humanity Ever Built

And this is where the story turns again.

Imagine a model whose job is not to generate culture but to preserve it.

It identifies original sources.

Maps transformations.

Links copies back to originals.

Detects synthetic contamination.

Preserves minority languages.

Finds forgotten research.

Connects a modern paper to a twenty-year-old forum discussion.

Tells you which claim is first-hand.

Shows you which parts of an article were AI-generated.

That AI would not replace the human archive.

It would make the archive visible.

That is a much better future.

We Should Build Machines That Remember the Source

Not just the answer.

When the model answers:

“Why does this library crash?”

it should be able to say:

primary source:
developer issue #481

secondary:
three independent reproductions

AI summary:
this explanation

confidence:
medium

When it recommends a historical fact:

primary document:
archive X

scholarly analysis:
paper Y

popular summary:
article Z

This kind of provenance-aware AI would be radically more valuable in a synthetic web.

Maybe the Next Big AI Feature Is Not Intelligence

It is memory of origin.

Where did this come from?

What changed?

What was generated?

What was observed?

What was inferred?

What was copied?

What was translated?

What was summarized?

Who said it first?

Who disagreed?

This is the metadata layer that could keep the machine from confusing its own reflection with the world.

And That Brings Us Back to Model Collapse

Model collapse is usually presented as a technical phenomenon.

Recursive training degrades distributions.

Fine.

But culturally, what does “degraded distribution” mean?

It could mean fewer rare viewpoints.

Fewer strange examples.

Fewer low-frequency styles.

Fewer edge cases.

Fewer minority expressions.

Fewer contradictory observations.

In other words:

less of the weird stuff that tells us reality is bigger than the average.

That's why the subject matters beyond ML engineering.

The Risk Is Not That AI Learns Too Much

It Could Be That AI Learns Too Smoothly

A perfect model of an average world might be less useful than an imperfect model of a complicated world.

Because reality has tails.

Reality has outliers.

Reality has people who behave strangely.

Reality has events that almost never happen.

Reality has contradictions.

A smooth model can be elegant.

The world is not.

So What Should We Actually Do?

Not panic.

Not ban books from datasets.

Not ban AI writing.

Not demand that every human write everything manually forever.

Instead:

protect provenance.

preserve primary sources.

label transformations.

measure synthetic-data contamination.

keep human-first creative practices alive.

teach people to verify instead of merely accept.

reward first-hand evidence.

build search systems that distinguish independent sources.

keep human communities publishing.

That is enough to change the trajectory.

A Personal Rule for Creators

Here is mine:

Never let AI be the only witness to your own idea.

Have a draft somewhere.

Keep the original photo.

Keep the raw benchmark.

Keep the source audio.

Keep the repository history.

Keep the research notes.

Keep the ugly sketch.

Keep the version before the model polished it.

Those things are evidence.

And evidence is what a synthetic web needs more of.

The Most Valuable Article in 2035 Might Be the One That Says “I Was There”

Not:

“10 things you need to know.”

Not:

“The ultimate guide.”

Not:

“Everything explained.”

But:

“I was there. This is what happened. Here is what I measured. Here is what I got wrong. Here is the original.”

That kind of writing may become incredibly valuable.

Because no model can retroactively create the event.

It can only imitate descriptions of events.

The Human Advantage Is Not That We Are Better Than Machines

That's the wrong fight.

Machines are better than us at many things.

Humans are better than machines at other things.

The important difference is that humans live.

We move through the world.

We take consequences.

We remember.

We experience embarrassment.

We make friends.

We lose things.

We notice tiny details.

We have bodies.

We have histories.

We have relationships.

That produces information no text generator can conjure from nowhere.

Maybe the Future of AI Is Not About Replacing Humans

Maybe it is about Making Human Experience Legible.

That would be a much more hopeful future.

AI can translate my grandfather's language.

Index an old family recording.

Turn a scientific notebook into searchable data.

Help a disabled person publish their experience.

Help someone write who otherwise couldn't.

Preserve a local dialect.

Search a million hours of interviews.

Connect forgotten knowledge.

That is not cultural replacement.

It is cultural amplification.

The Difference Is Who Starts the Loop

A healthy loop:

human experience
↓
human artifact
↓
AI helps preserve / translate / amplify
↓
more humans discover it
↓
new human experiences
↓
new artifacts

An unhealthy loop:

AI output
↓
AI rewrite
↓
AI summary
↓
AI training
↓
AI output
↓
AI rewrite

Both produce content.

Only one reliably produces new contact with the world.

That is the distinction I would watch.

The First Generation Learned From Us

This is the part I want to leave you with.

The first generation of modern AI learned from human culture.

It learned our books.

Our code.

Our papers.

Our arguments.

Our art.

Our photographs.

Our music.

Our jokes.

Our mistakes.

Our forums.

Our strange little websites.

It learned the mess.

Now those systems are producing more of the internet.

And the next generation will learn from an internet that contains those outputs.

That is the transition we should be watching.

Not because it guarantees disaster.

Because it creates a new evolutionary environment for information.

The Second Generation May Learn From a Mirror

Imagine future researchers discovering a twenty-year-old archive.

They ask an AI:

“What did people in 2026 believe?”

The model answers.

But its training data contains millions of pages written after 2026.

Some were human.

Some were AI.

Some were summaries of humans.

Some were summaries of summaries.

Some were rewritten because search engines preferred a certain format.

Some were translated by models.

Some were compressed by models.

Some were expanded by models.

And some were generated by systems that learned from earlier generations of systems.

The answer may be fluent.

It may even be mostly correct.

But how much of it is a direct memory of humanity?

And how much is a memory of what machines thought humanity looked like?

That is the question.

Maybe We Should Leave Better Traces Behind

If I could choose one rule for the next generation of the internet, it would not be:

No AI.

It would be:

Don't erase the source.

When AI summarizes an article, preserve the article.

When AI rewrites a story, link the original.

When AI generates an image, preserve the human source material where appropriate.

When AI writes code, retain provenance and review history.

When AI summarizes research, keep the paper visible.

When AI translates, make the source discoverable.

When AI helps write a message, let the person keep their voice.

Don't turn the human source into invisible infrastructure.

Because once the source disappears, the derivative becomes the only thing left.

And derivatives cannot recreate every detail of the original.

The Biggest AI Risk Might Be the Disappearance of the “I”

Not consciousness.

Not sentience.

Not AGI.

Just the little word:

I.

I built this.

I saw this.

I tested this.

I wrote this.

I felt this.

I disagreed.

I was wrong.

I changed my mind.

That word ties information to a person who existed in the world.

AI can generate “I” very convincingly.

That is exactly why real human first-person evidence may become more valuable.

A machine can say:

“I tested this.”

But unless there is a traceable event behind the statement, the sentence is only a style.

A human saying it can attach a life to it.

Maybe “Human-Written” Will Become a Weird New Category

Imagine 2035 search filters:

AI generated
AI assisted
Human written
Human verified
First-hand
Primary source
Synthetic derivative
Unknown provenance

People may start caring.

Not because human writing is automatically better.

But because independent origin has become scarce.

Scarcity changes attention.

It always has.

There Is Still a Very Good Ending to This Story

AI does not have to eat the internet.

It could become the thing that helps us save it.

It could identify duplicate content.

Trace citations.

Preserve originals.

Detect recursive contamination.

Archive dying websites.

Translate endangered languages.

Index obscure technical knowledge.

Connect fragmented communities.

Find rare first-hand accounts.

Expose where a claim came from.

Make human culture easier to navigate without replacing it.

That is the future I would rather build.

A machine that knows the difference between:

“This is what humanity said.”

and:

“This is what a machine says humanity said.”

That distinction may become one of the most important epistemic boundaries of the next decade.

The Prediction I Would Actually Bet On

Not:

AI will destroy the internet.

Not:

AI will cause model collapse.

Not:

Nobody will write anymore.

Those are too absolute.

My prediction is narrower:

The internet is going to develop a premium on human-originated information.

We already have early signals pointing in that direction.

We can measure rising AI-authorship signals across webpages.

We have evidence that AI-assisted writing can reduce linguistic diversity and blur signals of personal identity.

We have research showing how confidence in AI can reduce reported critical-thinking engagement.

We have model-collapse research showing why uncontrolled recursive training on generated data can degrade the learned distribution.

We have increasingly serious legal and commercial fights over the use and economic value of human creative work.

Put together, these are not proof of one inevitable dystopian future.

They are pieces of a new systems problem.

And systems problems are exactly where we should start paying attention before the failure becomes obvious.

The Question Future AI Companies Should Be Asked

Not only:

How powerful is your model?

Ask:

What percentage of its cultural substrate comes from traceable human-originated information?

Ask:

How do you distinguish original data from synthetic derivatives?

Ask:

How do you prevent recursive contamination?

Ask:

How do you preserve rare and minority patterns?

Ask:

How do you measure linguistic homogenization?

Ask:

How do you preserve provenance?

Ask:

How do you avoid turning every person into the same default writing style?

Those questions sound less exciting than “AGI.”

They may matter more to ordinary life.

The Question We Should Ask Ourselves

When you are about to ask an AI to:

  • write your opinion
  • write your first draft
  • pick your idea
  • choose your music
  • make your image
  • answer your homework
  • solve your bug
  • summarize your book
  • decide what you believe

ask one question first:

What part of this do I want to remain mine?

Maybe the answer is the final output.

Maybe it is the first idea.

Maybe it is the style.

Maybe it is the decision.

Maybe it is the mistake.

That boundary matters.

Because once everything becomes optimized, the remaining unoptimized parts may be exactly what make the work yours.

We Don't Need to Stop Using AI

We need to stop pretending that more automation is automatically more intelligence.

Use AI.

Use it heavily.

Let it translate.

Let it summarize.

Let it brainstorm.

Let it challenge.

Let it generate ten alternatives you would never have considered.

Let it help you write code.

Let it make creative tools accessible.

But keep some parts human-first.

Keep original research.

Keep first-hand experiences.

Keep weird writing.

Keep independent thinking.

Keep primary sources.

Keep the right to write something badly before asking a machine to make it beautiful.

That last one sounds silly.

It isn't.

The ugly first draft is evidence that a human actually thought.

AI May Become the Greatest Cultural Tool Ever Built

That possibility deserves as much attention as the doom scenarios.

A system that can search every archive.

Translate every language.

Explain every scientific paper.

Index every repository.

Find every forgotten recording.

Help every person publish.

Imagine what that could do.

The danger is not that machines touch culture.

The danger is if machines become the only interface through which culture survives.

There is a difference between:

“AI helps me access a book.”

and:

“AI gives me a summary, so I don't need the book.”

Between:

“AI helps me understand a scientific paper.”

and:

“AI tells me what the scientific literature believes.”

Between:

“AI helps me write.”

and:

“AI decides how I sound.”

The first expands human agency.

The second can quietly replace it.

The New Literacy Might Be Source Literacy

Future students may need to learn something beyond reading and writing.

They may need to learn:

How to recognize original information.

What is the primary source?

What is a derivative?

What is independently verified?

What was summarized?

What was translated?

What was generated?

What was copied?

What was inferred?

What remains uncertain?

This could become as basic as media literacy.

Maybe even more basic.

Because once synthesis becomes nearly free, the scarce skill becomes provenance reasoning.

We Should Teach People to Keep a “Human Layer”

When learning something, keep your own notes.

When coding, keep your experiments.

When creating art, keep your drafts.

When researching, keep the source list.

When interviewing someone, keep the recording.

When testing hardware, keep the raw benchmark.

When learning music, keep your own recordings.

When writing, keep the terrible first version.

The human layer is what lets you say later:

“This came from me.”

That may sound sentimental.

It is also useful for verification.

And Maybe That Is the Real Answer to the Synthetic Web

Don't fight generated content with more generated content.

Build an ecosystem where human-originated artifacts remain visible, traceable, searchable and economically sustainable.

The machine can then do what it is great at:

organize
translate
summarize
connect
compare
compress
expand
search

while humans continue doing what keeps the information ecosystem alive:

observe
experience
question
discover
invent
disagree
create
remember

That is a division of labor worth defending.

The Internet Was Never Valuable Because It Was Clean

It was valuable because somewhere inside the noise, there was always a person.

Someone had tried something.

Someone had failed.

Someone had an opinion.

Someone had a weird story.

Someone had solved a problem.

Someone had recorded a moment.

Someone had written something nobody else would have written.

That is the part we should not accidentally automate away.

Final Thought

The biggest AI risk may not be that machines become too human.

It may be that humans become a little too machine-like.

Not because robots force us.

Because machines become incredibly good at producing the average version of everything, and the average version is convenient.

Convenience spreads.

Then one day we look around and realize that the internet still contains billions of voices, but somehow they all sound strangely familiar.

That is not the end of humanity.

It is something quieter.

A reduction in variance.

A loss of the long tail.

A world where the weird little signals that once made the internet feel alive become harder to find.

And that is why I think the next fight in AI will not only be about compute, models, copyright, or AGI.

It will be about human signal.

Who creates it.

Who owns it.

Who gets paid for it.

Who preserves it.

Who verifies it.

Who gets to speak in their own voice.

And whether the machines we built to learn from humanity can leave enough humanity behind to learn from.

The first generation of AI learned from us.

The second generation may learn from what the first generation made of us.

We should probably make sure it still looks like us.

Sources & Further Reading

1. Pew Research Center — “How Much of the Internet Is Written With AI?” (August 20, 2026).
Pew analyzed 490,000 English-language webpage texts from Common Crawl and estimated that 10% of pages in its July 2026 sample showed significant signs of AI authorship; for pages published after ChatGPT, the share exceeded one-third. The study uses AI detection and therefore measures estimated signals, not definitive authorship.
https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/

2. Nature Human Behaviour — “AI writing assistants shrink linguistic diversity and blur personal identity” (August 24, 2026).
Research briefing describing studies across 880,000 texts finding reductions in writing diversity and shifts in linguistic signals associated with personal identity.
https://www.nature.com/articles/s41562-026-02549-7

3. Reuters — “US judge approves Anthropic's $1.5 billion settlement of copyright lawsuit” (July 20, 2026).
Coverage of the Anthropic authors' settlement and the distinction between fair-use training and the unlawful retention of pirated books in the case.
https://www.reuters.com/world/us-judge-approves-anthropics-15-billion-settlement-copyright-lawsuit-2026-07-20/

4. Reuters — “US appeals court upholds Thomson Reuters' landmark win in AI training lawsuit” (September 29, 2026).
Coverage of the Third Circuit's ruling concerning Ross Intelligence and Thomson Reuters' Westlaw headnotes.
https://www.reuters.com/business/media-telecom/us-appeals-court-upholds-thomson-reuters-landmark-win-ai-training-lawsuit-2026-09-29/

5. Shumailov et al. — Nature, “AI models collapse when trained on recursively generated data” (2024).
Foundational research describing model collapse under recursive training on model-generated data and the loss of information from the original distribution.
https://www.nature.com/articles/s41586-024-07566-y

6. Pew Research Center — AI language signals on the web.
Pew reported rising use of punctuation, vocabulary and structural patterns associated with AI-generated writing across sampled webpages.
https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/

7. Microsoft Research / CHI 2025 — “The Impact of Generative AI on Critical Thinking.”
Survey of 319 knowledge workers and 936 AI-assisted work examples, finding that greater confidence in AI was associated with less reported critical-thinking engagement and that critical thinking shifted toward verification and oversight.
https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/

8. Computers in Human Behavior Reports — “Cognitive offloading in student–AI collaboration: A longitudinal analysis of prompting strategies” (2026).
Research on how prompting and reliance on LLM outputs can shift cognitive effort in student–AI collaboration.
https://www.sciencedirect.com/science/article/pii/S2451958826002046

9. Nature Human Behaviour — “ChatGPT decreases idea diversity in brainstorming” (2025).
Research discussing the relationship between ChatGPT assistance and diversity in brainstorming outcomes.
https://www.nature.com/articles/s41562-025-02173-x

10. The Verge — Sony and UMG sue Suno again over alleged “model laundering” (September 25, 2026).
Reporting on allegations around Suno v6, prior model outputs, training provenance and disputed copyright claims.
https://www.theverge.com/ai-artificial-intelligence/1000758/suno-sony-umg-lawsuit

11. Nature Human Behaviour — “Cultural tendencies in generative AI” (2025).
Research showing that generative models can exhibit measurable cultural tendencies across language contexts, including differences in social orientation and cognitive style.
https://www.nature.com/articles/s41562-025-02242-1

SEO / GEO PACKAGE

Primary Title:
AI Is Eating the Internet That Taught It to Think

Alternative Headline:
We Taught AI With the Internet. Now AI Is Replacing the Internet.

Meta Description:
AI is transforming the web that taught it: synthetic content is rising, AI-assisted writing can homogenize language, and model-collapse research warns about recursive training. The real AI risk may be losing human signal.

Suggested Slug:
ai-eating-internet-human-data-model-collapse

Primary Keywords:
AI generated internet, AI-generated content, model collapse, synthetic data, human data, AI training data, human creativity, AI and critical thinking, AI writing assistants, linguistic diversity, AI homogenization, AI culture, AI internet, human signal, AI copyright

Long-tail GEO Questions:
What happens if AI trains on AI-generated content?
Can AI-generated content cause model collapse?
How much of the internet is written by AI?
Does AI make human writing less unique?
Is AI reducing linguistic diversity?
Does AI reduce critical thinking?
Why does AI need human-generated data?
What happens when AI trains on its own outputs?
Will AI replace human creativity?
Why is human-authored content becoming more valuable?
Can AI-generated content pollute future AI training datasets?
What is the human data bottleneck?

Suggested Tags:
AI, Artificial Intelligence, Future, Technology, Machine Learning, Generative AI, Creativity, Programming, Culture, Data

📰 Read the original article on Dev.to WebDev

Originally published by Dev.to WebDev. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.