AI authorship study highlights webpage content on a laptop screen

One in Three New Webpages Since ChatGPT’s Launch Shows Signs of AI Writing, Pew Says

Pew says AI authorship appears on 35% of webpages published after ChatGPT’s launch, highlighting the rapid rise of AI-written content online.

In short

Pew Research says more than one-third of webpages published after ChatGPT’s launch show signs of AI authorship. The study suggests AI-assisted writing is now widespread across the modern web, especially on commercial sites.

  • Pew found signs of AI authorship on 35% of webpages published after ChatGPT launched.
  • Commercial .com sites showed far higher AI-writing signals than .edu, .gov and .org domains.
  • The study used nearly 500,000 English-language webpages from Common Crawl and Open Pangram detection.
  • Pew warns the estimate is directionally useful but not perfect because AI detectors can misclassify pages.

More than one-third of webpages published after ChatGPT’s debut appear to show signs of AI authorship, according to a new Pew Research study released Thursday. The finding suggests that generative tools have quickly reshaped what gets published online, even as the web also fills with automated systems reading, indexing and comparing that content at scale.

The report adds fresh evidence to a broader shift already visible across the internet: newer webpages are increasingly likely to be written entirely by AI or heavily edited with AI assistance, while detection tools and bot traffic analyses continue to point toward a web that is becoming more machine-mediated on both the publishing and browsing sides.

What Pew found

Pew’s analysis centered on English-language webpages collected from the Common Crawl web archive, which is one of the largest public snapshots of the web. Using Open Pangram’s AI-detection technology, researchers examined nearly half a million pages published over roughly the past five years, including years before ChatGPT launched in November 2022.

In a July 2026 random sample of 10,000 webpages, Pew said about 10% displayed significant indicators that AI had helped write or substantially revise them. That figure, however, blends together pages from before and after consumer AI writing tools became widely available, so it undercounts the influence of generative AI on newer content.

When Pew removed older pages and looked only at webpages published after ChatGPT’s release, the estimated share with signs of AI authorship climbed to 35%. In other words, roughly one in three newer webpages in the sample appeared to have been produced with AI involvement.

Metric Pew finding What it means
Webpages analyzed Nearly 500,000 English-language pages Large-scale snapshot of the web over about five years
Random sample from July 2026 About 10% with significant AI signs Includes older pages from before AI writing tools were common
Pages published after ChatGPT launch 35% with signs of AI authorship Suggests rapid adoption of AI in recent web publishing
.com domains About 10 times the AI-authorship rate of .edu and .gov Commercial sites appear far more likely to use AI-generated text
.edu and .gov domains About 1% Institutional sites appear much less reliant on AI writing
.org domains 4.6% Nonprofit and advocacy sites fall between commercial and institutional sites

Why the timing matters

The study lands at a moment when concerns about the automation of the web are intensifying. Just days before Pew’s report, Cloudflare said bot traffic had overtaken human traffic across the internet, reaching that threshold sooner than the company expected. Pew’s work does not measure browsing habits, but together the reports point to a digital ecosystem in which automated systems are increasingly responsible for both creating and consuming content.

That matters because the internet has long depended on an assumption that human authorship is the norm. Search engines, social platforms, ad networks and web archives all rely on a certain amount of trust in the material they index. If AI-generated pages continue to proliferate, that trust could become harder to preserve, especially when the same systems that generate content are also used to rank, summarize or redistribute it.

How did Pew measure AI authorship?

Pew used a detection model from Open Pangram, a firm that specializes in identifying AI-written text. The research approach was not a direct audit of authorship records, because webpages rarely disclose whether a human, a machine or a hybrid workflow produced the final copy. Instead, the study looked for linguistic patterns associated with AI generation.

That means the results should be read as probabilistic rather than absolute. The report itself acknowledges that detection tools can make mistakes, including false positives where human-written pages are flagged as AI-generated. Even so, Pew argues the scale of the dataset makes the overall trend meaningful, at least directionally.

In practical terms, the study offers a broad estimate of how generative AI has changed the publishing environment since late 2022. It does not prove that every page with AI-like language was fully machine-written, but it does suggest that AI influence is now common enough to alter the texture of the modern web.

What “signs of AI authorship” really means

The phrase refers to pages that show linguistic patterns associated with AI generation or AI-assisted editing. That can include fully generated articles, rewritten drafts, or content polished by a model after a human wrote the initial version.

Pew’s framing is important because many publishers have moved beyond a simple human-versus-machine divide. In newsroom, marketing and SEO workflows, AI is often used for outlining, drafting, translation, optimization or cleanup rather than for full automation. The study’s estimate therefore reflects a spectrum of machine involvement, not just fully synthetic pages.

Pew’s core conclusion is that AI authorship has become visible enough across the post-ChatGPT web to be measured at scale, even if no detector can perfectly identify every page.

How much of the web is AI-written by domain type?

Commercial websites appear to be using AI content far more aggressively than institutional sites. Pew found that .com domains showed AI-authorship signals at roughly 10 times the rate seen on .edu and .gov webpages.

.edu and .gov sites both came in at around 1%, while .org sites registered a 4.6% rate. The gap likely reflects differences in purpose, editorial standards and risk tolerance. Universities and government agencies often operate under more formal review processes, while commercial publishers may face greater pressure to publish quickly, cheaply and at scale.

That said, domain name alone is not a perfect proxy for editorial quality or AI usage. A .com site may host a rigorously edited publication, while a .org domain might publish a broad mix of professionally written and AI-assisted material. Still, the pattern is consistent with what many content strategists and publishers have suspected: the financial incentives to automate are strongest in the commercial sector.

Why are em dashes and Oxford commas showing up more often?

Because AI-generated writing has left detectable stylistic fingerprints that are increasingly common in public web content. Pew says it found a rise in features often associated with machine-assisted text, including em dashes, Oxford commas and formulaic contrastive phrasing such as “it’s not X, it’s Y.”

These markers do not automatically mean a page is AI-written. Skilled human writers use them too. But when they appear more frequently across large samples, they can indicate that content pipelines are being shaped by the conventions of generative models. In effect, the web may be developing a more standardized voice as AI tools influence drafts, edits and rewrites.

That trend raises a separate issue for publishers and SEO teams: as AI content becomes more common, style itself becomes less reliable as an indicator of credibility or originality. Writers who use models as assistants may unintentionally produce pages that resemble machine output, while pure human writing can begin to look algorithmic because AI-generated prose has become a shared stylistic baseline.

Why the report matters for publishers, search and SEO

The study is important because it helps quantify a change that has often been discussed anecdotally. Since ChatGPT’s launch, publishers have debated whether AI would increase output, lower costs and improve efficiency—or flood the internet with generic pages that are harder to trust, rank and differentiate.

For search engines, a larger volume of AI-assisted text creates both opportunity and risk. On one hand, generative tools can help smaller publishers produce timely coverage and structured content more efficiently. On the other, they can also fuel low-quality mass publishing, duplicate answers, and thin pages designed primarily to capture search traffic.

For readers, the biggest concern is not whether a page was touched by AI, but whether it remains accurate, original and useful. Yet as AI use spreads, the burden on audiences to judge quality increases. Users may have to rely more heavily on source transparency, editorial reputation and fact-checking signals.

What publishers may do next

Many publishers are likely to respond by tightening disclosure policies, editorial review standards and internal rules for AI usage. Others may embrace AI more openly as a production tool while emphasizing human oversight and reporting.

Likely strategies include:

  • Clear labeling of AI-assisted content
  • Editorial review before publication
  • Fact-checking for model-generated claims
  • Style guidelines to preserve a distinct brand voice
  • Stronger content provenance systems

At the same time, the market may reward sites that can prove a human editorial process. As synthetic content becomes more common, authenticity itself may become a differentiator.

How reliable are AI detection tools?

They are useful at scale, but imperfect in individual cases. Detection systems can mislabel human writing as machine-generated and can also miss AI-assisted pages that have been heavily edited.

That limitation is especially important for a study like Pew’s, which is estimating a population-wide trend rather than making claims about any one article or website. The report does not suggest that every flagged page was entirely drafted by a chatbot. Instead, it points to a broad increase in AI-like language patterns that align with wider changes in publishing behavior.

In other words, the strongest value of the study is not courtroom-level proof. It is evidence of direction: the web after ChatGPT is measurably different from the web before it.

What the data says about the post-ChatGPT web

The most striking takeaway from Pew’s report is not simply that AI-authored pages exist, but that they are now widespread enough to be visible in a large public archive. That suggests AI is no longer a niche experiment used by a few early adopters. It is part of routine web production.

Because Common Crawl captures a wide range of publicly accessible content, the findings likely include pages from blogs, commercial sites, reference pages, marketing pages and other online properties. The study does not spell out a category-by-category breakdown in the source summary, but the domain-level differences imply that some sectors are adopting AI much faster than others.

That uneven adoption is likely to continue. Some publishers will use AI mainly behind the scenes. Others will embrace full automation. And many will settle on a hybrid model in which a person starts the work and a model finishes it, or vice versa.

Timeline: how AI-authored web content became visible

The shift Pew measured did not happen overnight. It accelerated alongside the consumer rollout of generative AI and the normalization of chatbot-based writing tools.

Date / Period Development Why it matters
Before November 2022 Older webpages predate mainstream consumer AI writing tools Sets the baseline for comparing pre- and post-ChatGPT publishing
November 2022 ChatGPT launches Marks the start of broad public access to generative writing tools
2023–2025 AI writing tools spread rapidly across publishing and marketing workflows Increases the amount of machine-assisted online text
July 2026 Pew samples 10,000 webpages from a larger archive Provides a recent snapshot of AI-authorship signals
August 2026 Pew releases its findings Quantifies the post-ChatGPT transformation of the web

What it means for trust online

If one-third of new webpages show signs of AI authorship, the challenge is no longer whether AI has entered web publishing. The challenge is how the internet distinguishes between helpful automation and content that exists mainly to occupy search results, generate ad impressions or manipulate discovery systems.

That distinction will likely become more important over time. Search engines are already under pressure to identify which pages add original reporting or expertise and which are merely reassembled from existing material. As AI lowers the cost of producing text, the value of human judgment, reporting access and subject-matter accountability may rise.

The study also hints at a broader cultural change. When machines increasingly write the internet and other machines increasingly consume it, the web becomes less like a library of human expression and more like a dynamic machine-to-machine information layer. That transformation may improve speed and scale, but it also makes provenance, verification and editorial standards more important than ever.

Bottom line

Pew’s research suggests that AI writing has moved from novelty to normality on the post-ChatGPT web. The estimate that 35% of webpages published after November 2022 show signs of AI authorship is one of the clearest large-scale indicators yet that generative tools are fundamentally changing online publishing.

Even with the limitations of detection technology, the direction of travel is hard to miss: more webpages are being produced with AI help, and the gap between human and machine content is getting harder to see.

Frequently asked questions

How much of the web appears to be AI-written after ChatGPT launched?

Pew estimates that 35% of webpages published after ChatGPT’s November 2022 release show signs of AI authorship. That figure comes from a large archive-based sample and reflects pages that appear to be written or heavily edited with AI assistance.

How did Pew measure AI authorship on webpages?

Pew analyzed nearly half a million English-language webpages from the Common Crawl archive and used Open Pangram’s detection technology to identify pages with linguistic patterns associated with AI writing or AI-assisted editing.

Which websites were most likely to show AI writing signals?

Commercial .com domains were most likely to show AI writing signals, at roughly 10 times the rate of .edu and .gov sites. Pew said .edu and .gov domains were around 1%, while .org sites were at 4.6%.

Are AI detection tools always accurate?

No, AI detection tools are not always accurate. They can flag human-written pages as machine-generated and can miss heavily edited AI content, so Pew’s findings should be treated as a large-scale estimate rather than a definitive authorship audit.

Why does this study matter for publishers and search engines?

This study matters because it shows AI is now a major part of online publishing, which affects trust, editorial standards and search quality. As more pages are generated or edited by AI, publishers and search engines face greater pressure to verify originality and usefulness.

Share this 🚀