What happens when AI learns from info increasingly shaped by AI itself?
As AI-generated material spreads and chatbots become gateways to information, a new challenge is emerging: distinguishing original evidence from content repeatedly summarised and rewritten by machines
)
artificial intelligence (AI)
Listen to This Article
Artificial intelligence (AI) is changing not just how information is produced, but how it moves across the internet. As people increasingly turn to AI assistants for answers, the web itself is filling with AI-generated and AI-assisted material. A Pew Research Center analysis of nearly 490,000 webpages found that 10 per cent of pages in its July 2026 sample showed significant signs of AI authorship, rising to more than one-third among pages published after ChatGPT's release.
This raises a problem beyond hallucinations: what happens when AI starts learning from a web increasingly shaped by AI itself?
From AI-generated content to an AI feedback loop
Consider an example. A technology website publishes a review of a new gadget based on hands-on testing. An AI system summarises that review. Another website uses the summary to create a buying guide. A third site rewrites that guide, perhaps adding information generated by an AI chatbot. Eventually, another AI system searches or retrieves these pages while answering a consumer's question about the gadget.
The answer may look well supported because several websites appear to say similar things. But those websites may not represent several independent sources. They may all trace back to the same original review.
Also Read
This is where the idea of an AI feedback loop becomes important. Information can move through several layers of websites and AI systems, with each layer making the material appear more established while adding little new evidence.
The problem becomes more serious when the original information is incomplete or wrong. A small error can be copied into several summaries, buying guides and explainers. Once the same claim appears repeatedly, a future AI system may encounter it in multiple places and treat that repetition as evidence of reliability.
Repetition, however, is not the same as verification. This is not necessarily how current AI models work in every case. Leading model developers use different datasets, filtering systems, human feedback and other methods to build and evaluate models. Many models also increasingly rely on live search or retrieval systems rather than depending only on their original training data. But the growing amount of AI-assisted material online creates a new challenge for the broader information ecosystem.
The web is becoming more machine-written
Pew's analysis provides some clues about how this change is taking place. The researchers found that signs of AI authorship became increasingly common after the release of ChatGPT. In July 2026, about one in 10 webpages with a dot-com domain showed significant signs of AI authorship, compared with 4.6 per cent of dot-org pages and around 1 per cent of dot-edu and dot-gov pages.
Pew also found that some language patterns associated with AI-generated writing have become more common online. For example, the use of em dashes in pages published after ChatGPT's release had increased substantially compared with earlier samples, while certain AI-associated words and phrases also became more frequent.
These signals are not proof that a particular webpage was written by an AI model. But collectively, they show how the language generated by AI is becoming part of the wider internet.
The shift is also not limited to articles. Product descriptions, travel guides, software documentation, marketing material, financial explainers, social media posts and other forms of online information can now be generated or heavily edited with AI. That changes the raw material from which future information systems operate.
AI is becoming a new gateway to information
At the same time, people are changing how they access that material. Instead of typing a question into a search engine and choosing from a list of websites, users can ask an AI assistant and receive a direct answer.
The Reuters Institute's Digital News Report 2026 found that 10 per cent of respondents globally said they used AI chatbots for news each week, up from 7 per cent the previous year. Among people under 35, the figure was 16 per cent.
That does not mean AI chatbots have replaced search engines. But it shows that chatbots are becoming another layer between people and the original information source. For publishers, this has a direct financial consequence.
The Reuters Institute said Google traffic from organic search to more than 2,500 websites fell 33 per cent globally between November 2024 and November 2025. In the US, the decline was 38 per cent. Publishers surveyed by the institute expected search traffic to fall further over the next three years, creating a difficult cycle for the web.
A publisher spends money on reporting, testing, research or analysis and puts the information online. An AI system may use that information to answer a user's question without requiring the user to visit the original site. The user gets the answer, but the publisher may receive less traffic.
If that continues, there is a risk that producing original information becomes harder to finance while producing derivative content becomes cheaper.
The problem is not only 'AI content'
The bigger concern, therefore, is not that AI-generated content exists. AI-assisted writing can be useful. It can help people translate, edit, summarise and organise information. Companies can use AI to produce routine material, while journalists and researchers can use it to analyse large volumes of information.
The problem begins when synthetic information becomes indistinguishable from independently produced information and is then fed back into systems that generate more information.
This is particularly difficult because AI models are built to produce plausible responses, not to prove that every statement has an independent source.
A model may see the same claim across several webpages. Unless it can establish where those claims originated and whether they were independently verified, the apparent agreement may be misleading. This is why provenance could become increasingly important in AI systems.
Instead of asking whether a piece of information exists online, future systems may need to establish where it came from, who verified it, whether multiple sources are genuinely independent and how the claim changed as it moved across the internet.
That could push AI systems towards greater use of source tracking, citations, retrieval systems and other ways of establishing information lineage.
The next problem may be information quality
The internet's first major challenge was getting enough information online. Its next challenge may be determining which information deserves to be trusted. The emergence of AI does not remove the need for original reporting, research or expertise. In some areas, it may make them more important.
As AI-generated material increases, the value of information that can be traced to a real experiment, interview, document, dataset or first-hand observation could rise.
The question for AI companies is therefore no longer only whether their models can produce better answers. It is also whether those answers can remain connected to reliable, independently produced information.
Otherwise, the web could enter a strange cycle where humans create information, AI learns from it, AI generates new information, new material enters the web and future AI systems learn from that material again.
The risk is not that machines will suddenly know nothing. It is that they may increasingly know the same things because they have repeatedly encountered one another's versions of them.
More From This Section
Don't miss the most important news and views of the day. Get them on our Telegram channel
First Published: Oct 07 2026 | 3:17 PM IST
