Study Shows Leading Chatbots Offer Narrower, More Homogeneous Knowledge Than Google
A University of Copenhagen analysis of 27 large language models across 155 topics reveals that even the most advanced chatbot, GPT‑5, provides at least 18.7% less epistemic diversity than a standard Google search.

Researchers at the University of Copenhagen have completed the most extensive comparative test of large language models (LLMs) to date. They evaluated 27 prominent LLMs on 155 distinct subjects, generating 200 query formulations for each subject. The experiment produced roughly 70 million individual statements, allowing the team to quantify how broadly each model draws from the pool of human knowledge and to compare that breadth with traditional search tools.
The core finding is stark: every LLM examined delivers less epistemic diversity than a conventional Google search. Even the best‑performing model, identified as GPT‑5, still lags behind Google by a minimum of 18.7% in terms of the variety of information it can surface, meaning that users receive a narrower slice of the available knowledge when they rely solely on a chatbot.
Growth in Diversity Over the Last Three Years
Contrary to the popular narrative of an inevitable decline in knowledge breadth, the study shows a substantial increase in epistemic diversity among LLMs over the past three years. This upward trend, however, does not close the gap with search engines, which continue to outperform chatbots across the board, especially on niche or highly specialized topics.
The improvement is uneven. Retrieval‑augmented generation (RAG) systems, which combine a language model with external document retrieval, show a measurable boost in diversity. In contrast, the largest models—those with the greatest number of parameters—are paradoxically less diverse than their smaller counterparts, suggesting that sheer scale does not automatically translate into broader coverage.
Language Bias and the Risk of Knowledge Collapse
A second, more subtle issue emerges from the analysis of parametric knowledge: English dominates the internal representations of all tested models, even when the query concerns topics that are specific to non‑English‑speaking countries. This linguistic skew reduces the visibility of local sources and cultural perspectives, effectively marginalising non‑English scholarship.
The authors attribute this bias to the way LLMs compress massive training corpora. By prioritising the most frequent patterns, the models inherently filter out less common information. If future training pipelines incorporate text generated by other AIs—a practice that is already gaining traction—the phenomenon known as "knowledge collapse" could accelerate, further homogenising the knowledge base.
Proposed Remedies and Measurement Tools
To mitigate these risks, the researchers introduced a quantitative metric for assessing epistemic diversity in chatbot responses. They urge AI developers to adopt this measurement as a standard benchmark during model evaluation, arguing that consistent monitoring will reveal when diversity begins to erode.
In parallel, they recommend that end‑users consult multiple information sources rather than relying exclusively on a single chatbot summary. This habit can help preserve a broader view of the topic at hand and act as a safeguard against inadvertent echo chambers.
- Adopt the new epistemic‑diversity metric in model testing
- Incorporate diverse, multilingual data during training
- Limit the proportion of AI‑generated text in training sets
- Encourage retrieval‑augmented architectures for better coverage
The study does not yet observe a concrete "knowledge collapse" in the current generation of chatbots, but it warns that the trajectory will depend heavily on training practices. The uneven progress across languages and model types suggests that without deliberate intervention, the gap with search engines may persist or even widen over time.
For English‑speaking organisations, the practical implication is clear: while chatbots can accelerate routine information retrieval, they should not replace comprehensive search strategies. Decision‑makers are advised to cross‑verify chatbot outputs with traditional search results and to maintain a diversified set of knowledge sources, especially when operating in multilingual or region‑specific contexts.
Academic institutions, too, are urged to scrutinise the composition of their training datasets, ensuring that non‑English literature and under‑represented viewpoints are not systematically excluded. Such vigilance will help preserve the richness of scholarly discourse in the age of AI.
Policymakers may consider guidelines that limit the proportion of synthetic text fed back into model training pipelines, thereby curbing the feedback loop that fuels homogenisation. Transparent reporting of dataset provenance could become a regulatory requirement in the near future.
Finally, the authors stress that the metric they propose is not a one‑off tool but a dynamic instrument that should evolve alongside AI capabilities. Continuous refinement will be essential to keep pace with the rapid advancements in model architecture and data acquisition.
In summary, the Copenhagen analysis paints a nuanced picture: LLMs are becoming more diverse, yet they remain markedly narrower than traditional search engines, and language bias persists. Proactive measures—both technical and policy‑driven—are needed to prevent a gradual erosion of the collective knowledge pool.
Sources
- AI chatbots give us a narrow slice of knowledge: Researchers warn of 'knowledge collapse'TechXplore · September 28, 2026
- [2510.04226] What and Whose Knowledge? Measuring Epistemic Diversity in Large Language ModelsarXiv · September 28, 2026



