Study: ChatGPT's Sources Only Overlap Google's Top 10 by 4%

Toronto researchers compared the sources cited by GPT-4o, Claude, Gemini and Perplexity with Google's top 10 results: just 4 to 15% overlap. What it means for AI visibility.

July 23, 2026 · 9 min

Overlap between Google's top 10 and the sources cited by AI engines — University of Toronto study, 2026

Four researchers at the University of Toronto ran a simple experiment. Across 1,500 shopping questions, they compared the sources cited by GPT-4o, Claude, Gemini and Perplexity with Google's top 10 results. The finding is striking: 4 to 15% of sources in common, no more. In other words, ranking well on Google is barely enough anymore to get cited by AI. And for the first time, a study also puts hard numbers on what really matters instead.

What does the study show, in four numbers?

The study comes down to four numbers. AI engines mostly cite sites that are absent from Google's top 10 results (4 to 15% overlap). They prefer news articles and independent sites (46 to 65% of citations) over forums. They cite content that is 40 to 70% more recent than Google's. And for well-known brands, they lean mostly on what they learned during training.

  • 4.0%: that's what the sites cited by GPT-4o and Google's top 10 results have in common, on the same questions. Gemini: 11.1%, Claude: 12.6%, Perplexity: 15.2%.
  • 46 to 65% of the sources cited by AI engines are news articles or independent sites, versus 41% at Google; forums and social media drop to 1-8%, versus 34% at Google.
  • 40 to 70% more recent: the articles cited by AI engines are noticeably younger than Google's, in every product category.
  • 16% of the products and brands recommended by the tested engine rested on no page found online: the model was answering purely from what it had learned during training.

The study is titled "Navigating the Shift: A Comparative Analysis of Web Search and Generative AI Response Generation." It is authored by Mahe Chen, Xiaoxuan Wang, Kaiwen Chen and Nick Koudas of the University of Toronto, and was posted on arXiv in January 2026. It covers 1,500 English-language questions such as "best reliable smartphones" or "X vs Y," spread across ten product families.

Is ranking first on Google enough to get cited by ChatGPT?

No. On the same questions, the sites cited by GPT-4o and Google's top 10 results share only 4% — 11 to 15% for Gemini, Claude and Perplexity. The researchers' conclusion is clear: AI engines don't just reshuffle Google's results into a new order. They draw on largely different sources.

This number confirms what GEO specialists have been seeing for two years: AI citation follows its own rules. A site can rank at the top of Google and stay completely absent from ChatGPT or Perplexity answers. And the reverse is true: a site that ranks poorly on Google can become an AI's favorite source. Being visible on one tells you nothing about the other. These are two separate playing fields, each to be tracked and improved separately.

Which sources do AI engines prefer?

AI engines give a lot of room to news articles, comparison sites and independent sites. Conversely, they almost completely ignore forums and social media. Here is the breakdown measured by the study:

EnginePress & independent sitesForums & social mediaOfficial brand sites
Google (top 10)41%34%26%
Claude65%1%
GPT-4o57%8%
Perplexity50%39%
Gemini46%46%

Two concrete lessons. First lesson: you often hear that Reddit dominates AI answers. That's not what the study shows. Here, forums and social media barely count in the final sources. What really matters is being cited by credible sites: press, comparison sites, studies. Second lesson: official brand sites keep a real place. They account for 46% of Gemini's citations and 39% of Perplexity's. And when the person wants to buy ("where to buy," "price of"), every engine climbs to 52-68% brand-site citations. A brand site built to be citable therefore captures those searches directly — the ones where the visitor is ready to make a purchase.

Why does freshness weigh so heavily?

Because AI engines are built to go looking for content that is more recent than Google's. Take consumer electronics. Look at the age of the cited articles — that is, how long ago they were published. Half of the cited articles are less than 62 days old for Claude, 80 days for GPT-4o and 90 days for Perplexity; at Google, that limit rises to 130 days. In automotive, the gap is even sharper: 148 to 217 days for the AI engines, against 493 days for Google — roughly three times more recent.

Google can surface a reference page two years old, if the site hosting it is judged trustworthy. AI engines, on the other hand, look for the most up-to-date information possible. When they go fetch pages from the web, they favor those published or updated recently. This confirms what we explained in our article on publishing frequency: keeping your important pages up to date, with a visible and honest update date, is no small detail. It's a signal AI engines genuinely look at.

Known brands vs niche brands: two very different games

This is the study's most surprising result. For a well-known brand, the ranking the AI gives comes mostly from what it learned during training; the pages it finds only confirm it. For a little-known brand, it's the opposite: the AI genuinely depends on the pages it finds, and both their content and their order change the answer.

To prove it, the researchers scrambled the pages fed to the model. They changed the order of the excerpts, swapped one brand for another, or forced the model to use only the excerpts provided. For well-known brands, the final answer barely moved. The authors sum it up: at that point, the AI confirms what it already knows and discovers nothing new. For little-known brands, the same shuffling test moves the ranking almost twice as much (4.15 positions of shift on average, against 2.30 for well-known brands). A sign that the model then actually starts to search and to rely on what it reads: this is the grounding mechanism working at full strength.

But this memory has a downside. 16% of recommended brands rested on no page found online. On SUV searches, 73% of the times Infiniti was cited, no page backed the answer, against 6% for Toyota. In other words, the AI sometimes recommends from memory, with no recent evidence. And if your brand is neither in its memory nor in the pages it finds, it simply doesn't appear in the answer.

"This is the best news in the study for small players: on little-known brands, AI engines really do read the pages they find. Every page that's easy to cite — a direct answer, sourced facts, well-structured information — has an immediate effect on the answer. A big brand, meanwhile, leans mostly on the reputation it already has," says Nicolas Meridjen, founder of LightSpot.

What should you take away for your strategy?

Three concrete consequences, from the most important to the least:

  1. Measure your AI visibility separately from your SEO. With 4 to 15% of sources in common, your position on Google tells you almost nothing about your presence in ChatGPT, Gemini or Perplexity answers. The first step is to take stock of this specific ground: being visible on ChatGPT, Gemini and Perplexity is not something you can read in Google's Search Console.
  2. Work both fronts. On one side, get cited by credible sites (46 to 65% of citations): that's won with original data, studies and press-relations work. On the other, your own site captures buying searches (52 to 68% brand citations), provided it's easy to cite: answers that stand on their own, complete and well-structured information, facts with their source.
  3. Keep your pages up to date. AI engines cite content 40 to 70% more recent than Google: it's a matter of consistency. Update your important pages with a visible date, publish regularly, and make those freshness signals readable by machines.

And if your brand is a small player rather than a leader, good news: yours is exactly the case the study points to as the most sensitive to the pages found online. Optimization work pays off faster for you than for the big, already-established brands.

Where do these numbers come from, and what are the limits?

These numbers come from a solid protocol. The researchers asked 1,000 ranking questions across ten product families, 200 head-to-head comparisons between two brands (100 well-known brands, 100 niche brands) and 300 electronics questions, sorted by intent. They fed them to Google Search, GPT-4o, Claude (Sonnet 4.5), Gemini 2.5 Flash and Perplexity Sonar Pro, all in web-search mode. Every cited address is reduced to its site, classified by source type, and dated using information hidden in the page's code. To measure overlap with Google, they look at what each AI's list of cited sites and Google's top 10 results share (the Jaccard index, for the specialists).

Three limits to keep in mind before generalizing. First, the questions are in English and revolve around product comparisons: the proportions may look different in other languages, or on general-knowledge questions. Second, the tests that scramble the supplied pages only covered GPT-4o. Third, the study is a snapshot taken in January 2026, and the way AI engines go about finding their sources changes fast. That said, the direction of the results lines up with other work on the subject, including the study that founded GEO (Princeton / ACM SIGKDD 2024).

To find out where your own site stands — big brand or small player — the LightSpot audit runs 46 checks (AI visibility and SEO) for free across 3 pages of your site. You get a score out of 100 and the top 3 fixes to make first.

FAQ

What does the University of Toronto study say about AI engine citations?
Posted on arXiv in January 2026, the study tested 1,500 shopping questions. It compares the sources cited by GPT-4o, Claude, Gemini and Perplexity with Google's top 10 results. Four findings stand out: only 4 to 15% of sources overlap with Google, a clear preference for news articles and independent sites (46 to 65% of citations), content that is 40 to 70% more recent, and a heavy reliance on what the model already learned during training when a brand is well known.
Does ranking well on Google help you get cited by ChatGPT?
Barely, according to this study. On the same questions, the sites cited by GPT-4o and Google's top 10 results share only 4% (11.1% for Gemini, 12.6% for Claude, 15.2% for Perplexity). AI engines don't simply reshuffle Google's results — they pull from other sources entirely. Ranking well on Google is still useful, but it no longer guarantees you'll show up in AI answers. These are two different things, each to be tracked and worked on separately.
Which sources do AI engines cite the most?
First, news articles and independent sites (magazines, comparison sites, review sites): they make up 65% of Claude's citations, 57% for GPT-4o, 50% for Perplexity and 46% for Gemini, against 41% at Google. Forums and social media are almost absent from AI answers: just 1 to 8%, versus 34% at Google. Finally, official brand sites carry real weight for Gemini (46%) and Perplexity (39%), especially when the person is trying to buy.
How can a little-known brand show up in AI answers?
Little-known brands are exactly where the pages the AI finds matter most. The study shows that when the model has learned almost nothing about a brand during training, it genuinely searches: its answer then depends directly on the pages it finds. For well-known brands, by contrast, it mostly confirms what it already knows. In short, for a small player, three things have maximum impact: pages that are easy to cite (a clear, direct answer, well-structured information, facts with their source), recent content, and mentions in the press and on independent sites.
How do I know if my site is citable by AI engines?
Run a LightSpot audit on your site's URL. It checks 46 points: 21 on classic SEO and 25 on AI visibility (can AI crawlers read your pages, is your information well structured, does each passage stand on its own, are your facts sourced, is your content recent). You get a score out of 100, a letter grade from A to F, and the top 3 fixes to make first. The 3-page audit is free, with no account and nothing to install.
Portrait de Nicolas Meridjen, Fondateur de LightSpot.ai — outil d'audit de visibilité IA (46 critères SEO + GEO)

Nicolas Meridjen

Fondateur de LightSpot.ai — outil d'audit de visibilité IA (46 critères SEO + GEO)

Je construis LightSpot.ai et j'analyse comment les moteurs de recherche IA (ChatGPT, Perplexity, Google AI Overviews) choisissent les sources qu'ils citent. J'écris sur le GEO et le SEO à partir de données d'audit réelles.

On the same topic