The 60% Problem: Where ChatGPT Really Finds Its Sources

A theory has taken over GEO circles in 2026. It goes like this: ChatGPT splits every prompt into smaller search queries behind the scenes, so if you find those queries and rank for them on Google, you get cited. Rank for the fan-out queries, win AI search.

Grow & Convert, a content marketing agency, tested that theory in a study published in March 2026. They took 100 buying intent prompts, captured the fan-out queries ChatGPT ran for each one, and checked whether the sources ChatGPT cited actually ranked for those queries on Google and Bing.

The answer: only about 40% of cited sources appeared anywhere in the first 10 pages of Google or Bing for the fan out queries. The other 60% came from somewhere else entirely.

That one number reshapes how a GEO strategy gets planned. Below is what the study found, the 4 most likely explanations for the missing 60%, and what we do differently at Pro Start Marketing because of data like this.

What is a fan out query?

A fan out query is one of the 2 to 4 smaller, more specific searches ChatGPT generates from a single user prompt when it decides to search the web. Ask ChatGPT for the best help desk software for small teams and it does not run one search. It breaks the request into narrower queries, pulls results for each, and assembles its answer from those sources.

Marketers noticed this and drew a straight line: fan out queries are visible in the ChatGPT interface, Google rankings are trackable, so ranking for the fan out queries looks like a repeatable path to citations. Entire tools and service packages now sell fan out query tracking on that logic.

The logic sounds clean. The data says it explains less than half of what ChatGPT cites.

How Grow & Convert tested the fan out theory

Grow & Convert compared the sources ChatGPT cited for 100 buying intent prompts against the Google and Bing rankings for every fan out query those prompts generated. The methodology matters here, because most GEO claims float around with no methodology at all. Seven details define the dataset:

  1. All 100 prompts carried buying intent, such as “best help desk software for small teams,” “top accounts payable automation tools,” and “best standing desk for a home office.”
  2. The prompts spanned B2B, B2C, software, services, and physical products.
  3. The tests ran through the browser interface rather than the API, so the data reflects what a real user sees.
  4. 83 of the 100 prompts triggered a live web search. ChatGPT answered the other 17 from its own knowledge, which matches the agency’s earlier finding that ChatGPT searches the web for more than 80% of product related prompts.
  5. Each search generated 2 to 4 fan out queries.
  6. Google results came through SerpAPI. Bing results were checked manually because the API returned unreliable data, so the Bing sample is smaller.
  7. The comparison went 10 pages deep into each SERP and ran at 3 levels: per prompt, per business category, and exact URL match versus domain match.

The buying intent restriction is deliberate, and we agree with it. Top of the funnel questions get answered by LLMs without brand mentions or links. Bottom of the funnel, product intent prompts are where AI visibility converts into pipeline, so they are the only prompts worth measuring.

The results: a 40% ceiling with wild swings

Only 27% of the sources ChatGPT cited ranked anywhere in Google’s first 10 pages for the fan out queries. Bing came in at 23%. The two engines overlapped by roughly 10%, which puts the combined figure at about 40% of citations appearing somewhere in either engine. Flip that around and you get the finding that matters: 60% of ChatGPT’s cited sources did not rank in the first 10 pages of Google or Bing for the queries ChatGPT itself ran.

The averages hide how erratic the prompt level results were. 20 of the 100 prompts showed zero overlap between citations and the fan out SERPs, even though ChatGPT searched the web for them. Exactly 1 prompt out of 100 showed complete overlap.

Category explained less of the variance than expected. Physical products had the lowest overlap, which Grow & Convert attribute to ChatGPT tapping shopping feeds the study did not track. B2C SaaS also ran weak, since app store listings often replace vendor websites as the source for those products. Outside those 2 cases, overlap swung prompt by prompt rather than category by category. The swing comes from how heavily ChatGPT leans on live search versus training data for a given answer, not from what the user wants to buy.

Rankings still mattered inside that 40% window. Among the cited sources that did rank, page 1 of Google alone held a third of the matches, and the first 3 pages held about 66%. Matches thinned out fast after page 3. Ranking higher raises your odds of citation. It just operates on a smaller share of the outcome than most GEO advice admits.

What the 40% number does not mean

The finding does not break the link between search rankings and AI visibility. ChatGPT still ran a live web search for 83 of the 100 prompts, and the agency’s earlier research measured web search triggering on more than 80% of product related prompts. Search remains the largest visible pipe feeding citations, and Google rankings remain the most controllable input into that pipe.

What the number breaks is the idea of precision targeting. Producing a single page to “rank” for a specific AI prompt by chasing its fan out queries earns a citation less than half the time, and the queries themselves change between runs. The same warning covers the shortcut tactics circulating alongside the fan out theory, such as bolting on FAQ pages, bullet point summaries, and llms.txt files. Each of those aims at the same fantasy of direct, page level control over a system that offers none.

The honest read: AI search rewards broad, durable inputs, such as domain authority, training data presence, and topic coverage, over narrow and precise ones. That distinction drives everything below.

Where the other 60% of citations come from

Nobody outside OpenAI can see exactly where the remaining citations originate, and Grow & Convert are honest that this part is inference. Four explanations fit the evidence in their data.

1. Searches you never see

The fan out queries visible in the browser are probably a subset of the queries ChatGPT actually runs. QueryBurst has documented why OpenAI would distribute extra searches across its own servers in parallel, using different engines and different query variations. Running everything sequentially inside the browser would make responses painfully slow. Any source pulled in through those server side searches enters the citation pool through queries no tracking tool can observe.

2. Cached answers from other people’s prompts

Millions of users ask nearly identical buying questions, and reusing stored results costs OpenAI far less than running fresh searches every time. OpenAI already gives API developers discounts for cached responses, so applying the same economics inside ChatGPT is a reasonable assumption. Caching has a hard limit as an explanation, though. OpenAI’s own documentation puts cache lifetimes at 5 to 10 minutes, occasionally stretching to an hour. Anything stale by months needs a different explanation, which leads to the third theory.

3. Memorized training data

ChatGPT retains specific pages, and sometimes specific URLs, from its training data. Research presented at NeurIPS demonstrated that large language models reproduce portions of their training data verbatim, and Grow & Convert found the fingerprints of that memory in their citation lists. Several cited URLs were real, live pages carrying outdated titles that no longer matched the current page. In the sharpest example, ChatGPT cited the exact same Zapier URL twice in a single source list, once with a title dated 2025 and once dated 2026, while the live article carried a third, different headline. A cache that expires within an hour cannot produce a title that is a year out of date. Training data can. The likely story is that one citation came from memory, the other from a live search, and ChatGPT stitched both into the same answer.

4. Citations that do not exist

Roughly 10% of the citations in the study led to error pages. These look legitimate, with plausible URLs and plausible titles, and point at nothing. Academic research puts the problem even higher: a February 2026 study measured hallucination rates above 14% for cited sources, and a 2023 paper in Nature documented the same pattern. Part of the missing 60% is not coming from anywhere. ChatGPT invented it.

For brands, this one cuts both ways. Your domain occasionally gets credited for pages that never existed, and users occasionally get routed to dead ends wearing your name. One more reason a real, deep content library beats a thin site: the more genuine pages you publish, the more of ChatGPT’s guesses about your domain resolve to something useful.

Why GEO refuses to behave like SEO

SEO runs on a repeatable loop: pick a keyword, publish a relevant page, handle the on page basics, earn links, rank. The same input produces roughly the same output for every searcher, every day. Nothing about LLM answers works that way, and the study stacks up 5 separate sources of chaos:

  1. SparkToro’s research found that ChatGPT’s brand recommendations barely repeat. Two users entering the identical prompt receive different brand lists, in different orders.
  2. The same prompt generates different fan out queries on different runs.
  3. ChatGPT sometimes skips web search entirely for a prompt it searched the day before.
  4. Free accounts and paid accounts run different model capabilities.
  5. Every response gets shaped by personal context: the user’s chat history, saved memory, business details, and phrasing habits.

We call that last point the Invisible Prompts problem. The literal prompt a user types is only a fraction of the effective prompt the model answers, because the model folds in everything it knows about that specific user. The closest analogy is word of mouth: a friend recommending software weighs everything they know about you, and so does the model. You cannot see the effective prompt. You cannot see all the fan out queries. And the fan out queries you can see explain only about 40% of citations.

A strategy built on ranking one page for one prompt collapses under those conditions. The unit of optimization has to move up a level, from prompts to topics.

Why classic SEO still carries your GEO results

None of this makes SEO optional. Two mechanisms in the data explain why publishing SEO content remains the strongest GEO lever available.

Mechanism 1: your content feeds the training data. OpenAI runs web crawlers for live search and for model training. Every detailed page you publish becomes a candidate for the training corpus, and the training corpus supplies a large slice of the 60% that never touches a SERP. No submission form exists for training data. Publishing substantive, crawlable content about your product is the entire eligibility requirement.

Mechanism 2: ChatGPT trusts domains more than pages. Grow & Convert reran their overlap analysis with a looser rule: count a citation as matched when any page on the cited domain ranked for a fan out query, instead of requiring the exact URL. The match rate jumped from 27% to about 50%. ChatGPT cites domains it treats as topical authorities even when the page that ranks and the page it cites are different. SparkToro’s data points the same direction: the 3 most mentioned brands appeared in 64% of responses to the same prompt, so a stable authority layer sits underneath the surface randomness.

Both mechanisms reward one behavior: deep, detailed coverage of your topic area on your own domain. That is Tier 1 of the Prioritized GEO framework we build client programs on, and this study is the clearest outside validation of it we have seen. Owned content that ranks on Google for bottom of the funnel keywords exposes your brand to LLMs 3 ways at once: through the fan out searches, through the training crawl, and through the domain authority signal.

Building that authority looks unglamorous. Meta tag tuning and schema markup play almost no part in it. The work is publishing detailed, substantive articles that answer the real pain points your customers bring to sales calls, week after week, until your domain covers the topic more completely than any competitor’s. The more comprehensively a domain covers a topic area, the more likely ChatGPT treats it as a relevant source for that topic, regardless of which fan out query fires on a given day.

The domain finding also explains why we measure AI visibility topic by topic instead of quoting one blended brand score. Authority lives at the topic level. A brand can dominate one topic in ChatGPT’s answers and be invisible in the neighboring one, and a single averaged percentage hides exactly the signal that tells you where to publish next.

The content ChatGPT cites looks boringly familiar

78% of the sources ChatGPT cited in the study were standard SEO and marketing formats: listicles, product pages, pricing pages, and homepages. The pages that sell products, in other words. Among the 60% of sources that ranked for no fan out query at all, the share barely moved, landing around 74%.

That second number deserves a pause. Even the citations arriving through caching, training data, or hidden searches looked like pages written to rank on Google. Wherever ChatGPT sources its material, it keeps reaching for the same familiar formats.

Grow & Convert point to a Calendly post as the archetype: a listicle ranking 13 appointment scheduling apps, written exactly the way an SEO team writes for a commercial keyword, with Calendly placed at number 1 on its own list. It ranked for one of the observed fan out queries and earned a citation. Put your own product first on your own list pages. Nobody else will.

The 78% figure also settles the format debate for us. FAQ blocks, llms.txt files, key takeaway boxes, and question style headings sit in Tier 3 of our priority pyramid because they help LLMs parse content that has already been found. They do not get content found. When plain listicles and product pages earn 78% of citations, substance and topical coverage decide the outcome, and formatting tweaks decorate it.

What happened when one brand simply published

Grow & Convert include one of their own clients, Toro TMS, as the live example. Toro TMS sells trucking management software, had a young domain, and had done little content marketing before the engagement. The agency produced detailed, product centric articles targeting real Google keywords in the trucking software space, and ChatGPT visibility followed the publishing schedule. Their topic level tracking shows the brand surfacing in AI answers for commercial terms such as “trucking management software” and “bulk hauling software.”

No fan out query reverse engineering. No prompt spreadsheets. Detailed content aimed at Google keywords where the product is a genuine answer, and the AI citations arrived as a byproduct.

The Pro Start Marketing playbook, built on this data

The study translates into 7 working rules for a GEO program:

  1. Drop prompt level targeting. You cannot see the effective prompts, the fan out queries change run to run, and targeting the visible ones pays off less than half the time.
  2. Map your product’s strengths to buying intent Google keywords. Real, trackable keywords with commercial intent, such as “best X for Y,” “X alternatives,” and “X vs Z” comparisons. A topic map built around your product’s genuine strengths beats any prompt list.
  3. Publish detailed, product centric content for those keywords. Cover features, use cases, pricing, limitations, and customer scenarios in depth, because that detail is what lets a model match you to prompts you will never see. Our shortcut for surfacing that detail: interview the sales team, since the objections and use cases they hear daily are the exact specifics an LLM needs before it recommends you.
  4. Rank your product first on your own list content. Calendly does it. Every serious vendor does it.
  5. Run citation outreach on the domains LLMs already cite for your topics. Earn mentions on the specific third party pages appearing in AI answers for your category, instead of blanket Reddit pushes based on generic charts.
  6. Treat on site tactics as experiments, never as strategy. Test llms.txt, schema, and FAQ sections after the owned content and outreach layers are running.
  7. Measure visibility per topic, every month. Track which topics your brand surfaces for in AI answers and which it misses, and let the gaps set the content calendar.

Then hold the schedule. Training data refreshes on model release cycles rather than publishing cycles, so GEO gains lag your content by months and then compound. The lag filters out competitors who quit early, which quietly turns patience into a moat.

The bottom line

The fan out query theory spread because it made AI search feel controllable, like SEO with one extra step. The data reveals a messier machine: part live search, part cache, part memory, part invention. You cannot rank for a memory. You can become the domain the memory is made of.

That is the whole strategy. Publish the deepest, most useful bottom of the funnel content in your category, earn mentions where the models already look, and measure by topic. Building exactly that for B2B and startup brands is what our GEO service at Pro Start Marketing does, and this study is a large part of why we build it that way.

Leave a Comment

Your email address will not be published. Required fields are marked *