ChatGPT Retrieval Rarely Opens Source Pages

·

·

·

4–7 minutes
ChatGPT Retrieval Funnel - From Pages to Answers

New research into ChatGPT retrieval shows that getting discovered by ChatGPT is only the first step toward earning visible AI search exposure. RESONEO traced more than 1,200 real answers and found a steep funnel between URLs retrieved, URLs cited and pages actually opened by the model.

What did the research find?

RESONEO reverse-engineered ChatGPT’s web-retrieval process using 1,249 answers captured in real conditions.

Its headline numbers were striking:

  • About 58,000 URLs were retrieved.
  • Around 7,600 were promoted as sources.
  • About 5,000 appeared as citations in answer text.
  • Only around 760 were opened and read by the model.

The research was carried out during July 2026 and subsequently updated in August as ChatGPT’s retrieval behaviour changed.

Search Engine Land published a detailed analysis of the work on August 17.

The biggest lesson is simple.

Being found is not the same as being cited.

And being cited is not the same as ChatGPT fully opening the page.

How does ChatGPT find pages?

When ChatGPT decides that a question needs current web information, it can retrieve several candidate URLs.

RESONEO found that the retrieval system often works with compact information containing elements such as:

  • Page title
  • URL
  • A short snippet

The July dataset suggested that roughly 200 characters of snippet text can represent a page during much of the selection process.

That means ChatGPT may decide whether a page is useful before it ever reads the complete article.

This is an important difference from the simple idea that an AI engine always visits five pages, reads them completely and writes an answer.

The actual process can contain a much larger candidate pool.

Most candidates do not receive the same level of processing.

Why do snippets suddenly matter?

Search marketers have spent years optimizing titles and descriptions because they influence discovery and clicks.

AI search adds another reason to care about how clearly a page can be represented in a small amount of text.

Imagine two pages about the same topic.

One begins with:

Our innovative solution brings the future of digital transformation to modern businesses.

The other begins with:

Customer churn rate is the percentage of customers who stop using a product during a defined period.

The second passage communicates a useful answer immediately.

That type of clarity gives retrieval systems more meaningful information to work with.

This does not mean marketers should mechanically optimize every first paragraph for a chatbot.

It means vague introductions have become even less useful.

What gets opened is more likely to be cited?

RESONEO found a strong difference between URLs that were actually opened and those that were merely retrieved.

Its research found that opened pages were cited much more frequently than pages that stayed only in the retrieval layer.

That creates a useful AI visibility funnel:

Retrieved → considered → opened → cited → clicked

A brand can lose at any stage.

Traditional AI visibility tools often measure only the final answer.

That tells marketers which brand received the citation.

It does not necessarily show which competitors entered the candidate pool and were rejected.

Understanding that difference may become increasingly important for LLM visibility.

Retrieval also differs by ChatGPT mode

RESONEO’s August update showed that ChatGPT does not use one fixed source pipeline.

In its tests, Free Think pulled about 35.3 URLs per conversation, while paid thinking pulled about 35.1.

However, the source mix differed sharply.

Free Think relied heavily on what RESONEO identifies as OpenAI’s own Labrador index, while paid thinking relied much more heavily on Google-derived results in the tested configuration.

The research also found major differences between instant and thinking behaviour.

This creates a serious measurement challenge.

A brand may be visible for the same prompt on one ChatGPT configuration but absent on another.

Marketers should therefore avoid treating one prompt test as a permanent ranking.

Why should marketers care?

Many AI SEO programmes currently focus on one question:

Was my website cited?

The research suggests teams need a wider model.

A page may fail because:

  1. ChatGPT cannot access it.
  2. It does not enter the retrieval pool.
  3. Its snippet is weak.
  4. Another source is considered more useful.
  5. The page is retrieved but not cited.
  6. The brand is mentioned without a citation.

Each problem needs a different solution.

A technical crawl issue cannot be fixed by rewriting a paragraph.

A weak answer cannot be fixed only by changing robots.txt.

AI visibility requires both technical access and useful content.

What should content teams change?

Answer important questions early

Do not hide the useful answer beneath 300 words of background.

Give readers the direct answer first.

Then explain it.

Make sections self-contained

A heading and the paragraphs below it should make sense even when extracted from the complete article.

Use precise language

Prefer:

Product X supports Windows 11 and macOS 15.

over:

Product X works across modern platforms.

Specific information is easier to understand and verify.

Add original evidence

Useful evidence can include:

  • First-party research
  • Testing
  • Statistics
  • Expert commentary
  • Product specifications

Generic summaries are easier for every competitor to reproduce.

What should technical teams check?

OpenAI’s official publisher guidance says public websites can appear in ChatGPT search and recommends allowing OAI-SearchBot when publishers want their content to be discovered, summarized and cited.

Teams should therefore review:

  • robots.txt
  • CDN bot rules
  • WAF settings
  • indexability
  • canonical URLs
  • rendered content

Do not confuse OAI-SearchBot with GPTBot.

OpenAI separates search discovery from controls associated with potential model training.

That distinction matters when creating crawler policies.

What are the limitations?

This is reverse-engineering research.

It is not official OpenAI documentation describing a fixed ranking algorithm.

ChatGPT can change:

  • Models
  • Search providers
  • Retrieval depth
  • Caching
  • Citation interfaces
  • Tool routing

RESONEO itself updated parts of its July conclusions after behaviour changed in August.

That is exactly why marketers should avoid turning one finding into a permanent optimization rule.

The useful insight is the architecture of the problem.

Retrieval and citation are separate stages.

What happens next?

Build AI visibility reporting around several layers.

Track:

  1. Crawler accessibility.
  2. Important prompt coverage.
  3. Brand mentions.
  4. Citation frequency.
  5. Cited URLs.
  6. ChatGPT referral traffic.
  7. Leads or revenue from those visits.

OpenAI automatically adds utm_source=chatgpt.com to ChatGPT search referral URLs, which can help publishers analyse direct visits.

Most importantly, stop thinking of AI optimization as simply “ranking in ChatGPT.”

There is no single visible ten-result list.

A page must first enter the right information pool and then survive several selection stages.

Clear information, strong evidence and reliable technical access give it a better chance of doing that.


Vatsal Makhija

Meet the Writer

Hi, I’m Vatsal. The SEO chief behind Get Search Engine, a small business SEO specialist who’s worked on hands-on campaigns for global brands and scrappy local businesses alike.


Free SEO AUDIT!

Smart brands are fixing SEO gaps before peak season hits. Are you?


Prefer Direct Contact?

Getsearchengine.com
📍 Business Hours: Monday – Friday | 9 AM – 6 PM IST
For urgent queries, email us at:
vatsalmakhija.work@gmail.com

Message Us

First Name
Last Name
Email
Message
The form has been submitted successfully!
There has been some error while submitting the form. Please verify all form fields again.