The ChatGPT search index cites small sites the same way it cites major publishers, and new data puts numbers behind that claim. A study from research firm Resoneo analyzed 1,249 ChatGPT responses and found no difference in how OpenAI’s in-house index served licensed and unlicensed sites. Small Italian publishers appeared through the same pipeline as Reuters and The Guardian. If you run a modest site and assumed you were locked out of ChatGPT results without a content deal, the evidence says otherwise.
This article breaks down what the research found, how OpenAI’s index actually pulls and stores your page, and the practical steps that help any site earn a citation. You will also learn why one recent change makes this analysis harder to repeat.
Key Takeaways
- No licensing gap. The ChatGPT search index served unlicensed sites the same way it served OpenAI’s licensed partners.
- Labrador does the heavy lifting. OpenAI’s own index (labeled “labrador”) accounted for the large majority of primary search sources.
- Your H1 matters. The index stores roughly 200 characters of a page, usually starting with the H1 heading.
- Small sites can compete. Visibility depends on clear, relevant content, not on a deal with OpenAI.
- The signal went dark. OpenAI removed the field that exposed which pipeline fetched each result, so this exact study is hard to rerun.
What The Data Actually Shows
Resoneo tagged the pipeline behind each cited result and found four sources feeding ChatGPT search: OpenAI’s own index (labrador), plus three external fetchers. Labrador dominated, handling the vast majority of primary search sources. The headline finding is simpler than the plumbing: licensing status did not change how a page was treated. A small publisher with no OpenAI agreement carried the same labrador designation as a national newspaper.
That matters because it removes a common excuse. You do not need an enterprise contract to show up. You need content the index can read and trust, which is the same principle behind getting cited by AI search with specific writing.

How Does The ChatGPT Search Index Pull Sources?
ChatGPT search does not rely on a single fetcher. The research observed four labeled sources, with OpenAI’s in-house index carrying the clear majority and external pipelines filling the rest. Two points stand out for site owners.
- The in-house index leads. Because labrador handled most primary sources, being present in OpenAI’s own crawl gives you the best odds of selection.
- External pipelines still contribute. The remaining fetchers mean a page can surface even when it is not yet in the in-house index, so broad crawlability helps.
The practical read is to stay easy to crawl and quote everywhere, not to chase one pipeline. That mindset connects to the wider shift covered in what ChatGPT means for traditional search traffic.
What Does The Index Store Of Your Page?
The index does not keep your whole article. Resoneo found it captures around 200 characters of a page, and those characters usually begin with the H1. Across 463 pages analyzed, 83.6% of stored snippets included the H1 heading.
That single detail changes how you write the top of a page. If the model mostly sees your headline and the first line or two, those elements have to state clearly what the page delivers. A vague or clever H1 wastes the exact space the index remembers. You can read the full methodology in Resoneo’s analysis of what ChatGPT retrieves. Whether that memory sticks over time is its own question, one we explore in whether answer engines remember your brand.

How To Earn A Citation From The ChatGPT Search Index
Since the field is open to sites of any size, the work is straightforward. Focus on the signals the index actually reads.
- Write a precise H1. Lead with a clear, specific headline that names the topic, because that is what the index most often stores.
- Front-load the answer. State the core point in the first sentences, inside the roughly 200 characters the index keeps.
- Stay crawlable. Keep pages accessible to bots and avoid blocking the crawlers that feed AI answers, an issue tied to the llms.txt question for 2026.
- Earn topical trust. Cover a subject thoroughly and accurately so the index has a strong reason to select you over a rival.
Why The Signal Just Went Dark
There is a catch for anyone hoping to verify this themselves. On 21 July, OpenAI removed the field that revealed which pipeline fetched each result. The labels that made Resoneo’s study possible no longer appear in the traffic, so this precise breakdown is hard to reproduce now.
The disappearance of the signal does not change the takeaway. The behavior it documented, equal treatment of small and large sites, still stands as the best current evidence. It simply means future analysis will lean more on citation outcomes than on internal pipeline labels. You can see how the feature presents sources for yourself in OpenAI’s own overview of ChatGPT search.
Frequently Asked Questions
Does ChatGPT cite small websites?
Yes. Resoneo’s analysis of 1,249 ChatGPT responses found the ChatGPT search index served small, unlicensed sites the same way it served large licensed publishers. Site size and licensing status did not change how a page was treated.
What is labrador in ChatGPT search?
Labrador is the label researchers observed for OpenAI’s in-house search index. It accounted for the large majority of primary search sources in the study, ahead of the external fetchers that made up the rest.
What does the ChatGPT search index store about my page?
It stores roughly 200 characters of the page, and those characters usually begin with the H1 heading. In the study, 83.6% of stored snippets included the H1, so a clear headline and opening line matter most.
Do I need a content deal with OpenAI to appear in ChatGPT?
No. The data showed no advantage for licensed partners in how the index served pages. Clear, crawlable, trustworthy content is what earns selection, not a commercial agreement.
Can I still see which pipeline ChatGPT used?
Not easily. On 21 July, OpenAI removed the field that exposed the pipeline behind each result, so the labels that powered this study are no longer visible in current traffic.
Size Is Not The Barrier
The evidence is encouraging for small publishers: the ChatGPT search index judges pages on clarity and relevance, not on the size of your logo or the existence of a licensing deal. Write a sharp H1, answer the question in your opening lines, keep your pages crawlable, and cover your topic with real depth. Do that, and you give OpenAI’s index every reason to pull your page into the answer, right alongside the household names.
