ChatGPT search index may be fairer to small sites than expected

Published:
August 14, 2026

Small sites do not appear to be shut out of ChatGPT Search by default. A new Resoneo analysis of 1,249 answers captured in July suggests OpenAI's in-house index served partner and non-partner pages in the same way on free accounts, with no visible difference in format, snippet length, or freshness. That changes the practical GEO question. For many brands, the issue is no longer, "Do we need a publisher deal to get in?" It is, "Can ChatGPT retrieve, understand, and reuse our page when it decides to search?"

That is good news, but only up to a point. The same research also suggests paid "thinking" mode relies much more heavily on Google scraping, which means visibility can still change depending on account type, search mode, and query class. Small sites may have a path in, but they are not competing in one clean, stable search system.

What did the new data actually show?

The main finding is simple: in this dataset, partner status did not change how pages were served inside the pipeline tagged "labrador," which Resoneo interprets as OpenAI's in-house index. Resoneo reviewed 1,249 ChatGPT answers and found that licensed and unlicensed pages looked the same on the dimensions it measured: format, length, and freshness.

  • On free accounts, the in-house index appeared to power most results.
  • Questions with settled answers, local businesses, and products were especially likely to come from that index.
  • News queries were more mixed, with results split more evenly between the in-house index and Google scraping.
  • In paid "thinking" mode, Google scraping accounted for about 75% of the 16,407 search results Resoneo recorded, while the in-house index accounted for about 24%.

That matters because it cuts against a popular assumption: that ChatGPT Search is mostly a closed loop for major publishers with OpenAI deals. It also lines up with a July correction from SEO researcher Suganthan Mohanadasan, who had initially read the same pipeline as a kind of licensed tier and later said that conclusion went too far after broader testing across accounts.

There is an important limit here. This is reverse-engineered evidence based on network traffic, not a formal product specification. Resoneo also says the pipeline labels stopped appearing around July 21, which means this window of visibility may not stay open forever. Still, as a directional signal, it is strong enough to matter.

Why does this matter for small sites?

It matters because it removes one lazy excuse. If your site is not showing up in ChatGPT answers, a missing publisher deal is no longer the most obvious explanation. A small brand, local business, niche SaaS company, or specialist blog may already be eligible for inclusion in the system that powers a large share of free-account results.

Take a simple example. Imagine a regional accounting software vendor with a well-structured comparison page called "Best bookkeeping software for US nonprofits." Under the old assumption, that team might believe ChatGPT will never surface them because they are not Reuters or The Wall Street Journal. These new findings suggest the real problem may be something more fixable: weak page structure, generic copy, poor snippet composition, or the fact that ChatGPT searched a slightly different sub-question and landed elsewhere.

This is also consistent with a broader pattern we covered in why AI search traffic does not follow organic search rules. AI discovery does not simply copy Google's winner list. A page can have modest organic visibility and still become useful in AI answers if it is clear, specific, and easy to reuse.

But there is no reason to get romantic about this. Inclusion is not the same as prominence. A small site may be in the pool and still never earn the citation because the answer surface is tight, the model prefers a stronger entity, or the query path never reaches the page. As we explored in ChatGPT Search is citing fewer domains, the system can be selective even when it is not exclusionary.

What does ChatGPT appear to keep from your page?

According to Resoneo's page-level review, the in-house index appears to store a title and a short snippet of about 200 characters. That is a small window, which makes the top of your page far more important than many teams realize.

Resoneo compared 534 pages cited by ChatGPT with the snippets stored in the index. Of the 463 pages that had an H1, 387 snippets included that H1, or 83.6%. The median H1 length was 51 characters, which leaves roughly 150 characters for body content after the heading. That means your first lines are not decorative. They are part of the retrievable summary.

The details here are more useful than they look. A section kicker appeared before the H1 on 29% of pages and consumed about 18 characters. A publication date appeared on 11% of pages and took 25 characters. The alt text of the first image appeared on 9% of pages and could consume around 50 characters on its own. In other words, template clutter can eat a meaningful share of the snippet before your actual answer even begins.

Picture a page that opens like this: "Industry News | August 2026 | Hero image alt text | Top 10 CRM tools for law firms." By the time the system reaches the first useful sentence, much of the snippet budget is already gone. Compare that with a page that starts with a direct H1 and a first sentence that defines the page's value immediately. The second page gives the model a cleaner summary to work with.

Another notable detail: about one in seven pages in the sample had no H1 at all. In those cases, the snippet reportedly started with whatever subheading the template exposed first. That is not a minor markup issue. It means the system may be summarizing your page from a structural accident rather than your intended headline.

Why do free and paid ChatGPT experiences tell different stories?

Because they seem to rely on different retrieval mixes. In Resoneo's data, free accounts leaned heavily on the in-house index, while paid "thinking" mode leaned much more on Google scraping. If you test only one account type, you may mistake one retrieval path for the whole product.

That distinction matters operationally. A page that performs well for free-user informational queries may not win the same way in a more research-heavy, multi-step mode that reaches outward through Google more often. The reverse can also happen. A page with strong search visibility and rich SERP signals may appear more often in thinking mode than in the free experience.

Imagine a B2B cybersecurity company testing the query "best incident response platforms for mid-market teams." On a free account, ChatGPT might resolve the question through its internal index and pull a compact snippet from a clear category page. In thinking mode, the system may branch through Google-sourced comparisons, reviews, or analyst pages before deciding which domains to cite. Same topic, different retrieval path, different winners.

This is why single-screenshot GEO is a dead end. If you only test one prompt on one account, you are not measuring visibility. You are recording a moment inside one interface condition. That is useful for debugging, but weak as a strategy.

BotRank's Take

The useful lesson here is not "small sites are fine." It is that eligibility and visibility are two different layers. A page can be eligible for ChatGPT's index and still fail to earn a mention because the wrong query branch was triggered, the snippet is weak, or a competitor owns the framing around the topic. That is exactly why teams need measurement beyond a yes-or-no ranking check.

BotRank's AI Visibility and Source Analysis features are useful in this kind of situation because they let teams run recurring prompts across models, compare outcomes over time, and inspect which pages and sources actually support the answer. If ChatGPT can find your page but still cites a reseller, summarizes you badly, or mentions a competitor first, you still have a GEO problem. The fix starts with seeing the retrieval path clearly, not guessing from one lucky result.

What should teams optimize if publisher deals are not the main gate?

They should optimize for retrievability. If the index stores a title and roughly 200 characters, and if different ChatGPT modes use different source pipelines, then the job is to make important pages easy to select, easy to summarize, and easy to trust.

  • Make the first 200 characters do real work. Your H1 and opening sentence should explain the page clearly. A line like "Best CRM for law firms: 7 tools compared by price, migration support, and compliance" is stronger than a vague brand-led opener.
  • Reduce low-value text before the answer. If your template puts category labels, dates, badges, or decorative alt text above the first useful sentence, that text can consume snippet space. This does not mean every date should disappear. It means every character above the fold should justify itself.
  • Test multiple experience types. Save prompts for free-account style checks, deeper reasoning checks, local intent, and product intent. A reusable library in Prompts Studio helps teams compare the same query set instead of improvising every week.
  • Audit the technical layer, not just the copy. If your page depends on JavaScript for core content, exposes weak heading structure, or blocks the wrong crawler, your content can be eligible in theory and absent in practice. That is where GEO Page Analysis becomes useful, especially when paired with the signals discussed in what 68.9 million AI crawler visits tell us about AI search visibility.
  • Track source ownership, not just brand mentions. If the answer cites your distributor, a review site, or a directory instead of your own page, you may still be visible but not in control. Source-level analysis helps separate nominal presence from real brand ownership.

A practical example makes this concrete. Suppose you run a small ecommerce site selling specialty espresso grinders. Your product category page might already be indexable, but if it opens with a generic slogan, hides key specs behind tabs, and spends the first paragraph talking about your founder story, it gives the model little to reuse. Rewrite the H1, move the comparative value up, expose the essential specs in visible HTML, and the page becomes easier to pull into a product answer.

This approach works especially well for local, product, and comparison intent. It is less decisive for pure breaking news, where source freshness and upstream discovery paths can dominate. That nuance matters. There is no universal ChatGPT optimization move.

What does this mean for publishers and OpenAI deals?

It does not prove that publisher deals are useless. It only suggests they are not the simple gate to free-tier inclusion that many people assumed. Resoneo's work focused on how pages were stored and served, not whether deals improve citation frequency, ranking priority, freshness latency, or legal distribution rights.

That distinction matters. A large publisher may still value a direct feed relationship for reasons that have nothing to do with whether a smaller site can appear in free-account results. A feed can affect ingestion mechanics even if the visible snippet looks similar. But for most marketing teams, the more urgent conclusion is simpler: do not explain weak ChatGPT visibility by blaming closed-door partnerships before you have tested your own retrievability.

A small legal blog, for example, should not assume it loses every answer because it lacks a deal. It may be losing because its articles bury the answer, use weak H1s, or fail to cover the sub-questions that the model actually branches into. That is harder news, but also better news, because it gives the team something they can change.

FAQ

Does this mean small sites can beat major publishers in ChatGPT Search?

Sometimes, yes, but not automatically. The study suggests small sites are not excluded from the in-house index by default, yet citation competition is still tight and query-dependent.

Should teams remove dates, kickers, or image text from the top of every page?

No. The better rule is to audit whether those elements add value or just consume snippet space before the core answer appears.

Does a content deal with OpenAI still matter?

Possibly, but this dataset did not test citation lift from deals. It only suggests that partner status did not change the basic way pages were served in the in-house index sample.

What is the first GEO test a small brand should run now?

Test the same high-value prompts across different ChatGPT experiences, then inspect which pages and sources are actually cited. If your own page is missing, start by fixing structure, opening copy, and technical accessibility before assuming the system is closed to you.

The big takeaway is refreshingly blunt: small sites are not necessarily locked out of ChatGPT Search, but they are still easy to ignore. If you want better odds, stop treating access as the whole game. Measure the retrieval path, tighten the page opening, and make your best pages easier for the model to trust and reuse.

Florian Chapelier

About the author
AI Search & GEO expert

After nearly 15 years in digital strategy on the client side (including 10 years at Olympique Lyonnais, where he was notably in charge of SEO).
Florian co-founded BotRank.ai in 2025, the GEO (Generative Engine Optimization) tool used by more than 2,500 companies to manage their visibility in AI-generated search results. He writes regularly about GEO and AI Search.