ChatGPT search index may be fairer to small sites than expected
New data suggests ChatGPT's in-house search index does not sideline small sites. For GEO teams, retrievability matters more than publisher deals.
Small sites do not appear to be shut out of ChatGPT Search by default. A new Resoneo analysis of 1,249 answers captured in July suggests OpenAI's in-house index served partner and non-partner pages in the same way on free accounts, with no visible difference in format, snippet length, or freshness. That changes the practical GEO question. For many brands, the issue is no longer, "Do we need a publisher deal to get in?" It is, "Can ChatGPT retrieve, understand, and reuse our page when it decides to search?"
That is good news, but only up to a point. The same research also suggests paid "thinking" mode relies much more heavily on Google scraping, which means visibility can still change depending on account type, search mode, and query class. Small sites may have a path in, but they are not competing in one clean, stable search system.
The main finding is simple: in this dataset, partner status did not change how pages were served inside the pipeline tagged "labrador," which Resoneo interprets as OpenAI's in-house index. Resoneo reviewed 1,249 ChatGPT answers and found that licensed and unlicensed pages looked the same on the dimensions it measured: format, length, and freshness.
That matters because it cuts against a popular assumption: that ChatGPT Search is mostly a closed loop for major publishers with OpenAI deals. It also lines up with a July correction from SEO researcher Suganthan Mohanadasan, who had initially read the same pipeline as a kind of licensed tier and later said that conclusion went too far after broader testing across accounts.
There is an important limit here. This is reverse-engineered evidence based on network traffic, not a formal product specification. Resoneo also says the pipeline labels stopped appearing around July 21, which means this window of visibility may not stay open forever. Still, as a directional signal, it is strong enough to matter.
It matters because it removes one lazy excuse. If your site is not showing up in ChatGPT answers, a missing publisher deal is no longer the most obvious explanation. A small brand, local business, niche SaaS company, or specialist blog may already be eligible for inclusion in the system that powers a large share of free-account results.
Take a simple example. Imagine a regional accounting software vendor with a well-structured comparison page called "Best bookkeeping software for US nonprofits." Under the old assumption, that team might believe ChatGPT will never surface them because they are not Reuters or The Wall Street Journal. These new findings suggest the real problem may be something more fixable: weak page structure, generic copy, poor snippet composition, or the fact that ChatGPT searched a slightly different sub-question and landed elsewhere.
This is also consistent with a broader pattern we covered in why AI search traffic does not follow organic search rules. AI discovery does not simply copy Google's winner list. A page can have modest organic visibility and still become useful in AI answers if it is clear, specific, and easy to reuse.
But there is no reason to get romantic about this. Inclusion is not the same as prominence. A small site may be in the pool and still never earn the citation because the answer surface is tight, the model prefers a stronger entity, or the query path never reaches the page. As we explored in ChatGPT Search is citing fewer domains, the system can be selective even when it is not exclusionary.
According to Resoneo's page-level review, the in-house index appears to store a title and a short snippet of about 200 characters. That is a small window, which makes the top of your page far more important than many teams realize.
Resoneo compared 534 pages cited by ChatGPT with the snippets stored in the index. Of the 463 pages that had an H1, 387 snippets included that H1, or 83.6%. The median H1 length was 51 characters, which leaves roughly 150 characters for body content after the heading. That means your first lines are not decorative. They are part of the retrievable summary.
The details here are more useful than they look. A section kicker appeared before the H1 on 29% of pages and consumed about 18 characters. A publication date appeared on 11% of pages and took 25 characters. The alt text of the first image appeared on 9% of pages and could consume around 50 characters on its own. In other words, template clutter can eat a meaningful share of the snippet before your actual answer even begins.
Picture a page that opens like this: "Industry News | August 2026 | Hero image alt text | Top 10 CRM tools for law firms." By the time the system reaches the first useful sentence, much of the snippet budget is already gone. Compare that with a page that starts with a direct H1 and a first sentence that defines the page's value immediately. The second page gives the model a cleaner summary to work with.
Another notable detail: about one in seven pages in the sample had no H1 at all. In those cases, the snippet reportedly started with whatever subheading the template exposed first. That is not a minor markup issue. It means the system may be summarizing your page from a structural accident rather than your intended headline.
Because they seem to rely on different retrieval mixes. In Resoneo's data, free accounts leaned heavily on the in-house index, while paid "thinking" mode leaned much more on Google scraping. If you test only one account type, you may mistake one retrieval path for the whole product.
That distinction matters operationally. A page that performs well for free-user informational queries may not win the same way in a more research-heavy, multi-step mode that reaches outward through Google more often. The reverse can also happen. A page with strong search visibility and rich SERP signals may appear more often in thinking mode than in the free experience.
Imagine a B2B cybersecurity company testing the query "best incident response platforms for mid-market teams." On a free account, ChatGPT might resolve the question through its internal index and pull a compact snippet from a clear category page. In thinking mode, the system may branch through Google-sourced comparisons, reviews, or analyst pages before deciding which domains to cite. Same topic, different retrieval path, different winners.
This is why single-screenshot GEO is a dead end. If you only test one prompt on one account, you are not measuring visibility. You are recording a moment inside one interface condition. That is useful for debugging, but weak as a strategy.
The useful lesson here is not "small sites are fine." It is that eligibility and visibility are two different layers. A page can be eligible for ChatGPT's index and still fail to earn a mention because the wrong query branch was triggered, the snippet is weak, or a competitor owns the framing around the topic. That is exactly why teams need measurement beyond a yes-or-no ranking check.
BotRank's AI Visibility and Source Analysis features are useful in this kind of situation because they let teams run recurring prompts across models, compare outcomes over time, and inspect which pages and sources actually support the answer. If ChatGPT can find your page but still cites a reseller, summarizes you badly, or mentions a competitor first, you still have a GEO problem. The fix starts with seeing the retrieval path clearly, not guessing from one lucky result.
They should optimize for retrievability. If the index stores a title and roughly 200 characters, and if different ChatGPT modes use different source pipelines, then the job is to make important pages easy to select, easy to summarize, and easy to trust.
A practical example makes this concrete. Suppose you run a small ecommerce site selling specialty espresso grinders. Your product category page might already be indexable, but if it opens with a generic slogan, hides key specs behind tabs, and spends the first paragraph talking about your founder story, it gives the model little to reuse. Rewrite the H1, move the comparative value up, expose the essential specs in visible HTML, and the page becomes easier to pull into a product answer.
This approach works especially well for local, product, and comparison intent. It is less decisive for pure breaking news, where source freshness and upstream discovery paths can dominate. That nuance matters. There is no universal ChatGPT optimization move.
It does not prove that publisher deals are useless. It only suggests they are not the simple gate to free-tier inclusion that many people assumed. Resoneo's work focused on how pages were stored and served, not whether deals improve citation frequency, ranking priority, freshness latency, or legal distribution rights.
That distinction matters. A large publisher may still value a direct feed relationship for reasons that have nothing to do with whether a smaller site can appear in free-account results. A feed can affect ingestion mechanics even if the visible snippet looks similar. But for most marketing teams, the more urgent conclusion is simpler: do not explain weak ChatGPT visibility by blaming closed-door partnerships before you have tested your own retrievability.
A small legal blog, for example, should not assume it loses every answer because it lacks a deal. It may be losing because its articles bury the answer, use weak H1s, or fail to cover the sub-questions that the model actually branches into. That is harder news, but also better news, because it gives the team something they can change.
Sometimes, yes, but not automatically. The study suggests small sites are not excluded from the in-house index by default, yet citation competition is still tight and query-dependent.
No. The better rule is to audit whether those elements add value or just consume snippet space before the core answer appears.
Possibly, but this dataset did not test citation lift from deals. It only suggests that partner status did not change the basic way pages were served in the in-house index sample.
Test the same high-value prompts across different ChatGPT experiences, then inspect which pages and sources are actually cited. If your own page is missing, start by fixing structure, opening copy, and technical accessibility before assuming the system is closed to you.
The big takeaway is refreshingly blunt: small sites are not necessarily locked out of ChatGPT Search, but they are still easy to ignore. If you want better odds, stop treating access as the whole game. Measure the retrieval path, tighten the page opening, and make your best pages easier for the model to trust and reuse.