AI search technical signals most SEO teams still miss

Published:
September 3, 2026
Update:
September 3, 2026

Most SEO teams are still optimizing AI search as if access were the whole game. It is not. In a June 12, 2026 audit of 50 major websites across retail, SaaS, travel, publishing, and finance, retrievability scored an average 74.4%, but attribution and meaning fell to 38.5%, and agent transaction and discovery collapsed to 2.1%. That gap matters because a page AI can fetch is not automatically a page AI can interpret correctly, trust confidently, or act on.

For brands doing Generative Engine Optimization (GEO), the takeaway is blunt: citations are only layer one. If you want AI systems to quote the right facts, describe your brand accurately, and eventually transact on your site, you need to optimize the technical signals behind understanding, not just crawling.

  • Most large sites are decent at AI access, but weak at AI understanding.
  • Structured data and content-use signals are still missing on too many important sites.
  • Agent-ready protocols are barely implemented across major brands.
  • A low AI-readiness score is not always bad if it reflects a deliberate business choice.
  • For GEO, the next gains come from meaning, control, and execution, not from citations alone.

Why are most SEO teams still optimizing the wrong layer?

Because classic SEO trained us to equate visibility with discoverability. That logic still matters, but AI search adds two extra jobs: machine understanding and machine action.

The audit used a three-layer framework built around 27 elements. Layer 1, retrievability, asks whether AI can fetch and parse the page. Layer 2, attribution and meaning, asks whether AI can tell what the page means and who it belongs to. Layer 3, agent transaction and discovery, asks whether AI can securely interact with site capabilities such as forms, APIs, or transactions.

The scoring only covered 12 established signals, rated on a 0/1/2 scale, while emerging and frontier protocols were tracked separately. That matters because it keeps the benchmark grounded in what teams can act on now, rather than inflating the picture with speculative standards.

That framing lines up with something we have argued before: AI visibility is a three-layer problem, not a content volume problem. Publishing more pages can help Layer 1. It does almost nothing if the real failure is that a model cannot distinguish your full-price product from a promotional bundle, or your company entity from a similarly named topic.

The examples in the cohort make the point clearly. Airbnb led the scoring at 79.2%. Amazon sat far lower at 29.2%, but that does not automatically mean Amazon is bad at AI. In this case, deliberate crawler controls appear to be part of the strategy. A low score can reflect a business decision. A high score usually reflects a deliberate one too.

Which technical signals does AI search actually care about?

The short answer is that AI search still cares about many technical SEO basics, but it puts extra weight on whether those basics make meaning explicit. The easiest way to see it is layer by layer.

LayerWhat it checksAverage scoreWhat it means
RetrievabilityRobots rules, accessibility tree, ARIA labels, semantic HTML, clean delivery, machine-usable forms, sitemap signals74.4%Most teams have the basics partly covered
Attribution and meaningJSON-LD, semantic richness, content-use policy38.5%AI can read many pages, but often lacks confidence in what they mean
Agent transaction and discoveryOAuth metadata and agent-facing access protocols2.1%Most sites are not ready for AI to do useful actions on the user's behalf

Layer 1 includes familiar work: robots directives, accessibility tree integrity, ARIA labeling, semantic HTML, server-rendered delivery, machine-usable forms, and sitemap declaration. Most enterprise teams are at least partly here already, which explains why the average was relatively healthy. This is still important work. If AI cannot fetch or parse the page cleanly, nothing else matters.

But Layer 1 is also where many SEO teams stop. That is the mistake. A well-crawled page can still be a semantically muddy page. If a site ships a clean DOM but leaves product attributes, brand ownership, review context, or FAQ intent ambiguous, a language model still has to infer too much.

Layer 2 is where the drop starts. The audit found JSON-LD on 35 of 50 homepages, or 70%, which sounds decent until you flip the number around: nearly one-third of major sites were still giving machines no structured semantic help at all on the homepage. That is not a minor gap. It is the difference between AI reading words and AI knowing what those words refer to.

Layer 3 is the real wake-up call. Of 48 sites where endpoint testing was possible, 46 scored zero. Only Airbnb and Vercel showed OAuth authorization server metadata, and even they lacked protected resource metadata. In other words, most sites are not prepared for AI to do anything useful on the user's behalf yet.

Why is attribution and meaning the biggest missed opportunity?

Because misunderstanding is more dangerous than invisibility. If a model cannot tell which figure on the page is the base price, which is the sale price, which organization owns the page, or which product the review refers to, it will still try to answer. That is where confident but flawed AI responses start.

Schema markup is the obvious piece here, but it is not the only one. Structured data helps define products, brands, authors, prices, FAQs, and other entities in machine-readable terms. It does not guarantee citations. It does reduce ambiguity. That distinction matters.

The audit also found that only five of the 50 sites had implemented Cloudflare's Content Signals Policy. That matters because it lets site owners express more specific preferences than the blunt allow-all or block-all logic many teams still use in robots.txt. A brand can signal one preference for search use, another for training, and leave other cases unspecified. That is a smarter starting point for AI governance.

This is exactly why technical SEO audits now need an AI-readiness layer. A crawler-access audit can tell you whether the page is reachable. It cannot tell you whether the machine is likely to assign the right meaning to your claims, your prices, your authorship, or your brand entity.

Think about a travel brand with dynamic pricing, loyalty discounts, and multiple regional versions of the same offer. A human can usually infer the correct context from surrounding design cues. An AI system working from raw HTML, rendered DOM, or extracted passages cannot rely on visual intuition. It needs clean structure, explicit labels, and consistent entities.

The same goes for brand ambiguity. If your company name overlaps with a broader concept, a person can use context to resolve it. A model can too, but only if you make the context machine-readable enough to trust. This is where a surprising amount of GEO performance is won or lost.

What does the near-empty agent layer mean in practice?

It means most brands are still treating AI as a reader, not a doer. That is fine for some business models today. It will not be fine for all of them tomorrow.

Agent transaction and discovery is the layer that lets AI move from summarizing a site to interacting with it. In the audit, this layer averaged just 2.1%. That is not because every SEO team is asleep. It is because the standards are newer, the commercial implications are bigger, and the ownership usually sits across SEO, engineering, security, product, and legal.

For publishers, deliberate resistance can make sense. News organizations in the cohort, including major media brands, block many AI bots because their economics still depend on direct visits. For retail, SaaS, travel, and finance, the calculus is different. If customers increasingly ask assistants to compare vendors, pull live availability, or complete a purchase, the brands that expose clean, secure machine pathways will have an edge.

That is why Google's Universal Commerce Protocol guide matters even if you are not ready to implement it yet. It signals where commerce interfaces are heading: structured, machine-readable workflows that support discovery, selection, checkout, and follow-up inside AI-driven experiences. The same pattern is appearing around APIs, authentication, and tool access more broadly.

The important nuance is timing. Layer 3 is the area to watch, not necessarily the area every team should prioritize first. If your core pages are still ambiguous, or your robots rules are accidental, chasing agent protocols before fixing the foundation is like wiring a smart checkout into a store with no shelf labels.

BotRank's Take

The biggest mistake teams make with AI search is measuring only appearance, not usability. A brand can be mentioned often and still lose. The reasons are familiar once you inspect the underlying pages: missing structured data, weak template semantics, inconsistent naming between regions, or source pages that models cite even though they barely mention the brand. That is why technical AI readiness should be monitored like an ongoing system, not a one-off checklist.

GEO Page Analysis is useful in this exact context because it tracks the pages you actually care about, audits them repeatedly, and shows score history instead of a single snapshot. It looks at technical accessibility, robots.txt and llms.txt handling, machine-readable structure, and the issues that block LLM discovery or interpretation. Pair that with Recommendations, and the audit stops being a pile of observations and becomes a prioritized worklist. That is the difference between knowing AI search is changing and actually adapting your site to it.

What should SEO teams audit first?

Start with the signals that remove ambiguity fast. The audit results make clear that most teams do not need another generic AI SEO checklist. They need a sequence.

  • Make crawler access a decision, not a default. The audit found that 29 of 50 sites had made no deliberate decision on AI agent access. Review robots.txt, bot-specific directives, and security policies together. Airbnb, eBay, and Tripadvisor illustrate what deliberate policy looks like, even when the choices differ.
  • Compare raw HTML with the rendered page. If price, availability, or primary claims only appear late in client-side rendering, some AI systems may miss or misread them. This matters especially on JavaScript-heavy product pages and SaaS templates.
  • Validate structured data on core templates. Homepage, product, article, FAQ, organization, and breadcrumb markup should be complete and consistent. Nearly one-third of the audited homepages still lacked JSON-LD, which is a large avoidable gap for enterprise sites.
  • Audit how models describe you, not just whether they mention you. If different engines summarize your brand differently, AI Visibility tracking helps you compare that drift across models, while Source Analysis helps you inspect which pages are actually shaping the answer.
  • Treat llms.txt as a helper, not a shortcut. The audit found 11 of 50 sites had published one. That is a useful signal, but it is not a substitute for architecture. Expedia is the cautionary example: a polished llms.txt did not compensate for a weak overall technical setup, missing JSON-LD, or limited server-side delivery.
  • Watch the agent layer where intent is highest. For ecommerce and booked services, start mapping which actions an assistant may need in the next 12 months: price lookup, inventory, sign-in, booking, support, or reorder. That exercise alone often reveals whether engineering needs cleaner APIs, clearer authorization metadata, or both.

If you want the strategic version of this exercise, our piece on why AI visibility is an operations problem before it is an SEO problem is worth reading next. The technical layer is real, but it rarely lives inside SEO alone.

FAQ

Does better schema guarantee more AI citations?

No. Structured data improves machine understanding, but it does not force a model to cite you. It helps AI interpret the page correctly, which is a prerequisite for reliable citation, not a guarantee of it.

Should every site allow AI crawlers?

No. The right choice depends on the business model. Some publishers and large platforms deliberately block or limit AI access, but the key is to make that choice explicitly rather than inherit it from old defaults.

Is llms.txt worth publishing today?

Usually as a lightweight aid, yes. But it is still a signal, not a control mechanism, and it will not rescue weak structured data, poor server-side delivery, or broken discoverability.

What is the fastest win for most enterprise sites?

Audit your money pages for machine clarity: raw HTML delivery, heading structure, form labels, JSON-LD, and consistent entity naming. Then compare how major models describe the brand before and after the fixes.

The technical signals behind AI search are no longer a niche concern. If your brand wants to be understood, cited accurately, and ready for the next wave of agent interactions, start by auditing meaning, not just access, and turn the gaps into an execution plan with BotRank.

AI Search & GEO expert

After nearly 15 years in digital strategy on the client side (including 10 years at Olympique Lyonnais, where he was notably in charge of SEO).
Florian co-founded BotRank.ai in 2025, the GEO (Generative Engine Optimization) tool used by more than 2,500 companies to manage their visibility in AI-generated search results. He writes regularly about GEO and AI Search.