What Is Generative Engine Optimization (GEO)?
GEO is the practice of making a site retrievable and citable by AI answer engines. Here is what it is, what it is not, and what the evidence supports.
Generative Engine Optimization (GEO) is the practice of improving how AI answer engines retrieve, interpret and cite a website. It is a young discipline with a thin evidence base, a great deal of vendor noise, and a small core of practices that hold up. This article separates the three, and states the evidence tier for every recommendation it makes.
A working definition
An AI answer engine is any product that answers a question in generated prose rather than a ranked list of links: Google's AI Overviews and AI Mode, ChatGPT with search, Perplexity, Claude with web search, and Copilot. Most of them retrieve web content at answer time and attach some form of attribution to the answer. GEO is whatever a site owner can legitimately do to make their pages more likely to be retrieved for relevant questions and, once retrieved, more likely to be used and cited.
Three things in that sentence carry weight.
Retrieved is not the same as crawled. A crawler fetching your page proves it can be read. It says nothing about whether the page will be selected when a question is asked. Retrieval happens at prompt time against an index, and the index is not the crawl log.
Cited is not the same as retrieved. A retrieved page may inform an answer without being named in it, or may be discarded because a competing passage answered the question more directly.
Legitimately rules out a whole category of "GEO tactics" that amount to producing content for machines rather than people. Google's spam policies define scaled content abuse as generating many pages for the purpose of manipulating rankings rather than helping users, and the policy is explicit that AI-generated volume is covered. A technique that would be spam in Search is spam in AI features too.
The full chain, which every claim on this site has to respect, is: crawled, then retrieved, then cited, then recommended. Each step is a separate event with a separate failure mode.
What GEO is not
Much of what is sold as GEO is a relabelling of things that already existed, or a promise the evidence cannot support.
It is not a new ranking system. Google states that there are no additional requirements to appear in AI Overviews or AI Mode, and no special optimisations are necessary. A page must be indexed and eligible to be shown with a snippet, exactly as for ordinary Search. Google also says explicitly that you do not need new machine-readable files, "AI text files", or markup to appear in these features. That is a first-party statement and it should shape every expectation you hold about Google's surfaces.
It is not llms.txt. The llms.txt proposal is a plain-text file that offers a curated summary of a site for language models. No major provider documents reading it, and no provider ties it to retrieval or citation. Publishing one is cheap and harmless, but llms.txt presence is not ingestion. It belongs in the EXPERIMENTAL tier.
It is not a schema trick. Structured data helps machines understand page content and is an established standard. It does not cause citation. Schema that misrepresents a page or exists only to win a rich result is against Google's policies regardless of which surface you are aiming at.
It is not "ranking in ChatGPT". A generated answer has no stable ordered result list. One prompt, run twice, can name different sources. Any tool that reports a numeric position for your brand in a generative answer is reporting something it invented. The honest unit of measurement is an observation under recorded conditions, aggregated into a rate with a confidence interval. See how LLM visibility monitoring works for the mechanics.
It is not separate from AEO. Answer Engine Optimization and GEO are the same practice under two names. Treating them as distinct disciplines with distinct evidence is marketing.
How GEO relates to SEO
The overlap is large, and the direction of dependency is one way: GEO builds on SEO, and nothing in GEO substitutes for it.
| Layer | What it governs | Who documents it | Evidence tier |
|---|---|---|---|
| Crawl access | Whether a fetcher can read the page | Robots Exclusion Protocol (RFC 9309), each provider's crawler docs | ESTABLISHED STANDARD / OFFICIAL PROVIDER GUIDANCE |
| Indexability | Whether the page can enter an index at all | Google, Bing | OFFICIAL PROVIDER GUIDANCE |
| Snippet eligibility | Whether text may be shown or used | Google (nosnippet, max-snippet, noindex) |
OFFICIAL PROVIDER GUIDANCE |
| Passage structure | Whether a retrieval system can isolate a self-contained answer | Retrieval-system design in general | EMERGING PRACTICE |
| Entity clarity | Whether machines can tell who and what the page is about | Structured data (standard); effect on citation (unconfirmed) | ESTABLISHED STANDARD for markup; EMERGING PRACTICE for effect |
| Citation-worthiness | Whether the page offers something a model will attribute | No provider documentation | EMERGING PRACTICE at best |
The first three rows are conventional SEO. If a page fails there, nothing below matters. Google's guidance on AI features points site owners at the same controls they already use for snippets: nosnippet, data-nosnippet, max-snippet and noindex limit what the AI features can show from a page. The controls that keep you out of AI Overviews are the controls that keep you out of Search snippets.
The lower rows are where GEO adds something, and they are also where the evidence thins out. A detailed comparison is in SEO vs GEO.
Where AI systems actually read your site
Each provider documents separate crawlers for separate jobs, and the distinction matters for policy. As of the date these sources were checked:
| Provider | Training | Search index | User-triggered fetch |
|---|---|---|---|
| OpenAI | GPTBot |
OAI-SearchBot |
ChatGPT-User |
| Anthropic | ClaudeBot |
Claude-SearchBot |
Claude-User |
| Perplexity | none documented | PerplexityBot |
Perplexity-User |
Google-Extended (a token, not a separate crawler) |
Googlebot |
— |
Several first-party statements shape what you can expect from blocking or allowing them.
- OpenAI states that sites opted out of
OAI-SearchBotwill not be shown in ChatGPT search answers, and thatChatGPT-Useris not used for automatic crawling. - Anthropic states that blocking
Claude-SearchBotmay reduce a site's visibility and accuracy in user search results, and that blockingClaudeBotsignals that future material should be excluded from training. - Perplexity states that
PerplexityBotis not used to collect content for training foundation models, and thatPerplexity-Usergenerally ignores robots.txt because a user requested the fetch. - Google states that
Google-Extendedgoverns whether crawled content may be used for training Gemini models and for grounding in Gemini Apps and Vertex AI, and that it does not affect inclusion in Google Search nor act as a ranking signal.
All four are OFFICIAL PROVIDER GUIDANCE. Note what none of them says: that allowing a crawler produces citations. Access is a precondition, not a cause. This is the point of crawled is not cited.
The evidence problem
The evidence base for GEO is weaker than for SEO for a structural reason. Search engines publish documentation about what they index and how site owners can control it, and they have done so for two decades. AI answer engines publish crawler documentation and product terms, and little else about how they select what to cite.
The 2023 academic paper that coined the term "generative engine optimization" reported that adding citations, statistics and quotations to content increased visibility in a synthetic generative engine the authors built. That is an interesting result about one experimental system. It is not evidence about ChatGPT, Gemini, Perplexity or Claude as shipped, and it should not be cited as if it were.
What this means in practice:
- A recommendation about crawl access, indexability or snippet controls can be made with confidence, because a provider documents it.
- A recommendation about passage structure, entity clarity or authorship transparency has a plausible mechanism and is widely practised, but no provider has confirmed that it changes citation behaviour. It sits at EMERGING PRACTICE.
- A recommendation about a specific file, phrase pattern or markup "for AI" with no documented consumer sits at EXPERIMENTAL or HYPOTHESIS, and must not be scored as a deficiency when it is absent.
SiteRank AI enforces this in its audit: findings at the bottom two tiers are excluded from status calculations, so a site is never marked "needs attention" because it lacks something no provider has asked for. The methodology page describes the weighting.
Practical actions, each with its tier
These are the things worth doing, in roughly the order to do them. The tier is the honest ceiling of what can be claimed for each.
1. Make the page indexable and snippet-eligible — OFFICIAL PROVIDER GUIDANCE
Google's eligibility rule for AI Overviews and AI Mode is indexing plus snippet eligibility. Check for accidental noindex, nosnippet or restrictive max-snippet values, a canonical pointing elsewhere, or a robots.txt rule that blocks the page. A page that cannot appear in Search cannot appear as a supporting link in AI features.
2. Decide crawler policy deliberately — OFFICIAL PROVIDER GUIDANCE
Separate training crawlers from search crawlers from user-triggered fetchers. Blocking GPTBot or ClaudeBot expresses a training preference; blocking OAI-SearchBot or Claude-SearchBot removes you from those products' search answers by the providers' own statements. Serve robots.txt with a 200 status: RFC 9309 says a crawler may treat a 4xx as "no restrictions" but must treat a 5xx as complete disallow, so an erroring robots.txt is worse than none. See robots.txt for AI crawlers.
3. Structure content as answerable passages — EMERGING PRACTICE
Retrieval systems generally index passages, not documents. A heading that states a question or a topic, followed by a paragraph that answers it completely without depending on surrounding text, is the unit most likely to be retrieved intact. Very long unbroken sections and very short fragments both work against this. No provider documents a passage-length threshold; this is inference from how retrieval systems are built.
4. Make the entity unambiguous — ESTABLISHED STANDARD for the markup, EMERGING PRACTICE for the effect
State who the organisation is, in text, on an about page and a contact page. Mark it up with Organization and WebSite JSON-LD that agrees with the visible text. Use one canonical name. Schema.org is a published standard, so the markup itself is standard practice; whether it moves citation behaviour is not confirmed by any provider. Entity clarity goes deeper.
5. Show who wrote it and when — OFFICIAL PROVIDER GUIDANCE for Search quality, EMERGING PRACTICE for AI citation
Google's helpful-content guidance asks whether it is self-evident who authored the content, and recommends disclosing how automation or AI generation was used. Bylines, author pages and visible dates are conventional quality practice. Their effect on AI citation is plausible but unconfirmed.
6. Add something the model cannot get elsewhere — EMERGING PRACTICE
Original data, first-hand observation, a worked example, a precise definition. Google's self-assessment questions ask whether content provides original reporting or analysis beyond the obvious; this is the same test. A page that restates what ten other pages say gives a retrieval system no reason to choose it.
7. Publish llms.txt if you want to, and expect nothing — EXPERIMENTAL
It costs a few minutes, the format is simple, and it does not hurt. It is not read by any documented provider crawler for retrieval, and its absence is not a deficiency. What is llms.txt covers the format and its actual status.
8. Measure, rather than assume — methodology, not a tactic
Run a controlled prompt set against provider APIs, record every response with its conditions, and compute mention and citation rates with confidence intervals. Then change something and observe again. This is the only way to know whether any of the above did anything for your site. It is what the AI visibility feature does, with the caveat that an API observation is not what a person sees in a consumer chat product.
What not to assume
- That a crawler visit means your page is in a retrieval index.
- That a mention means a citation, or that a citation means a recommendation.
- That any tier below STRONG EVIDENCE is a requirement, a ranking factor or a signal.
- That an
llms.txtfile is read by anyone. - That one query on one day tells you anything about a trend.
- That a consumer chat product behaves like the API you can measure.
Key takeaways
- GEO is optimisation for retrieval and citation by AI answer engines. It is an emerging discipline, and every claim in it carries an evidence tier.
- Google states there are no additional requirements or special optimisations for AI Overviews and AI Mode beyond ordinary Search eligibility.
- Crawl access, indexability and snippet eligibility are documented and come first. Passage structure, entity clarity and authorship are plausible and widely practised but unconfirmed.
- Crawled is not retrieved, retrieved is not cited, and a generative answer has no rank.
- Measure with controlled prompts and confidence intervals, or you are guessing.
Official sources & further reading
- AI features and your website — Google Search Central
- Spam policies for Google web search — Google Search Central
- Creating helpful, reliable, people-first content — Google Search Central
- Google's common crawlers, including Google-Extended — Google Search Central
- Overview of OpenAI crawlers — OpenAI
- Does Anthropic crawl data from the web, and how can site owners block the crawler? — Anthropic
- Perplexity crawlers — Perplexity
- RFC 9309: Robots Exclusion Protocol — IETF
Related reading
Frequently asked questions
If Google says no special optimisation is needed, is there anything to do at all?
Yes, but most of it is work you would already recognise as SEO. Google's statement rules out a separate AI checklist for its own features; it does not say that indexability, snippet eligibility, clear entity information and well-structured pages stop mattering, and those are the same properties that decide whether any retrieval system can use a page. What is genuinely new is deciding a crawler policy for the providers that run their own crawlers, and measuring what those systems actually say about your site. Anything sold as a distinct AI-only technique should be held to the tier its evidence supports, which is usually EMERGING PRACTICE or lower.
Should I block AI crawlers to protect my content?
That depends on whether you are trying to keep content out of training or trying to appear in AI answers, and those are governed by different user agents. Blocking GPTBot or ClaudeBot expresses a preference about training material; blocking OAI-SearchBot or Claude-SearchBot removes you from those products' search answers, by the providers' own statements. There is no universally correct choice — it is a business decision about reuse of your work against discoverability in those products. Whatever you decide, decide it per user agent rather than with one blanket rule, and see should you block AI training crawlers for the trade-offs.
A recommendation is marked EXPERIMENTAL. Should I skip it?
Not necessarily, but never treat its absence as a fault. An evidence tier says what may be claimed for a practice, not whether the practice is worth a few minutes: a cheap, harmless action with no evidence behind it is a defensible bet as long as nobody bills it as a fix. The rule that keeps this honest is that budget and reporting follow evidence — spend the first hour on the documented items, and never let an unconfirmed practice make a working site look broken.
How do I know whether any of this made a difference?
You measure before and after with the same controlled prompt set, and you accept that the result is an association rather than a cause. Model updates, index refreshes and competitors publishing move these numbers at the same time your changes do, so a single before-and-after pair of answers proves nothing at all. Run enough samples that each rate carries a usable confidence interval, hold the prompts and the model fixed between runs, and treat any movement smaller than the interval as noise. That is slower than most GEO reporting, and it is the version that survives scrutiny.