Generative Engine Optimization (GEO): How to Get Cited by ChatGPT, Claude and Perplexity
A growing share of buying research now happens inside AI assistants instead of Google's ten blue links. When someone asks ChatGPT "what's the best attribution tool for Shopify," the brands that get named win the consideration set before a website is ever visited. Optimizing for that moment is called Generative Engine Optimization (GEO) — sometimes Answer Engine Optimization (AEO) — and it is learnable.
It is also, by mid-2026, a field with a widening gap between what gets recommended and what the evidence supports. This guide covers both: the moves with real support behind them, and the popular tactic that the data has quietly stopped backing.
How AI assistants actually choose sources
Assistants cite content through two paths. First, training data: what the model absorbed about your brand before its cutoff. Second — and far more actionable — retrieval: live web search that ChatGPT, Claude, Perplexity and Google's AI Mode run at answer time, then summarize with citations.
The retrieval path has become more complex than a single lookup. A user's question is now typically decomposed into multiple sub-queries before any document is fetched, so the engine assembles an answer from several searches rather than one. The practical consequence is that you are rarely competing for one keyword. You are competing to be the best available source for a facet of a question — the definition, the comparison, the counter-argument, the number — and a page that owns one facet cleanly can get cited even when it would never have ranked first for the parent query.
Retrieval favors pages that are crawlable, fast, factually dense, and structured so a machine can extract an answer in one pass. That much has been stable since the discipline was named. The term GEO itself comes from a 2023 research paper out of Princeton, Georgia Tech and the Allen Institute for AI, which tested optimization strategies across thousands of queries and found the largest visibility gains came from adding quotations, statistics and citations to content, with fluency improvements close behind. Those findings have held up better than most things published about GEO since.
What has changed is the overlap with classic search. Industry analysis summarized by Search Engine Land indicates that the set of pages winning AI citations and the set winning organic rankings now share well under a fifth of their members, down from roughly 70% two years earlier. Ranking is no longer a proxy for being cited. That divergence is the whole reason this discipline exists as something separate from SEO.
The llms.txt question, answered honestly
An earlier version of this article recommended shipping an llms.txt file as a cheap asymmetric bet. The bet was reasonable at the time. The evidence has since come in, and it is not kind.
Google has said no on the record: Gary Illyes confirmed in July 2025 that Google does not support llms.txt and has no plans to, and John Mueller has publicly compared it to the long-discredited keywords meta tag. On the crawler side, the log data is worse than ambivalent. Limy reported analyzing over 500 million AI bot traffic events across a 90-day window and found only 408 requests targeting /llms.txt directly — GPTBot, ClaudeBot, PerplexityBot and OAI-SearchBot overwhelmingly crawl HTML and skip the file. OtterlyAI ran a similar test on its own domain, logged 84 llms.txt requests out of 62,100 AI bot visits, and subsequently removed the file from its GEO audit checklist on the grounds that it was drawing attention away from things that actually move citation frequency.
There is a counterpoint worth recording. Profound, which specializes in GEO tracking, has reported that Microsoft and OpenAI crawlers do actively fetch both llms.txt and llms-full.txt. And adoption is no longer trivial — an SE Ranking study of 300,000 domains put implementation at roughly one site in ten.
So where does that leave an operator? Three honest conclusions:
- If you publish developer documentation, ship both files. This is where the standard genuinely lives. Coding assistants and MCP-based integrations can be pointed at an
llms-full.txtto pull clean, current docs into a session. Note what makes this case different: the file is fetched on demand by a tool the user configured, not discovered by a crawler deciding what to cite. That demand side is real. - If you run a content or commerce site and you are shipping it to improve AI citations, the current evidence says it will not. Ship it anyway if it costs you an hour and a build step — it is a machine-readable surface for the agentic web, and that is a different and more plausible bet than the SEO one. Just do not count it as GEO work, and do not let it displace anything below.
- Never confuse it with access control.
llms.txtblocks nothing. That isrobots.txt, and getting that right is not optional.
The technical layer that does matter
Let the right crawlers in. Explicitly allow GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended in robots.txt. This is the single most common own-goal in the category: default bot protection on a CDN or WAF quietly blocks AI crawlers, and the site vanishes from AI answers with no error anyone will ever see. Check your actual server logs for these user agents rather than assuming.
Server-render everything. Retrieval crawlers are impatient with JavaScript. Static or server-rendered HTML gets read; client-rendered single-page applications frequently get skipped entirely. If your content only exists after hydration, it does not exist.
Use structured data everywhere. Schema.org JSON-LD — Article, FAQPage, ItemList, Organization — tells machines exactly what each page asserts. FAQPage markup in particular maps one-to-one onto how assistants answer questions.
Keep pages fast and clean. Retrieval has a time budget. A page that takes four seconds to deliver its main content loses to one that takes four hundred milliseconds, independent of quality.
The content layer
Write answers, not essays. Assistants extract self-contained passages. Lead every section with a direct one-to-three-sentence answer, then elaborate. A page structured as question-heading followed by immediate answer gets quoted; a page that buries the answer in paragraph six does not. This is also why FAQ sections outperform their reputation — they are the purest possible form of an extractable passage.
Be factually dense. Numbers, definitions, comparisons and named entities give a model something concrete to quote. "Klaviyo is the default lifecycle platform for Shopify brands" is citable; "there are many great email tools" is not. This is the most durable finding from the original Princeton/Georgia Tech/Allen AI work, and it has survived every subsequent shift in how these systems retrieve.
Publish information gain, not summary. The failure mode of AI-assisted content production is that it converges on the consensus already present in the training distribution. A page that restates what ten other pages say gives a retrieval engine no reason to prefer it. Original data, first-hand operator experience, a contrarian read on a common claim, a comparison nobody else has run — these are what make a page the marginal source rather than a redundant one. If you have proprietary experience, that is your only genuinely defensible input.
Cover the comparison queries. "X vs Y," "best tools for Z," "alternatives to X" are the highest-commercial-intent prompts users give assistants. Honest comparison content — with real trade-offs, not affiliate cheerleading — is what retrieval engines prefer to cite, and it is also what a reader arriving from a citation actually wanted.
Build entity consistency. Your brand name, description and category should read identically across your site, LinkedIn, Crunchbase, Wikidata and any directory that lists you. Models resolve entities across sources, and inconsistency dilutes you. The 5W AI Platform Citation Source Index for 2026 found ChatGPT drawing a large plurality of its top citations from Wikipedia and encyclopedic sources — which is less an instruction to go edit Wikipedia than a signal about how heavily these systems weight corroborated, cross-referenced entity data.
SEO, AEO and GEO: what each is actually optimizing
| Classic SEO | AEO | GEO | |
|---|---|---|---|
| Goal | Rank in the results list | Own the direct answer | Be cited in a synthesized answer |
| Unit of success | Position + click | Featured snippet / AI Overview | Citation, mention, recommendation |
| Content shape | Comprehensive page per keyword | Question → immediate answer | Extractable, factually dense passages |
| Key signal | Links + relevance | Structure + schema | Information gain + entity authority |
| Measured by | Rankings, organic sessions | Snippet ownership | Citation share, AI referral traffic |
| Traffic profile | High volume, mixed intent | Often zero-click | Low volume, very high intent |
The three overlap by roughly 80% at the level of actual work — crawlable, fast, structured, genuinely useful content serves all of them. The 20% delta is where GEO earns its own name.
Measuring it
GEO measurement stopped being purely improvised in 2026. Microsoft added AI-visibility metrics to Bing Webmaster Tools in June 2026, free in preview — worth connecting given that Bing's index feeds ChatGPT search. Beyond that, three things every operator should run:
- Segment AI referral traffic. Isolate
chatgpt.com,perplexity.ai,claude.aiandgemini.google.comas referral sources in analytics. Expect low volume and unusually high engagement — visitors arrive pre-sold by the citation rather than shopping. - Log your money prompts monthly. Pick your 20–30 highest-intent prompts, run them through ChatGPT, Claude, Perplexity and Google AI Mode, and record whether you were cited and in what position within the answer. This is manual, it takes an hour, and it is the only measurement that maps directly to commercial outcome. Tools like Profound automate share-of-voice tracking if the manual version outgrows you.
- Track citations, not just clicks. You can be the source of truth for an answer and receive zero traffic. That still moves the consideration set. Judge the program on citation share first and referrals second, or you will kill work that is working.
Attribution here is genuinely hard, and it is the same structural problem covered in our MTA vs. MMM measurement guide: a channel that influences demand without leaving a click trail will be systematically undervalued by any click-based model. If AI citations become a meaningful share of how buyers find you, the honest measurement path runs through the same place it always does — incrementality testing, not attribution reporting.
The strategic point
Classic SEO and GEO are about 80% the same work. The 20% delta — answer-first formatting, entity consistency, comparison coverage, and above all publishing something a model cannot get anywhere else — is cheap to implement now and compounding, because early citations feed the next generation of training data. The brands that get cited today become the default answers of 2027 models.
What the last year has clarified is which parts of that delta are real. The file-based tactics were oversold. The content and entity work was undersold. If you have limited hours, spend them on being the most quotable source on a narrow set of questions rather than on machine-readable plumbing that the machines are not currently reading.
Bottom Line for Operators
Allow the AI crawlers, server-render your HTML, and structure every page as question-then-immediate-answer with real numbers attached. Publish something nobody else has — proprietary data, honest comparisons, first-hand operator judgment — because information gain is the only durable citation moat. Ship llms.txt if it costs you an hour, but classify it as agentic-web infrastructure rather than GEO, and never at the expense of the content work. Then measure it properly: monthly prompt logging plus segmented AI referral traffic, judged on citation share rather than sessions. For the tooling side, our directory tracks the AEO and AI-visibility platforms currently worth evaluating.
Frequently asked questions
What is Generative Engine Optimization (GEO)?
GEO is the practice of optimizing a website and its content so AI assistants — ChatGPT, Claude, Perplexity, Google AI Mode — find, trust and cite it when generating answers. The term originates from a 2023 Princeton, Georgia Tech and Allen Institute for AI paper that tested which content characteristics affect citation selection. It overlaps heavily with SEO but optimizes for a different outcome: being extracted into an answer rather than ranking in a list of links.
Is llms.txt actually used by AI companies?
Largely not, on current evidence. Google's Gary Illyes confirmed in July 2025 that Google does not support it and is not planning to. Limy reported only 408 direct llms.txt requests across more than 500 million AI bot events, and OtterlyAI found 84 requests out of 62,100 AI bot visits on its test domain before dropping the file from its audit checklist. Profound has reported that Microsoft and OpenAI crawlers do fetch it. The one clear-cut case is developer documentation, where coding assistants consume llms-full.txt on demand. For a content site hoping to improve citations, the evidence does not support it as a priority.
How is GEO different from SEO?
Roughly 80% of the work is identical — crawlability, speed, structured data, quality content. GEO adds explicitly allowing AI crawlers, answer-first formatting designed for passage extraction, entity consistency across the web, comparison coverage, and tracking citations rather than rankings. The distinction matters more than it used to: analysis summarized by Search Engine Land indicates the pages winning AI citations and those winning organic rankings now overlap by well under 20%, down from around 70% two years ago.
How do I check if my brand appears in AI answers?
Run your 20–30 highest-intent prompts through ChatGPT, Claude, Perplexity and Google AI Mode monthly and log whether you are cited and where in the answer. Connect Bing Webmaster Tools, which added free AI-visibility metrics in June 2026 and whose index feeds ChatGPT search. Segment referral traffic from AI assistant domains in analytics, but do not judge the program on that number alone — citations that produce no click still shape the buyer's shortlist.