LLMO, or large language model optimisation, is the work of making your brand the source that ChatGPT, Claude, Perplexity and Google's AI Overviews retrieve, trust and cite when a buyer asks a question in your category. In practice that means letting the right crawlers in, being indexed where those engines search, stating facts plainly enough to be lifted into an answer, and being discussed by sources the engines already trust. Most of it is disciplined SEO. The smaller, newer part is where we usually find the gaps.

The stakes are easy to state. When Pew Research Center tracked the browsing of 900 US adults in March 2025, 18% of their Google searches produced an AI summary, and people who saw one clicked a traditional result in 8% of visits, against 15% for those who did not.[1] Clicks on links inside the summary happened in just 1% of visits.[1] The answer is increasingly the destination, which makes being named in it the new first position.


LLMO, GEO and AEO: Three Names for One Job

Generative engine optimisation, or GEO, was coined in a 2023 paper by Pranjal Aggarwal and colleagues, later published at KDD 2024, as a way for content creators to raise their visibility in AI-generated responses.[2] Answer engine optimisation (AEO) is the older term, inherited from the era of featured snippets and voice assistants. LLMO is the umbrella for the lot, and it is what our LLMO & SEO service is built around.

Google's guide to generative AI features, updated in July 2026, says that from its point of view optimising for AI search is simply SEO.[3] That holds for AI Overviews, which sit on Google's own index. It is less complete for ChatGPT, Claude and Perplexity, which run their own crawlers and rules.


How AI Answer Engines Choose Their Sources

Two Routes In: Training and Retrieval

A model can mention you because your brand was in its training data, or because the product fetched your page while composing the answer. The first is slow and largely out of your hands. The second is where citations come from. Google describes its AI features as retrieval-augmented generation over the Search index, plus query fan-out: the model issues a batch of related sub-queries, gathers results for each, then writes a response linking to the pages that support it.[3] So the useful question is less "does the model know us" than "when it goes looking, can it reach us, and does it pick us".

Crawlability: Know Which Bot Does What

Each major engine now separates its training crawler from its search crawler, and both from the agent that fetches a page when a user asks.

  • OpenAI: GPTBot collects potential training data. OAI-SearchBot powers ChatGPT search; block it and you will not appear in ChatGPT search answers beyond bare navigational links. ChatGPT-User makes user-triggered visits.[4]
  • Anthropic: ClaudeBot gathers training data, Claude-SearchBot indexes for search, and Claude-User fetches pages for people using Claude. Restricting the latter two may reduce your visibility in Claude's results.[5]
  • Perplexity: PerplexityBot surfaces and links sites in results and, the company says, is not used for training. Perplexity-User handles live fetches and generally ignores robots.txt.[6]
  • Google: AI Overviews and AI Mode draw on the ordinary Googlebot index. Google-Extended governs Gemini training and grounding, and has no effect on Search inclusion or ranking.[7]

So "block AI" is no longer one decision. A publisher whose content is the product might disallow the training crawlers while allowing the search ones:

User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /

User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

Most brands that want to be recommended should allow both, since training data shapes what future models know about your category. Either way, robots.txt is only half the picture. Cloudflare says that from 15 September 2026, new domains on its network block training and agent bots by default on pages that display ads, and customers who block training crawlers, including via its older "Block AI bots" setting, also block multi-purpose crawlers such as Googlebot and BingBot.[8] Google's new Search generative AI control in Search Console, live for all sites since 31 August 2026, lets a site opt out of AI Overviews and AI Mode.[9] Inclusion is the default; check nobody has changed it.

Entity Clarity and Structured Data

Before an engine can cite you, it has to be confident who you are. Google says it uses structured data to understand pages and the companies named in them,[10] and its Organization markup includes a sameAs property pointing to your profiles elsewhere.[11] That is the entity layer: one consistent name, description and set of facts, marked up in JSON-LD and matched on LinkedIn, Crunchbase, review platforms and, where you legitimately qualify, Wikipedia.

Keep expectations proportionate: Google says no special schema is required for its AI features.[3] Schema is disambiguation, not persuasion. It will not make a weak page quotable.

Passages That Survive Being Quoted

The most useful evidence on content is still the GEO study. Using a benchmark of 10,000 queries, Aggarwal and colleagues found that adding citations, quotations from credible sources and statistics lifted a source's visibility in generated answers by 30–40% on their word-count measure, while keyword stuffing offered little to no improvement.[2] Lower-ranked pages gained most: citing sources raised visibility for fifth-ranked results by 115%, while the top-ranked result lost ground. On Perplexity itself, improvements reached 37%.[2]

Treat that as direction rather than law; it was measured on 2023-era engines. Google adds a counterweight, telling site owners not to chop content into fragments for machines and warning that mass-producing pages for fan-out queries breaches its scaled content abuse policy.[3] The positions reconcile easily. Write for people, but make each claim specific, sourced and able to stand alone when lifted out of the page.

Third-Party Mentions Outweigh Your Own Site

This is the part brands underestimate. A 2025 study by Mahe Chen and colleagues found AI search shows a "systematic and overwhelming bias towards Earned media" over brand-owned and social content, a sharper skew than Google's classic results.[12] Pew found Wikipedia, YouTube and Reddit together accounted for 15% of the sources listed in Google's AI summaries.[1]

So the work extends beyond your domain, into trade press, analyst coverage, podcasts, review sites and the communities your buyers read. Google notes that manufactured mentions help less than they appear to.[3]

llms.txt: Useful, Not Magic

The llms.txt proposal, published by Jeremy Howard in September 2024 and revised in August 2026, puts a short Markdown file at your site root: a title, a summary and curated links to clean versions of key pages.[13] Howard says agents mainly use it at inference rather than for training,[13] and Chrome's Lighthouse now checks for one in its agentic browsing audits.[14]

Google Search ignores it and says the file neither helps nor harms visibility there.[3] We set one up as standard because it takes an afternoon and agents do read it. Anyone selling it as a ranking lever is overselling.


A Practical LLMO Playbook

Our advice to clients runs in this order, which front-loads the cheapest fixes.

  1. Audit access. Read robots.txt, CDN bot settings and the Search Console control together. Confirm search crawlers can fetch key pages, and that the content sits in the HTML rather than behind heavy JavaScript.
  2. Write the prompt set. List the 30 to 50 questions buyers ask before shortlisting a supplier, in their words. Run them across ChatGPT, Claude, Perplexity, Gemini and Google AI Mode, and record who gets cited.
  3. Fix the entity layer. One canonical company description, Organization schema with sameAs links, consistent facts on every profile you control.
  4. Rebuild the pages that should be cited. Lead with the answer, back claims with named sources, add FAQs that mirror the prompt set, and show authorship and dates. Google's AI features also surface images and video,[3] so original visual content earns its place.
  5. Earn the third-party layer. Commentary, research, interviews and reviews on sites the engines already trust. It is slow, and it is where durable citation share comes from.
  6. Publish the machine-readable extras. llms.txt, Markdown versions of priority pages, clean sitemaps.

How to Measure Citation Share

Citation share is the proportion of all citations, across a fixed set of prompts and engines, that point to your domain: share of voice for AI answers. Track it alongside whether you are mentioned at all, and whether what is said is accurate.

First-party data improved sharply this year. Bing Webmaster Tools launched an AI Performance report in February 2026 showing where a site is cited across Copilot, Bing and partner AI experiences, and in June added Citation Share: your percentage of all citations shown for a given query.[15] Google Search Console's Generative AI performance report, covering impressions in AI Overviews and AI Mode, reached all sites on 31 August 2026.[16] ChatGPT appends utm_source=chatgpt.com to outbound links, so its referrals can be separated in analytics.[17]

For ChatGPT, Claude and Perplexity, you are still measuring from the outside. Run the prompt set on a schedule, several times per prompt because answers vary between runs, and log cited, mentioned or absent for you and your rivals. Tools such as Profound, Otterly and Goodie automate the grind, though Google warns that no third-party tool has access to its internal systems.[3]

Then judge the visits, not the count. Google says clicks from results pages with AI Overviews tend to be higher quality, with people staying longer on site, and suggests weighting conversions over raw clicks.[18] Add AI assistants to the "how did you hear about us" field on your forms; it is crude, and often the most honest number you will get.


The Second Front Door

Search has not been replaced. It has grown a second front door, and the brands being named there did the unglamorous work first: open to the right crawlers, unambiguous about who they are, quotable on the page and cited elsewhere. None of it needs a trick. All of it needs an owner.

If you want that handled across Google and the AI engines at once, that is what our LLMO & SEO service is for, or you can start a conversation.


Frequently asked questions

What is the difference between LLMO, GEO and AEO?

They describe the same job with different emphasis. GEO (generative engine optimisation) comes from a 2023 academic paper on improving visibility in AI-generated answers, AEO (answer engine optimisation) dates from the featured-snippet and voice-assistant era, and LLMO (large language model optimisation) is the umbrella term. All three mean making your content easy for AI search engines to find, trust and cite.

How do I get my website cited by ChatGPT?

Start by allowing OAI-SearchBot in robots.txt and your CDN settings, because OpenAI says sites that block it will not appear in ChatGPT search answers. Then make key pages state specific, sourced facts that can be quoted on their own, keep your company details consistent across the web, and earn coverage on third-party sites. Track the results using the utm_source=chatgpt.com parameter ChatGPT adds to its links.

Should I block GPTBot and ClaudeBot?

Only if keeping your content out of model training matters more to you than being known to future models. GPTBot and ClaudeBot are training crawlers, while visibility in ChatGPT and Claude search depends on OAI-SearchBot and Claude-SearchBot, which you can allow separately. Most brands that want AI engines to recommend them should allow both kinds.

Does llms.txt help with Google AI Overviews?

No. Google says Google Search ignores llms.txt files, so they neither help nor harm visibility in AI Overviews or AI Mode. The file is still worth publishing because AI agents use it at inference time to find a site's key content, and it takes very little effort to maintain.

How do you measure citation share in AI search?

Define a fixed set of real buyer prompts, run them repeatedly across ChatGPT, Claude, Perplexity, Gemini and Google AI Mode, and calculate the percentage of all citations that point to your domain. Supplement that with first-party data: Bing Webmaster Tools reports citation share for Copilot and Bing, and Google Search Console reports impressions in AI Overviews and AI Mode.


References