Generative Engine Optimization
Generative Engine Optimization : get AI assistants to cite your brand
More and more searches end in a generated answer: an AI Overview on Google, a conversation in ChatGPT, a summary in Perplexity. Those answers are built from pages the system retrieves and considers trustworthy. We work to make yours one of them.
What GEO is
From ranking links to being the source of the answer
Generative Engine Optimization (GEO), sometimes called SEO for AI, is the work of getting your content and brand to appear, accurately represented, in the answers produced by search engines built on large language models. It doesn't replace SEO; it extends it. Assistants that search the web rely on search indexes, and pages that aren't crawled or aren't considered trustworthy rarely make it into the answer.
What changes is the unit of competition. In classic results, a whole page competes for a position. In a generated answer, passages compete: a paragraph that defines something clearly, a table comparing options, a figure with its source. And something new matters too: whether the model recognises your brand as an entity and associates it with the right topics.
We're careful with promises. Generative systems change quickly, don't publish their full criteria, and their answers vary from one query to the next. We work with what the providers themselves document and with what we can observe and measure.
- Correct access for AI crawlers
- Content structured to be extracted and cited
- Entity signals and brand authority
- Tracking of mentions and citations in answers
How it works
How generative engines choose their sources
Each system has its own quirks, but those that answer with current information share a general pattern.
Retrieval from an index
When a question needs current or specific information, the assistant runs searches against a web index (its own or a search engine's) and retrieves a set of candidate pages. AI Overviews draw on Google's index; Copilot on Bing's. If your page isn't in those indexes, it can't be a candidate.
Query fan-out
A complex question is broken into several sub-queries that are searched separately. Google has described this technique for its AI features. The practical consequence: a page can be cited for answering one specific sub-question well, even if it doesn't rank for the original question.
Preference for clear, extractable passages
The model works with fragments. A paragraph that answers directly, a precise definition, a list of steps or a comparison table are easier to extract than an idea spread across several meandering paragraphs. Content structure matters as much as quality.
Entity and brand signals
Models learn which brands are associated with which topics from what's published on the web: your own site, the press, directories, reviews, profiles. A brand described consistently across many trusted sources is more likely to be mentioned. We work on this through entity SEO.
Freshness
For questions about prices, regulations, products or news, systems tend to prefer recent information. Visible update dates and content that has genuinely been revised, not just re-dated, help your page stay eligible.
Trust and authority
As in classic SEO, sources with clear authorship, references, reputation and quality links carry more weight, especially for health, money or legal topics. E-E-A-T signals remain relevant.
AI crawlers
Which agent does what, and what blocking it means
The main providers distinguish between crawlers that gather training data and crawlers that search and cite. Blocking one doesn't have to mean blocking the other.
| User agent | Provider | What it's used for | If you block it in robots.txt |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Indexing pages so they can be shown and linked in ChatGPT search | Your pages may stop appearing as results or citations in ChatGPT search |
| GPTBot | OpenAI | Collecting content that may be used to train models | Your content is no longer collected for training; according to OpenAI, this doesn't affect OAI-SearchBot |
| Google-Extended | A control token (not a separate crawler): decides whether content crawled by Google may be used to train Gemini models and to ground answers in Gemini apps and Vertex AI | No effect on inclusion or ranking in Google Search, and it doesn't remove you from AI Overviews | |
| Googlebot | Crawling for Google Search, including AI Overviews and other search features | You disappear from Google Search. To limit snippets, use nosnippet or max-snippet rather than blocking | |
| PerplexityBot | Perplexity | Indexing pages so they can be shown and cited in Perplexity results | Reduces the chance of Perplexity citing you as a source |
| Bingbot | Microsoft | Crawling for Bing, whose index feeds Copilot's answers | You disappear from Bing and, with it, from much of Copilot |
robots.txt
Decide deliberately what you allow and what you don't
Many sites block AI crawlers without knowing it: inherited rules, security plugins or CDN settings that filter unfamiliar agents. Others allow everything without ever having decided to. Both are business decisions and should be made with full information.
A common approach is to allow search crawlers (the ones that generate citations and visits) and decide separately about training crawlers, as the accompanying example shows. Bear in mind that robots.txt is a convention the major providers say they respect, not a technical barrier, and that agents acting on a user's direct request (such as ChatGPT-User or Perplexity-User) follow rules each provider documents separately.
We also review the infrastructure layer: firewalls, bot protection and CDN rules, because a block there doesn't show up in robots.txt but has exactly the same effect.
Technical foundations for AI
What your site needs to be readable by assistants
Rendering without relying on JavaScript
Google executes JavaScript, but many AI crawlers, judging by what shows up in server logs, mainly read the initial HTML. If your main content, prices or links are loaded by JavaScript in the browser, they may not see them. Server-side rendering or static generation removes that risk.
Learn moreStructured data
Schema.org helps search engines interpret unambiguously who you are, what you offer and how your pages relate: Organization, Product, Article with author, LocalBusiness, BreadcrumbList. It isn't a direct route to being cited, but it strengthens understanding of the entity and feeds the search systems that answers rely on.
Learn morellms.txt
A proposal that emerged in 2024: a Markdown file at the root of the domain that summarises the site and links its most useful pages for language models. It isn't a standard today, and no major search engine has confirmed using it to choose sources. We implement it where it's cheap and makes sense (technical documentation, for example), without presenting it as a guaranteed lever.
Process
How we approach GEO
Four phases built on healthy SEO foundations. If those foundations are broken, we fix them first.
AI visibility diagnosis
We define a set of questions representative of your business: how people would look for you, how they'd compare options, what they want to know before deciding. We run them across several assistants and record whether you appear, how you're described, which sources are cited and who takes your place.
Because answers vary between queries and over time, we repeat the tests and look at trends, not isolated screenshots.
Activities
Outputs
- Map of presence and gaps in AI answers
- Sources that cite your competitors
Technical access
We review robots.txt, CDN and firewall configuration, rendering and logs to confirm that AI search crawlers can reach your content. We also check indexing in Google and Bing, on which AI Overviews and Copilot depend.
We agree a policy for training crawlers with you and document it.
Activities
Outputs
- Documented crawler policy
- Access issues resolved
Citable content
We rewrite and expand key pages so they contain clear, extractable answers: definitions up front, numbered steps, comparison tables, figures with source and date, and sections that address the sub-questions assistants tend to generate.
We add what only you can provide: first-hand experience, real examples from how you operate, decision criteria. That's what separates a cited source from one of many.
Activities
Outputs
- Restructured pages
- Content plan for AI questions
Authority and monitoring
We strengthen external signals: mentions in relevant media and directories, consistent profiles, reviews and entity structured data. Assistants reflect what the web says about you, not just what you say about yourself.
We track the same question set regularly and combine it with referral traffic from assistants in GA4, to see whether presence turns into visits and enquiries.
Activities
Outputs
- Monthly AI visibility report
- Priorities reviewed every quarter
Zero-click searches
Measuring when the answer is read without a visit
Some queries are resolved within the answer itself: the user reads the summary and doesn't click. That doesn't make the visibility worthless. Being the brand mentioned in the answer influences consideration, later branded searches and direct visits.
So we measure on several layers: presence and type of mention across the question set, referral traffic from assistants (ChatGPT, Perplexity, Copilot and others show up as sources in GA4 when they send visits), the trend in branded searches in Search Console, and assisted conversions. No single metric tells the whole story; together, they let you make decisions.
- Presence across a fixed question set, tracked over time
- Referral traffic from assistants in GA4
- Trend in branded searches
- Linked conversions
FAQ
About generative engine optimisation
Can you guarantee ChatGPT will recommend my company?
No. Nobody can: answers vary, and providers don't publish their full criteria. What we can do is remove obstacles, improve the content and signals that available documentation and observation point to as relevant, and measure progress honestly.
Is GEO different from SEO?
It's an extension. Most of the work (indexing, useful content, authority) is SEO. What's specific is AI crawler access, passage-oriented content structure, entity signals and measurement inside assistants.
Should I block GPTBot?
That depends on your policy on your content being used to train models. According to OpenAI, blocking GPTBot doesn't stop you appearing in ChatGPT search, which relies on OAI-SearchBot. We help you make the decision with its implications clearly laid out.
Does blocking Google-Extended remove me from AI Overviews?
No. Google-Extended controls whether content is used to train Gemini models and to ground answers in Gemini apps and Vertex AI. AI Overviews are part of Google Search and depend on Googlebot; to limit how your snippets are shown, there are directives such as nosnippet.
Is an llms.txt file worth having?
It's a proposal without confirmed adoption by the major search engines. It costs little to implement and can be useful on certain sites, but we don't treat it as a primary lever or sell it as one.
How long before it shows?
Access fixes are reflected once systems recrawl. Content and authority changes, as in SEO, take months. Because answers vary, we assess trends across several measurements rather than results on a single day.
Related
Completing your AI visibility
Answer Engine Optimization
Answer Engine Optimization: we structure your content for featured snippets, People Also Ask, voice assistants and Google's AI Overviews.
Entity SEO
Entity SEO: Knowledge Graph, knowledge panels, Organization schema and sameAs, consistent NAP and topical authority so Google and AI recognise your brand.
How to appear in ChatGPT and AI Overviews
How ChatGPT, Perplexity, Gemini and Google AI Overviews choose and cite sources: practical GEO steps, how to measure results and what is still unknown.
Do AI assistants mention you?
We analyse how your brand appears in ChatGPT, Gemini, Perplexity and AI Overviews, and what to do to improve it.