AI search

How to appear in ChatGPT, Perplexity and Google's AI Overviews

A practical guide to GEO: how AI search engines pick their sources, what is within your control and how to measure it without fooling yourself.

AI searchUpdated: 14 min readBy the UrbanElevate team

AI assistants have not replaced search; they have wrapped themselves around it. Nearly every system that cites sources first queries a web index and then writes an answer from what it finds. Understanding that process is the foundation of any Generative Engine Optimisation (GEO) strategy that goes beyond the latest fad.

What GEO is and how it differs from SEO

GEO (Generative Engine Optimisation) is the set of practices aimed at getting a brand, product or piece of content to appear in, and be cited as a source by, the answers generated by systems such as ChatGPT with search, Perplexity, Gemini, Copilot and Google's AI features (AI Overviews and AI Mode). It overlaps heavily with Answer Engine Optimisation (AEO), which focuses on answering questions directly, and with classic SEO, on which it depends.

The main difference lies in the outcome you are after. In traditional SEO the goal is a position in a list of ten links and the click that follows. In generative search the user receives a written answer and, at most, a handful of citations. Three things matter there:

  • Being retrieved: your page is among the documents the system consults for that question.
  • Being used: the model finds a clear, reliable passage on your page that helps it build the answer.
  • Being cited or mentioned: your brand or URL appears visibly, and the way you are described is accurate.

None of the three is achieved with isolated tricks. They are achieved with the same things that work in organic search (useful content that can be crawled and is backed by external signals) plus particular attention to how text is structured and to the consistency of your brand as an entity.

How AI search engines select sources

The internal details of each system are not public, but the general architecture is well documented by the providers themselves and in the technical literature. Almost all of them follow a retrieval-augmented generation (RAG) pattern:

  1. Query interpretation. The system rewrites the user's question and often breaks it down into several sub-queries. Google has described this technique for AI Mode as query fan-out: running related searches in parallel to cover different aspects of the question.
  2. Retrieval. Those queries run against a web index. Google uses its own index; other systems combine their own indexes with external search providers. If your page is not in the index being queried, it cannot be a source.
  3. Passage selection. Specific passages, not whole pages, are extracted from the retrieved documents because they appear to answer each sub-query.
  4. Synthesis. The language model writes the answer by combining those passages with its prior knowledge.
  5. Attribution. The system links to some of the sources it used. Which sources are shown, and how many, varies by product and by query.

There is also a second route to visibility that does not go through retrieval: what the model "knows" about your brand from its training data. If a brand appears frequently and consistently in public sources, the model may mention it even without searching. That route is slower and more indirect, and you cannot influence it without building a genuine presence beyond your own website.

The practical consequence: in most cases the way into a generative answer is still being indexed and ranking in a search engine. That is why technical SEO is not optional in a GEO strategy.

What we know about each platform

These products change often; what follows describes the landscape at the time of writing and is worth checking against each provider's official documentation.

PlatformWhere its sources come fromWhat you control
Google AI Overviews and AI ModeGoogle's index. Google states that a page only needs to be indexed and eligible to be shown with a snippet in Search; there are no additional technical requirements.Indexing, content quality, snippet controls (nosnippet, max-snippet, data-nosnippet).
GeminiCombines the model with Google Search when it needs up-to-date information.As for Google; the Google-Extended token in robots.txt governs the use of your content for training and grounding Gemini models, but does not affect Search.
ChatGPT with searchIts own crawling and index alongside external search providers, according to OpenAI's documentation.Allowing the OAI-SearchBot crawler (search), which is separate from GPTBot (training).
PerplexityReal-time web search with numbered citations in every answer.Allowing PerplexityBot; clear, up-to-date content.
Microsoft CopilotThe Bing index.Being indexed in Bing (Bing Webmaster Tools, IndexNow).

Two ideas follow from the table. First, presence in Bing matters more than many businesses assume, because it feeds several assistants directly or indirectly. Second, blocking AI crawlers in robots.txt is a legitimate choice, but you need to distinguish between crawlers used to train models and crawlers used to search and cite. Blocking the latter means giving up on appearing as a source.

What makes a page get cited

There is no official list of "citation factors". What does exist is the logic of the process described above, Google's public documentation on helpful content, and the practical experience of observing answers systematically. Together they point to a set of sensible principles:

Be retrievable

  • The page is indexed, its main content does not depend on JavaScript to load, and it is not blocked for the relevant crawlers.
  • It ranks reasonably well in organic search for the queries, and sub-queries, that matter to you. If you are not among the results the system retrieves, the quality of your writing is irrelevant.

Be easy to extract

  • Each section answers a specific question, with the answer in the first sentences and the detail afterwards.
  • Headings describe the content ("How long a migration takes" rather than "Timelines").
  • Comparable data sits in tables or lists, not buried in long paragraphs.
  • Definitions stand on their own: a passage lifted out of context still makes sense.

Contribute something of your own

If ten pages say the same thing, any of them will do as a source and yours is interchangeable. What sets you apart is information others do not have: your own data, detailed procedures, the real terms of your offer, first-hand experience, honest comparisons. It is the same principle Google sums up in its concept of E-E-A-T.

Be endorsed beyond your own site

These systems tend to trust brands that appear consistently in third-party sources: media, industry directories, associations, forums, reviews. Digital PR and quality mentions influence both classic rankings and the likelihood that a model recognises you as a reference.

Practical steps: a seven-phase work plan

  1. Define your query set. List the questions your customers ask before buying: comparisons ("best X for Y"), process questions ("how does it work", "how long does it take"), local questions and questions about your brand. Group them by intent and funnel stage.
  2. Establish a baseline. Run those queries on each relevant platform and record whether you appear, how you are described, which competitors appear and which URLs are cited. Use a fixed protocol (same wording, language, location, a session without history), because answers vary.
  3. Check the technical foundations. Indexing in Google and Bing, robots.txt rules for AI search crawlers, rendering of the main content, valid structured data, speed. Any blocker here undermines everything else.
  4. Tidy up your entity. Identical name, description, location, services and contact details on your website, your Google Business Profile, LinkedIn and the directories where you are listed. Organization markup with sameAs links to your official profiles.
  5. Rewrite key pages so they answer. Start with pages that already rank or that assistants already cite for your competitors. Add direct answers, tables, definitions and genuine FAQs; strip out filler.
  6. Create the missing content. If your baseline shows questions without a good source, there is an opportunity. A solid content cluster covers the main question and its sub-questions, which is exactly what query fan-out looks for.
  7. Earn external mentions. Be present where your industry gets its information: specialist media, independent comparisons, associations, events, review platforms. No paid links and no undisclosed sponsored content.

How to measure visibility in AI search

Measurement is the least mature part of GEO, and it is worth accepting that from the start so you do not make decisions the data cannot support. These are the sources available:

SourceWhat it tells youLimitations
Google Search ConsoleImpressions and clicks in Search, including AI features.AI Overviews and AI Mode data are included in the "Web" search type; at the time of writing they cannot be isolated.
Google Analytics 4Visits referred from assistant domains (e.g. chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com).Some traffic arrives without a referrer and counts as direct; you only see people who click.
Server logsWhich pages AI crawlers visit and how often.A crawl is not a citation; requires access to logs.
Answer samplingPresence, mentions, cited sources and how your brand is described.Answers are non-deterministic; you need to repeat and average.
Indirect signalsBranded searches, enquiries that mention "I saw you on ChatGPT".Approximate attribution.

In GA4, a useful first step is to create a custom channel group for "AI assistants" using a regular expression on session source. That separates this traffic from generic referrals and shows you which pages receive visits and whether they convert. We set this up as part of our analytics service.

For sampling, define a fixed panel of queries (from a few dozen to a few hundred, depending on the size of the business), run it at a consistent interval and record simple indicators: the share of queries where you appear, the share where you are cited with a link, your position relative to competitors, and how accurately you are described. Commercial AI tracking tools automate this, but it pays to understand their methodology before trusting a number.

Common mistakes

  • Treating GEO as separate from SEO. If the site does not rank or index properly, "AI optimisation" has nothing to build on.
  • Blocking every AI bot by default without distinguishing training from search, and then wondering why you never appear as a source.
  • Mass-producing AI content to "cover questions". Google treats content produced at scale primarily to manipulate rankings as abuse, however it is made, and assistants gain nothing by citing text that adds nothing new.
  • Relying on a single screenshot. Answers change between sessions, users and days. A one-off appearance is not a trend.
  • Treating llms.txt as the answer. It is an interesting proposal for giving models a summary of your site, but at the time of writing no major provider has confirmed using it to select sources. You can publish one; just do not expect it to change anything on its own.
  • Ignoring how your brand is described. Appearing with wrong information (services you do not offer, an old address) can be worse than not appearing. Fix the source of the error, which is usually an old page or an outdated directory listing.

What is still unknown

Anyone who claims to know "the ChatGPT algorithm" is selling smoke. There are open questions worth keeping in mind:

  • How each system weighs its sources. We know retrieval matters a great deal, but not the relative weight of domain authority, freshness or text structure.
  • How much traffic citations generate. Generative answers resolve many queries without a click. The value of a mention may lie more in brand exposure than in the visit, and that is hard to quantify.
  • How formats will evolve. Google's AI features, their availability by country and language, and the way links are displayed have changed several times; more changes are a reasonable expectation.
  • What data the platforms will provide. There is barely any official reporting on visibility in assistants today. If it arrives, it will change how we measure.

Given that uncertainty, the sensible strategy is to invest in what works under any scenario: genuinely useful content, clean technical foundations, a consistent brand entity and a reputation beyond your own website.

Conclusion

Appearing in ChatGPT, Perplexity or AI Overviews does not depend on a new trick; it depends on being a source that is retrievable, easy to cite and endorsed by third parties. Start with an honest baseline, fix the technical foundations, rewrite your key pages so they answer questions, and measure with the tools available while knowing what they cannot tell you. If you would like an assessment of your starting point, our SEO and AI search audit includes answer sampling across the main platforms. The technical terms in this article are defined in our glossary.

Are AI assistants citing you?

We analyse how you appear in Google, ChatGPT, Perplexity and Gemini for your key queries and deliver a prioritised plan.