An AI answer typically names three to eight sources. Behind each of those numbered links sits a selection process that most publishers never see, and about which the platforms themselves disclose only partial detail.
This article sets out what is documented, what is consistently observed, and where the honest answer is that nobody outside these companies knows.
That distinction matters more here than in most SEO topics, because the gap has been filled with confident claims about ranking factors that do not exist.
What an AI Citation Actually Is
A citation is an attribution attached to a generated answer, indicating that a particular source contributed to a particular claim. It is not a ranking. There is no first position in an AI answer in any meaningful sense, and the order in which sources appear generally reflects the order of the claims they support rather than a judgment about which site is best.
It is also worth separating two things that get conflated. A citation is a linked source. A mention is your brand being named in the text, with or without a link. A recommendation that names your company without linking to your site is commercially valuable and completely invisible in your analytics. Both belong in any assessment of generative engine optimization performance.
Retrieval Happens Before Selection
The single most useful thing to understand is that selection operates on a shortlist. Before any model decides what to cite, a retrieval system has already narrowed the entire web down to a few dozen candidate documents. If you are not in that candidate pool, nothing about your content quality is relevant.
How the pool is built differs by platform, which is why the same question yields different sources across services.
Google’s documentation states that its generative AI features are rooted in its core Search ranking and quality systems, and that retrieval-augmented generation relies on those systems to pull relevant pages from the Search index. It also sets out an explicit eligibility requirement: a page must be indexed and eligible to be shown in Google Search with a snippet.
That has a blunt implication. If you use a noindex tag, or a nosnippet directive, or your page simply is not indexed, you are outside the pool regardless of how good the content is.
ChatGPT
OpenAI’s publisher guidance states that any public website can appear in ChatGPT search, and that content must not be blocked to OAI-SearchBot in robots.txt to be included in summaries and snippets. OpenAI also states that ChatGPT ranks search results using multiple factors intended to surface relevant, reliable information, and that placement is not guaranteed. It does not publish those factors.
ChatGPT search also draws on external search providers rather than an index OpenAI built alone, which is why visibility in other search engines can matter more than many UK businesses expect.
Perplexity
Perplexity performs real-time searches and returns numbered footnotes linking to the sources used, with an interface that distinguishes between sources it selected and sources it reviewed. It is comparatively transparent at the output level while disclosing little about the selection mechanics.
The Factors That Appear to Matter
None of what follows is a published ranking factor. These are the characteristics that repeatedly separate cited pages from uncited ones, drawn from platform guidance and observed behaviour.
Relevance to the Expanded Query
Because systems such as Google’s fan queries out into multiple related searches, relevance is judged against sub-questions you may not have targeted. A page about commercial lease negotiation may be retrieved for a fan-out query about break clauses even though that phrase never appeared in your keyword research.
The practical reading is that breadth of genuine coverage within a topic improves retrieval odds more than exact phrase matching does. Google’s guidance reinforces this, noting that its systems understand relevance without exact query-to-page matching.
Extractability of a Specific Claim
Generated answers are assembled from claims. A page that contains a clean, self-contained statement of fact gives the system something to lift and attribute. A page that expresses the same information across three hedged paragraphs does not.
Compare these two treatments of the same information:
Weak: “There are various considerations when thinking about how long a rebrand might take, and it really depends on a number of factors specific to your organisation.”
Stronger: “A full rebrand for a mid-sized UK business typically runs twelve to twenty weeks: four to six weeks for strategy and research, six to eight for identity design, and two to six for rollout across digital and print. Multi-territory rebrands with legal trademark clearance usually add another two months.”
The second version is not longer for the sake of it. It contains numbers, stages and a named condition, all of which can be attributed. Note that this is about clarity rather than formatting, and Google’s guidance is explicit that there is no requirement to chop content into small pieces for AI systems.
Corroboration Across Sources
Systems generating factual answers appear to favour claims that are supported in more than one place. A figure that appears only on your website, contradicted by everything else available, is a risk for the answer rather than an asset.
This has a counterintuitive consequence. Original research is valuable, but original research that is subsequently referenced elsewhere is considerably more valuable, because the corroboration strengthens the claim. Publishing data and then doing nothing to circulate it wastes most of its potential.
Credibility Signals
Author identification, organisational transparency, editorial standards and evidence of genuine expertise all contribute. We treat this properly in our article on what makes a website trustworthy to AI search engines, since it is a large topic in its own right.
Freshness, Where the Topic Warrants It
Currency matters on subjects that change and matters very little on subjects that do not. A guide to a certification syllabus needs to reflect the current syllabus. A guide to the principles of typography does not need an annual refresh to remain accurate. Updating the date without updating the substance is not a signal, it is decoration.
Entity Recognition
Systems need to work out that “BrandingX”, “BrandingX UK” and “brandingx.co.uk” refer to one organisation, and what that organisation does. Inconsistent naming, addresses and descriptions across your website, Companies House, LinkedIn and directories make that harder. This is unglamorous data hygiene with a disproportionate effect on whether you are recognised as an entity worth citing.
First-Party Versus Third-Party Information
These are not weighted equally, and the split is predictable. For factual details about your own business, such as pricing, specifications, opening hours and service scope, your own site is the natural authority. For evaluative claims about whether you are any good, third-party sources carry the weight. No AI system is likely to cite your homepage as evidence that you are the best agency in Bristol.
This is why review platforms, industry directories, trade press and genuine community discussion matter to AI visibility. It is also why manufacturing those mentions is a poor idea. Google’s guidance names the pursuit of inauthentic mentions as one of the tactics site owners can ignore, noting that its spam systems and quality systems both feed its generative features.
Why a Top-Ranking Page Is Sometimes Not Cited
This is the question publishers ask most often, and there are several plausible explanations.
The answer needed a different claim. Your page may rank first for the head term while the system was resolving a fan-out sub-question your page does not address.
The information was harder to extract. A page that ranks well because of links, brand strength and engagement may still be poorly structured for extraction.
Another source stated it more plainly. A model assembling an answer needs a supporting passage, not the best page.
You are outside that platform’s pool. A page can dominate Google and be invisible in ChatGPT because of crawler access or indexing elsewhere. This is the most commonly missed explanation and the easiest to check.
Snippet restrictions are in place. Directives such as nosnippet or restrictive max-snippet values limit what can be displayed and, on Google, affect eligibility for generative features.
Why Some Websites Appear Repeatedly
Certain domains show up across an unusually wide range of AI answers. The pattern usually reflects a combination rather than a single advantage: broad topical coverage, strong entity recognition, consistent structure, frequent updating, and a large volume of third-party references pointing at them.
Community platforms are a distinctive case. Forums and discussion sites are frequently cited for questions involving experience, preference or comparison, because they contain first-hand accounts that publisher content often lacks. For a business, the lesson is not to spam those platforms but to recognise that genuine participation in the places where your category is discussed has a visibility dimension.
What Publishers Can Realistically Influence
Ordered by how much control you have:
- Access and eligibility. Fully within your control. Check robots.txt for accidental blocking of AI search crawlers, confirm indexing, and remove snippet restrictions on pages you want surfaced.
- Claim clarity. Within your control. Make the important facts specific, self-contained and stated where a reader would look for them.
- Topical depth. Within your control, over time. Cover the surrounding questions properly rather than publishing isolated posts.
- Originality. Within your control, if you are willing to share what you know. Google’s guidance is unusually direct that unique, non-commodity content matters more than any other suggestion it makes.
- Entity consistency. Mostly within your control. Audit how you are described everywhere and fix the discrepancies.
- Third-party presence. Influenceable, not controllable. Earn coverage, encourage genuine reviews, participate honestly in your industry’s public conversations.
Our article on how interior design businesses get cited by AI search engines shows this hierarchy applied to a single sector, where the third-party layer turned out to matter more than most firms expected.
What Nobody Can Tell You
There is no published weighting of these factors. There is no universal algorithm shared across AI platforms, and describing one would be misleading, since Google’s approach is rooted in its Search index while OpenAI combines external providers with its own crawler and Perplexity runs its own retrieval.
Nor is there any tool with access to internal ranking systems. Google’s own guidance warns site owners to be wary of third-party tools claiming to use internal Google metrics, on the straightforward basis that none of them have access. Visibility tracking tools are useful for spotting patterns across many prompts. They are not reading the machine.
Where This Leaves You
Citation selection is a two-stage process, and most businesses lose at the first stage without realising it. Retrieval eligibility is technical and checkable. Claim quality is editorial and improvable. Third-party credibility is slow and cannot be bought without risk.
Work in that order. There is little point refining your phrasing for extractability on pages that a platform’s crawler cannot reach.
Frequently Asked Questions
Do AI search engines all use the same method to pick sources?
No. Google’s generative features draw on its own Search index and ranking systems. ChatGPT combines external search providers with its OAI-SearchBot crawler. Perplexity runs its own real-time retrieval. Assuming one universal algorithm leads to poor decisions, particularly around crawler permissions.
Does structured data help me get cited?
Google states that structured data is not required for its generative AI features and that no special schema is needed, while still recommending it for rich results in Search. Treat it as good practice with benefits elsewhere rather than as a citation lever.
Why does my competitor get cited when their content is worse than mine?
Usually because their claim was easier to extract, better corroborated elsewhere, or because they are inside a retrieval pool you are outside of. Check your crawler access and indexing before concluding it is a content quality problem.
If an AI answer mentions my brand but does not link to it, does that count?
Commercially, yes, often more than a link does, since the reader receives a recommendation. It will not appear in your analytics, which is why prompt-based monitoring is necessary alongside referral tracking.
Can I ask to be cited or submit my site to an AI search engine?
There is no submission process comparable to an XML sitemap for AI answers. The route in is the same as for search: be crawlable, indexed where the platform retrieves from, and worth using as a source.
Published by BrandingX UK.