Ask ten consultants how to rank inside AI Overviews and you will get ten frameworks, four acronyms, and at least one person telling you to publish an llms.txt file.
The advice market is global, and it moves quickly. A large share of the world’s search work is executed offshore, so a playbook that trends on LinkedIn on Monday is being applied to live client accounts by an agency in London and by an SEO consultant in India before Friday. Scale is what makes bad advice expensive.
So it is worth asking the blunt question. What do we actually know about how AI search picks its sources, and what are we guessing?
More is documented than the discourse suggests. And if you have any background in machine learning, that documentation is easier to read than most SEO material, because the mechanism it describes is one you already know.
The mechanism is retrieval, not magic
Google’s official guidance on optimizing for generative AI search, last updated at the end of June 2026, names two techniques sitting behind AI Overviews and AI Mode.
The first is retrieval-augmented generation, which Google also calls grounding. Core Search ranking systems retrieve relevant pages from the Search index, and the model generates a response over the information in those retrieved pages, with clickable links back to the sources that support it.
The second is query fan-out. Rather than treating your question as a single lookup, the model issues a set of concurrent, related queries. Google gives its own example: someone asks how to fix a lawn full of weeds, and the system also fires off searches about herbicides, removing weeds without chemicals, and preventing weeds in future.
None of this is exotic. It is a RAG pipeline with a query expansion step, running on an index that already existed.
That last clause is the one that matters, and it carries a consequence most AI SEO products quietly need you to miss. There is no separate AI index. To be eligible as a supporting link in AI Overviews or AI Mode, a page has to be indexed and eligible to be shown in Google Search with a snippet. That is the entire entry requirement. Google is explicit that there are no additional technical requirements.
If retrieval draws from the same index, and candidates are ordered by the same core ranking systems, then the set of things that can influence it is largely the set of things that were already called SEO.
The list of things Google says you can skip

The guide also includes a mythbusting section that is unusually blunt for Google. A significant slice of the AI SEO industry is currently selling work that Google states has no effect on Google Search. In summary:
- llms.txt and other special files. Google Search does not use them. Publishing one will neither help nor harm your visibility. Maintain one for other systems if you like, but Google ignores it.
- Chunking your content. There is no requirement to slice pages into small fragments so a model can understand them. Google states there is no ideal page length.
- Rewriting content for AI systems. You do not need a separate register for machines. The systems handle synonyms and general meaning, so you are not obliged to capture every phrasing of a question.
- Chasing mentions. Manufacturing inauthentic mentions across the web is less useful than it appears, and the spam systems are part of the same stack.
- Over-investing in structured data. It is not required for generative AI search and there is no special schema for it. Keep using it for rich results, which is what it was built for.
There is also a warning worth reading twice. Spinning up a separate page for every fan-out variation you can imagine falls under scaled content abuse. A tactic that appears in several GEO playbooks is a tactic Google classifies as spam.
Why folklore gets expensive at scale
That matters more than it first looks, because of how search work actually gets done.
Execution is a global supply chain. Agencies in the US and Europe sell the strategy, and a great deal of the implementation is carried out by delivery teams in India and elsewhere. A tactic that trends in March reaches a delivery playbook by May and is running on thousands of client sites by July.
Folklore does not stay theoretical. It gets operationalized, at volume, on other people’s websites. Where the tactic is merely useless, the cost is wasted hours. Where it crosses into scaled content abuse, the cost is somebody’s rankings.
Where the guidance stops being useful
Now the part the Google guide will not tell you, because it is not Google’s job to.
The document is scoped to Google Search. ChatGPT, Perplexity, Claude, and Copilot each run their own retrieval stacks, with different crawlers, different corpora, and different selection logic. Some of them read files that Google ignores. Some lean heavily on a small set of sources they have learned to trust. Their behavior is undocumented, unstable, and changes without notice.
So “Google says llms.txt does nothing” is a strong claim about Google and a much weaker claim about everything else.
The honest framing is that you are trying to be retrieved by several systems that use different indexes and different rankers, and exactly one of them publishes documentation. Everything you read about the others is inference from observed outputs. Some of that inference is careful. Most of it is not.
For a technical reader the smell test is simple. Ask whether a claim about AI search is falsifiable, and whether anyone has bothered trying to falsify it.
What the retrieval model actually implies
Strip out the vocabulary and the practical implications are unglamorous.
- Be retrievable before anything else. If a page is not crawlable, not indexed, or not snippet eligible, no content strategy will save it. This is a precondition, not an optimization.
- Stop writing interchangeable content. Google draws a hard line between commodity content, such as a generic tips listicle, and content built on first-hand experience. Read that in retrieval terms. If your passage is substitutable for ten others in the candidate set, the ranker has no reason to prefer it and the generator has no reason to cite it. Uniqueness here is not a branding concern. It is a selection concern.
- Answer the sub-questions inside the page. Fan-out means the useful unit is closer to a passage than a page. Cover adjacent questions properly, in one place, rather than publishing a thin page for each variation.
- Keep the technical surface dull. Crawlability, JavaScript rendering, duplicate content, page experience. None of it is new, and all of it still gates retrieval.
- Take agents seriously if they touch your business. Browser agents inspect the DOM, read visual renderings, and parse the accessibility tree in order to complete tasks. Semantic HTML and accessibility work, long treated as a compliance chore, now has a second payoff. Protocols for agent-driven commerce are already appearing.
Measuring any of this is genuinely hard
Worth being honest about the hardest part of the job.
A citation in an AI answer is not a ranking position. It is a sampled, non-deterministic output that varies with phrasing, user, region, and model version. Run the same prompt twice and you can get two different source sets. That breaks the instinct most search teams have, which is to track a number daily and watch it move.
A few things help.
- Treat vendor AI visibility scores as proxies with unknown error bars, not as metrics.
- Sample across many prompt variations instead of tracking one query. The distribution is the signal.
- Use Search Console as ground truth for Google surfaces, because that data is measured rather than inferred.
- Expect fewer clicks per impression, and judge the clicks you do get on engagement and conversion rather than volume.
If a vendor shows you a clean upward line for AI visibility, ask how it was computed. Non-deterministic systems rarely produce clean lines.
How to evaluate anyone selling you AI search work

Google maintains separate guidance on assessing third-party SEO advice, and one point from it is worth carrying into every vendor call: no third-party tool has access to Google’s internal ranking or AI systems. Any product implying otherwise is describing a proxy, not a measurement.
Four questions cut through most sales decks.
- Which retrieval system are you optimizing for, and how do you know it behaves the way you claim?
- What is the measurement baseline, and how do you separate AI surface visibility from ordinary organic movement?
- Which of your recommendations come from published documentation and which are hypotheses? Label each one.
- Show me a page that was cited, and the specific change that preceded the citation.
Whether you build the capability in-house, retain an agency, or hire an offshore team, the questions do not change. Anyone who cannot answer them is selling vocabulary.
The irony nobody in the industry wants to name
The organizations best placed to settle these arguments empirically are the large delivery teams, a great many of them in India, that run hundreds of client accounts at once.
They have the sample size a solo consultant will never have. They have account access across industries, geographies and site architectures. They could hold a control group, run a genuine holdout test, and publish the result. That single act would move this field from folklore towards evidence faster than any amount of commentary.
Almost none of them do it. The industry is sitting on the data and shipping the guesswork instead.
The unsatisfying conclusion
AI search looks like a new channel. Underneath, it is a retrieval system reading an index that has existed for years, ordered by ranking systems that have existed for years.
The lever has not moved. Content worth retrieving, on a site a crawler can reach, written by someone who actually knows the subject. That was the answer before the acronyms arrived, and Google’s own documentation now says it is still the answer.
Generative engine optimization will become a real discipline on the day it starts producing falsifiable claims. Until then, most of what is sold under that name is a glossary with an invoice attached.