Dev.to AI πŸ€– Ai πŸ‘ 0 πŸ“– 12 min read

How AI search picks fragments: query fan-out and RAG explained

Hi again πŸ‘‹ In my last post, How to become an SEO (AI) specialist in the coming 2027 #2, I wrote that without understanding the mechanics of AI search there is no "AI SEO", only guessing. A lot of you asked me to go deepe

Hi again πŸ‘‹ In my last post, How to become an SEO (AI) specialist in the coming 2027 #2, I wrote that without understanding the mechanics of AI search there is no "AI SEO", only guessing. A lot of you asked me to go deeper. So today we open the hood πŸ”§

In this post I will explain, step by step, what happens between the moment a user types a question into AI Mode, ChatGPT or Perplexity and the moment an answer with citations appears on the screen. And, most importantly, where in that process your content wins or loses.

  1. Why an SEO needs to understand RAG
  2. The big picture: from question to answer
  3. Step 1: query understanding and fan-out
  4. Step 2: retrieval
  5. Step 3: reranking and fragment selection
  6. Step 4: generation, grounding and citations
  7. Not every AI searches the same way
  8. What this means for SEO in practice
  9. Myths I keep seeing on LinkedIn

πŸ€” Why an SEO needs to understand RAG

RAG stands for Retrieval-Augmented Generation. In plain words: the model does not answer only from memory. First it searches, then it reads what it found, and only then writes an answer based on it.

Almost every AI search product works this way: Google AI Overviews and AI Mode, ChatGPT Search, Perplexity, Copilot, Gemini. The details differ, but the skeleton is the same.

Why does it matter for us? Because classic SEO optimized for one moment: ranking a URL for a keyword. RAG has four moments where your content can drop out:

  • the system never asks a question your page answers (fan-out),
  • your page is not retrieved for that question (retrieval),
  • your page is retrieved, but another fragment scores higher (reranking),
  • your fragment is read, but the model does not cite it (generation).

Ranking #1 helps with only some of these. That is why you can rank first and still be invisible in the AI answer 🀯

πŸ—ΊοΈ The big picture: from question to answer

Before we go into details, here is the whole pipeline in one place. Every AI search system has its own version, but they all follow roughly these steps:

  1. Query understanding. The model reads the question, the conversation history and the context (location, language, sometimes the user's past chats).
  2. Fan-out. One question becomes several or many sub-queries.
  3. Retrieval. Each sub-query goes to the index (web, Knowledge Graph, shopping, news, maps) and returns candidates.
  4. Chunking and reranking. Candidate pages are cut into fragments, and a stronger model picks the best ones.
  5. Generation. The LLM gets the question plus the selected fragments and writes the answer.
  6. Citation. The system attaches sources to the claims in the answer.

And here is what decides whether you win at each stage:

Stage What it looks for How your content wins
Fan-out the questions behind the question you cover the sub-topics users really care about
Retrieval documents that match a sub-query the page is indexed, accessible and textually relevant
Reranking the best fragment for a sub-query a paragraph that answers precisely, on its own
Generation facts the model can use clear, specific, verifiable statements
Citation a source that supports a claim your fragment is the one that actually contains the fact

Now let's go through it step by step πŸ‘‡

🌿 Step 1: query understanding and fan-out

Google states it directly in its Search Central documentation: both AI Overviews and AI Mode may use a "query fan-out" technique, issuing multiple related searches across subtopics and data sources to build a response.

In practice it looks like this. A user asks:

"Which running shoes are best for a beginner with flat feet who runs on asphalt?"

The model does not search for that sentence. It breaks it into sub-queries, for example:

  • running shoes for flat feet (stability vs neutral)
  • best beginner running shoes 2026
  • cushioning for asphalt running
  • reviews of specific models
  • how to choose running shoe size

Each of these goes to the index separately. The answer is then assembled from fragments that answered different sub-queries 🧩

A few things worth knowing:

  • Not every query fans out. A simple fact ("capital of Spain") does not need it. Complex, comparative and "it depends" questions fan out the most.
  • The sub-queries are hidden. You will not see them in Search Console, and no third-party tool has access to Google's internal ones. Tools that "show fan-out queries" are simulating them with their own LLM. Useful as inspiration, not as data.
  • Fan-out is not query rewriting. The original question still matters. The system searches it and the sub-queries.
  • Context changes the fan-out. Location, language, conversation history and personalization mean two users with the same question can trigger different sub-queries.

πŸ’‘ The SEO consequence: a page can be cited for a question it does not rank for, because it was the best answer to one of the sub-queries. And the opposite: your #1 page for the main keyword may not answer any sub-query precisely enough.

πŸ”Ž Step 2: retrieval

Retrieval is the moment the system goes to its index and brings back candidates for each sub-query. Two important facts first:

  • AI does not search the live web from scratch. It searches an index: Google's own, Bing's, or the AI company's own crawl. If you are not in that index, you do not exist. For Google's AI features the rule is simple: the page must be indexed and eligible to show with a snippet. Nothing more, nothing less.
  • What the bot could not read, the index does not have. Content loaded by JavaScript that a given bot does not render, content behind a login or blocked in robots.txt never reaches retrieval.

Chunking: pages become fragments

LLMs do not work on whole pages. Pages are split into chunks: paragraphs, sections or fixed-length pieces of text. Each chunk is scored separately. Your 4,000-word guide is, for the retrieval system, maybe 30 separate candidates competing with each other and with the rest of the web.

Lexical vs semantic search

There are two main ways to find matching chunks:

Method How it works Strength Weakness
Lexical (e.g. BM25) matches the words in the query to the words in the text precise for names, numbers, product codes misses synonyms and paraphrases
Semantic (embeddings) turns query and text into vectors and compares meaning finds an answer even with different wording can miss exact terms and rare names
Hybrid combines both scores the standard in modern systems β€”

An embedding is a list of numbers that represents the meaning of a text. Texts with similar meaning have similar vectors, so "cheap flights to Rome" and "budget airfare Italy capital" land close to each other, even though they share almost no words.

πŸ’‘ The SEO consequence: keywords did not die, they just stopped being enough. Exact terms (brand names, model numbers, prices, places) still help lexical matching. Clear, topical paragraphs help semantic matching. A paragraph that mixes three topics produces a "blurry" vector that matches nothing well.

πŸ† Step 3: reranking and fragment selection

Retrieval is fast but rough. It may return hundreds of candidate chunks per sub-query. The model generating the answer cannot read them all, because its context window is limited and every extra token costs money and time ⏱️

So there is a second, more precise stage: reranking. A stronger (and slower) model looks at the sub-query and each candidate fragment together and scores how well that specific fragment answers that specific question. Only the top few go further.

What typically helps a fragment survive reranking:

  • It answers the question directly, preferably in the first sentence.
  • It is self-contained. It makes sense without the paragraph above it. "As mentioned earlier, it costs 20% less" is a weak chunk, because what costs less?
  • It is specific. Numbers, names, dates, conditions. "It depends" without explaining on what loses.
  • It comes from a source the system trusts. Many systems mix relevance with quality and authority signals at this stage. Google has its whole classic ranking toolkit available here.
  • It is fresh, when freshness matters. For prices, versions and news, a newer fragment often wins.

Here is the difference in practice:

❌ Weak chunk βœ… Strong chunk
"There are many factors to consider and every runner is different, so it's best to try a few options." "Runners with flat feet usually do better in stability shoes, which limit inward rolling of the foot (overpronation). For a beginner on asphalt, choose a stability shoe with medium-to-high cushioning."

πŸ’‘ The SEO consequence: the unit of competition is now a paragraph, not a page. Your page can be great overall and still lose every individual "duel" with sharper fragments from other sites.

✍️ Step 4: generation, grounding and citations

Now the LLM gets a prompt that roughly says: here is the user's question, here are the selected fragments, write an answer based on them. This is called grounding: the answer is anchored in retrieved sources instead of only in the model's memory.

A few mechanics that matter for us:

1. The model synthesizes, it does not copy. It combines facts from several fragments into its own sentences. That is why your exact wording rarely appears in the answer, but your facts can.

2. Citations are attached to claims. The system links sentences in the answer to the fragments that support them. A fragment that contains a clear, checkable fact ("the battery lasts 30 hours") is easy to cite. A fragment full of opinions and fluff gives the model nothing to attach a citation to.

3. Read is not the same as cited. The model may read your fragment, use it as background and still cite someone else. Interfaces show a limited number of sources, so there is competition even at this last step.

4. Consensus wins. When many independent sources say the same thing, the model treats it as safe. When your fragment contradicts the majority, it needs strong support (data, a source, first-hand evidence) to be used.

5. Memory still leaks in. Grounding reduces hallucinations, but does not remove them. The model's training knowledge can still shape the answer, which is how outdated prices or old offers about your brand end up in AI answers ⚠️

πŸ’‘ The SEO consequence: write fragments that contain citable facts. Original data, specific numbers, clear definitions and first-hand experience are exactly what a model needs to justify a sentence, and what it cannot get from everyone else.

πŸ”€ Not every AI searches the same way

The pipeline is similar, but each system has its own index, its own bots and its own taste in sources. That is why your brand can look great in Perplexity and be missing in AI Overviews.

System Where it retrieves from What it means for you
Google AI Overviews / AI Mode Google's index (crawled by Googlebot), Knowledge Graph, Shopping, Maps classic Google SEO is the entry ticket; no extra files or markup needed
ChatGPT Search OpenAI's own crawler (OAI-SearchBot) plus partner search providers blocking OAI-SearchBot removes you from ChatGPT search results
Perplexity its own index (PerplexityBot) plus live fetching usually shows many sources, so niche pages have a real chance
Copilot Bing's index Bing Webmaster Tools and IndexNow are worth your time again

Memory vs live search

There is one more distinction many people miss. An assistant can answer in two modes:

  • From training data (memory). No search happens, no citations appear. The answer reflects what the internet said about you when the model was trained, sometimes a year or more ago.
  • With retrieval (search). The pipeline from this post runs, and citations appear.

The model often decides on its own whether to search. Questions about current prices, news, comparisons and "best X in 2026" usually trigger search. General questions ("what is CRM software?") are often answered from memory.

πŸ’‘ The SEO consequence: you need two strategies. Retrieval you influence with technical SEO and strong fragments, today. Memory you influence slowly, through a consistent picture of your brand across many sources that end up in future training data.

βœ… What this means for SEO in practice

Theory is nice, but here is the checklist I actually use. One block per pipeline stage πŸ‘‡

Fan-out: cover the questions behind the question

  • List the sub-questions a real user has around your topic: comparisons, costs, risks, "for whom", "how to choose", alternatives.
  • Check People Also Ask, Reddit threads, sales and support questions. They are the closest thing to real fan-out data you have.
  • Cover the sub-topics in one strong, well-structured page where it makes sense, not in 50 thin pages. Google warns that mass-producing pages for every query variation can fall under its scaled content abuse policy.

Retrieval: make sure you are in the index

  • Check robots.txt and your CDN/WAF for Googlebot, OAI-SearchBot, PerplexityBot and Bingbot. Verify in the server logs, not only in the file πŸ”
  • Make key content (descriptions, prices, specs, FAQs) available in the raw HTML, not only after JavaScript runs.
  • Use the words people use: product names, model numbers, locations. Lexical matching still counts.

Reranking: write fragments that win duels

  • Start each section with a direct answer, then expand.
  • Make each paragraph self-contained: one topic, no "as mentioned above".
  • Use descriptive headings that sound like the sub-question they answer.
  • Replace "it depends" with the conditions it depends on.

Generation: give the model citable facts

  • Add numbers, dates, definitions and conditions that can support a sentence.
  • Publish original data, tests and first-hand experience. That is the one thing other sites cannot copy.
  • Keep facts about your brand (prices, offer, locations) consistent and current everywhere: your site, Google Business Profile, comparison sites, reviews.

Measurement

  • Track your brand's citations for key questions in AI Overviews, ChatGPT and Perplexity. Treat the results as a trend, not a rank, because answers are non-deterministic.
  • Segment AI referral traffic in GA4 (chatgpt.com, perplexity.ai, gemini.google.com, copilot.microsoft.com).

πŸ™… Myths I keep seeing on LinkedIn

Myth Reality
"You need llms.txt or special AI markup to show up in AI Overviews." Google says no new machine-readable files, AI text files or special schema are needed. Indexing and eligibility for a snippet are the requirement.
"Create a separate page for every fan-out query." Fan-out queries are hidden, and mass-producing pages for query variations risks scaled content abuse. Cover sub-topics well instead.
"Rank #1 and you'll be cited." Ranking helps retrieval, but reranking and generation work on fragments. #1 pages are regularly not cited.
"This tool shows Google's fan-out queries." It shows queries generated by the tool's own LLM. Useful for ideas, not a data source.
"Short chunks of 40–60 words always win." There is no universal ideal length. A fragment wins by being complete and precise, not by a word count.
"AI answers from memory, so SEO doesn't matter." Answers with citations come from retrieval, and retrieval runs on indexes built by crawlers. Technical SEO is the entry ticket.

🧾 Summary

AI search is not magic. It is a pipeline: fan-out β†’ retrieval β†’ reranking β†’ generation β†’ citation. At every stage your content can win or drop out, and each stage rewards something slightly different:

  • fan-out rewards covering the real questions behind the query,
  • retrieval rewards being indexed, accessible and textually relevant,
  • reranking rewards precise, self-contained paragraphs,
  • generation rewards specific, citable facts and a consistent brand picture.

The good news? None of this replaces SEO fundamentals. It sits on top of them. The specialists who understand the pipeline stop guessing and start testing πŸ§ͺ

In my next post (#3) I will show step by step how to measure a brand's visibility in AI answers and how to turn it into a client report.

About the author ✍️

Łukasz, SEO AI Specialist at Afterweb. He has worked in SEO for years and today focuses on brand visibility in Google and in AI answers: from technical SEO, through data analysis and automation, to content and brand-mention strategy. He shares what he knows in the "How to become an SEO specialist" series, because he believes the best way to learn is practice and testing on real data.

Want to talk about SEO and your brand's visibility in AI? Visit afterweb.pl πŸš€

Sources

πŸ“° Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes β€” full credit and traffic to the original publisher.