Home/ Blog /GEO

AI and SEO Glossary

Turan Doğan
Turan Doğan
SEO & GEO Specialist
GEO April 7, 2026 14 min read
AI and SEO Glossary
SUMMARY
AI search did not replace the classic SEO vocabulary; it added a new layer on top of it. This glossary groups 38 terms such as GEO, RAG, query fan-out, grounding and citation rate into five sections and gives each definition in a form that can be read on its own. Most of the terms are links in the same chain: crawlability, retrieval, passage selection, citation and brand mention.

Anyone who has read the sentence "retrieval rate is good but citation rate is low" in a visibility report and got stuck needs the same thing: to know what the term means, which term it gets confused with and what it measures. The 38 terms below are grouped into five sections; each definition can be read on its own, and there is no need to read from start to finish. The groups follow the order in which generative search works: first how the model works, then search surfaces, the content side, measurement and, finally, technical access.

Generative AI search fundamentals

LLM (Large Language Model)

An LLM is an AI model trained on very large bodies of text that generates the continuation of a text by calculating probabilities. It is the generation layer behind ChatGPT, Gemini and Claude. The critical point is this: a language model on its own is not a search engine; it cannot know current information that is not in its training data, and it can only use that information when it is brought in from outside.

Prompt

A prompt is the instruction a user or system gives to the model. Its effect on search is directly visible: while users type two or three words in classic search, a prompt gives the goal, context and constraints in the same sentence. That is why queries reaching answer engines are noticeably longer and more specific than classic queries.

Retrieval

Retrieval is when a system pulls relevant documents and passages from sources outside itself before generating an answer. It is not the same as ranking: a page ranking first in organic results may never enter the candidate pool, while a page further down the list may. The first concrete goal of AI visibility work is not the top spot but getting into this pool.

RAG (Retrieval-Augmented Generation)

RAG is the name of the architecture in which the model first fetches relevant sources and then writes the answer by looking at those sources. It stands for Retrieval-Augmented Generation. Almost all answer engines that show sources work on this logic, meaning the links under an answer are not the model's memory but documents fetched at that moment for that query.

Grounding

Grounding is the process of tying a generated answer to a source that can be shown. Google uses this term for the process of basing its own generative answers on search results. In a grounded answer, there is a traceable link between the model's claim and the source; in an ungrounded answer, that link cannot be established.

Hallucination

Hallucination is when a model generates information that does not exist in any source as if it were true. Non-existent study names, made-up statistics and quotes attributed to the wrong person are typical examples. Grounding and RAG reduce this risk but do not eliminate it, which is why sources that state numbers and definitions clearly are safer material for these systems.

Vectors and Embeddings

Embedding is the process of converting a text's meaning into a sequence of numbers (a vector). Two texts with similar meanings are positioned close to each other in this space, so the system can ask "does it say the same thing" instead of "does it contain the same word". Retrieval in answer engines relies heavily on this similarity calculation.

Semantic Search

Semantic search is a form of search that matches a query by meaning rather than by word matching. Its practical consequence is that a page can be selected as the answer to a query without ever using the target keyword in its text. This is the technical explanation for the declining return on sprinkling keywords through a text.

Context Window

The context window is the amount of text a model can process at once. The query, the retrieved sources and the previous conversation share this window, so the system does not read every page it finds from start to finish. From a long page, usually only a few chunks make it into the window, and the structure of the content determines which chunk gets in.

Answer engines and search surfaces

Answer Engine

An answer engine is a system that generates the answer to a question directly instead of sending the user to a list of links. ChatGPT, Perplexity, Gemini and Google's AI Overviews fit this definition. Thinking of them all as one behavior is misleading: in a Profound analysis of 3.25 billion citations, the sources chosen by ChatGPT and Perplexity overlapped by only 11%, meaning being a source in one engine does not make you one in another.

AI Overview (Google AI Overviews)

An AI Overview is the generative summary that appears at the very top of Google's search results page and is compiled from multiple sources. Source links are shown next to the summary, and these links do not have to be the same as the classic top ten results. Because it launched for Turkish queries on February 18, 2026, the measurement history for Turkish content is still short, and the first observations cover the period after that date. How to appear on this surface is covered separately in our Google AI Overview guide.

AI Mode

AI Mode is Google's conversational search tab. Instead of a classic results list, it offers answer text and source cards, and users can narrow the topic by asking follow-up questions in the same session. Its most fundamental difference from the classic results page is that it works in the background by splitting a single query into multiple sub-queries.

Query Fan-Out

Query fan-out is when a single user query is split in the background into many sub-queries that are searched in parallel, with the results combined into one answer. From a visibility perspective, the consequence is that a page may be selected as a source not for the main query but for one of its sub-queries. This is exactly where content that also covers the side questions of a topic gains its advantage.

Zero-Click Search

A zero-click search is when a user gets the answer on the results page and leaves without clicking on any site. Featured snippets and AI summaries amplify this behavior, so clicks can fall while impressions stay flat. In an environment where clicks decline, the weight of traffic shifts to non-search surfaces, and Google Discover is the most visible example of this.

GEO (Generative Engine Optimization)

GEO is the work of making content something generative answer engines can select and cite as a source. It stands for Generative Engine Optimization, and the term comes from an academic study. It does not replace classic SEO but builds on top of it: ranking visibility continues to be an input to the retrieval stage. How it works in detail is explained in our article What is GEO?

AEO (Answer Engine Optimization)

AEO is work focused on making a question answerable directly and definitively. In practice, it means meeting the question with a clear sentence at the top of the page, not leaving the definition vague and giving the answer along with its conditions. AEO and GEO look at the same goal from different sides: one is concerned with the clarity of the answer, the other with being selected as a source.

Content and citation terms

Citation

A citation is the source reference an answer engine shows within or below the text it generates. It is not synonymous with ranking: the page in first place may get no citation at all, while a definition further down the list may get one. It is the basic unit of studies that measure AI visibility, and writing content that gets cited is a separate practical topic.

Passage and Chunk

A chunk is a piece of text into which retrieval systems split a page, and a passage is the name for that piece when it is used in an answer. Systems evaluate a page not as a whole but through these pieces. That is why the unit of competition is the passage, not the page; a poorly structured section of a good page is left unselected.

Self-Contained Explanation

A self-contained explanation is a piece of text that keeps its meaning even when taken out of context. The sentence "This method is cheaper" says nothing on its own; the sentence "Local SEO work is cheaper than an advertising budget" does. Because answer engines take text in pieces and use it, this difference directly affects the chance of being cited.

Information Gain

Information gain is the amount of new information a page adds to existing sources on the same topic. By this measure, repeating the three points everyone writes in different words carries zero value. Content with its own data, tests, calculations or original comparisons stands out here.

Entity

An entity is a defined thing that search systems recognize and link to others: a brand, person, product, place or concept. Modern search works not through strings of words but through relationships between entities. That is why a brand name appearing in the same form and with the same description across every channel is more valuable than keyword repetition.

Knowledge Graph

A knowledge graph is a data structure made up of entities and the relationships between them. Links such as "this company operates in this industry and is headquartered in this city" are stored here. If a brand does not exist as a clear entity in this structure, it becomes harder for answer engines to mention it in the right context.

E-E-A-T

E-E-A-T stands for experience, expertise, authoritativeness and trustworthiness, and is an evaluation framework found in Google's quality rater guidelines. It is not a directly measured score, and there is no score you can see in a dashboard. In content, it shows up as author identity, the use of verifiable sources and traces of real experience.

Topical Authority

Topical authority is a site covering a specific subject area not with a single page but with an interconnected body of content. It is not a single metric but the sum of breadth of coverage and consistency. Sites that also answer the side questions of a topic become candidates more often in the sub-queries generated by query fan-out.

Brand Mention

A brand mention is when a brand is named in an answer without a link. It is measured separately from citations: a brand may be mentioned without a link, or its page may be shown as a source while its name never appears in the answer. Reducing the two to a single number distorts the visibility picture.

Earned Media

Earned media is content third parties produce about a brand through their own editorial decisions: news stories, reviews, list articles, forum discussions. In an analysis by Muck Rack, about 84% of citations in AI answers came from sources of this kind rather than from brand sites. It should not be confused with paid or manipulative link practices; they differ in both method and outcome.

Measurement terms

Citation Rate

Citation rate is the proportion of queries in a fixed test query set in which a page or domain is actually shown as a source. If the test set changes, the rate changes too, so measurements are not comparable until the set is fixed. Drawing conclusions from a single run is also a mistake, as the same prompt can select different sources at different times.

Retrieval Rate

Retrieval rate is how often a page is seen among candidate sources, even if it does not make it into the answer. A high retrieval rate alongside a low citation rate changes the diagnosis completely: the problem is not discoverability but selectability. Keeping the two rates separate determines whether the next job is content or technical access.

Share of Voice

Share of voice expresses how often a brand appears compared with competitors in the same query set. The reason for looking at share rather than absolute numbers is that the visibility space in answer engines is narrow; usually only a few sources are mentioned in an answer. Because the share shifts whenever the competitor list or the query set changes, keep both fixed.

Answer Influence

Answer influence measures whether information from a source is actually carried into the answer. A page can appear in the source list without contributing anything to the content of the answer. It answers a different question from whether a citation exists: whose information is the answer text conveying?

AI Referral Traffic

AI referral traffic is when a visitor comes to the site by clicking a link in an answer engine. In analytics, it is distinguished by referring domains such as chatgpt.com and perplexity.ai. Its volume remains small next to classic organic traffic, and citations and brand mentions do not show up as traffic at all, which is why measurement cannot be done by looking at this channel alone.

AI Visibility Test Set

An AI visibility test set is the fixed list of prompts used in measurement. It contains informational queries, comparison queries, recommendation queries and brand queries together; if they are all of the same type, the result is misleading. Because answers can change even for the same prompt, run critical queries several times, at different times.

Technical access terms

Crawlers and Bot Types

A crawler is software that visits sites and collects their content, but not all crawlers come for the same job. Search-side bots such as Googlebot and OAI-SearchBot crawl to find sources for answers, while bots such as GPTBot and Google-Extended concern model training and permission to use data. The distinction has practical consequences: disallowing a training bot does not directly end your chance of being cited, but blocking a search bot does.

robots.txt

robots.txt is a text file in the site's root directory that states which bot can crawl which paths. It is a notice, not a technical lock; compliance is voluntary, and major providers generally comply. Blocking a page here removes it from the source pool of both classic search and answer engines.

llms.txt

llms.txt is a file proposed as a way to present a site's content as a simplified list for language models. It has remained at the proposal stage: there is no verified finding that major search and answer engines use this file for citations, and Google has stated explicitly that it does not support it. Those who want to try it can generate the file with the llms.txt generator, but it should not be seen as a setup that earns citations.

Schema and Structured Data

Schema is a standard for marking up the entities on a page in a machine-readable way, usually added in JSON-LD format. In classic search, it has a real and measurable function for rich result displays. However, there is no strong evidence that adding structured data directly increases citations in AI answers, and lumping the two benefits together creates false expectations.

Indexability and Main Content Access

Indexability is a page's ability to be crawled and added to the search index. For answer engines, one more condition is added: the main content must exist as readable text in the page source. If the content is loaded only later by scripts running in the browser, the text may never reach the retrieval layer, even if the page looks fine to the human eye.

Snippet Controls

Snippet controls are directives that determine how much text from a page can be shown in search results: nosnippet, max-snippet and data-nosnippet used within the page. Google states that these controls also apply to AI Overviews. A restriction added to protect content can, without you realizing it, keep the page entirely out of generative answers.

The chain that links the terms together

Most of the concepts in this glossary are not independent headings but successive links in the same chain. The chain works in this order: the page must be crawlable and its main content readable (indexability, robots.txt, crawler types), it must enter the candidate pool for the query's sub-queries (retrieval, query fan-out), a chunk within it must be understandable on its own (chunk, self-contained explanation), it must be found worth selecting among the other candidates (information gain, earned media) and, finally, it must appear in the answer as a source or as a brand (citation, brand mention).

Finding where the chain breaks is the real job of measurement. If the retrieval rate is high but the citation rate is low, the problem is on the content side; if both are low, look first at technical access and brand awareness; if there are citations but the brand name never appears in the answer, what is missing is entity clarity. Work that addresses the entire chain together falls under the GEO heading.

Was this article helpful?
Add Seobaz as a preferred source on Google to see us more often in your search results and AI answers.
Add as preferred source
Share this article
Turan Doğan
Founder · SEO & GEO Specialist
Publishing up-to-date guides on SEO, GEO and AEO since 2014, helping brands get seen on both Google and AI engines.
WhatsApp Online · Quick reply
Gift Wheel A discount on every spin
View Cart