Home/ Blog /GEO

Is AI Search Volume Data Reliable?

Turan Doğan
Turan Doğan
SEO & GEO Specialist
GEO September 11, 2026 10 min read
Is AI Search Volume Data Reliable?
SUMMARY
AI search volume data is a third-party reconstruction of prompt demand that chat platforms don't share. These estimates are produced through panel sampling, derivation from classic search volume or synthetic variation generation. The right way to use them is to rank topics against each other and tie budget decisions to measured data.

Where does AI search volume data come from?

None of the numbers sold as AI search volume come from the chat platforms' own records. ChatGPT, Gemini, Claude and Perplexity don't publish the prompts users write at the query level, don't open them to third parties and don't release them through an interface like Google Ads. Ahrefs says this plainly when describing its own calculation method: none of the major AI platforms share query data with anyone. Semrush acknowledges the same limit and states that no platform can provide exact numbers on visibility.

What platforms publish is not queries but aggregated behavioral research. OpenAI's usage research is a typical example. A representative sample of conversations is run through automated classifiers, nobody on the research team sees message content, and only category-level shares are released. A study like this shows what people use the chat interface for. It doesn't show how many times a particular keyword is asked each month.

When you see a four-digit number labeled monthly prompt volume on a tool's screen, what you have is not a measurement but a reconstruction. The right question is not whether the number is correct. The right question is: from which inputs, and under which assumption, was this number produced? This distinction affects almost every budget decision made on the generative engine optimization side.

How are these numbers produced?

The data on the market rests on three very different approaches. Showing up the same way on the same screen doesn't make them the same kind of information.

Opt-in panel measurement and scaling. Some providers collect real conversation text from double opt-in consumer panels. Profound says it uses this approach and scales the raw panel data with statistical modeling that corrects for demographic and geographic skew. There is a real measurement here, but it is a measurement on a sample. The final number is still a product of modeling.

Derivation from classic search volume. Ahrefs finds the parent keyword with the highest Google volume for each prompt, then applies a platform-specific ratio to that volume. The ratio comes from comparing AI platform traffic with Google organic traffic in its own traffic dataset. Ahrefs also writes that this metric can't be used to estimate the real size of the AI search market. The method is described honestly, but what it measures is Google demand, not AI demand.

Synthetic variation generation. This is the weakest link. Some tools derive algorithmic variants from a seed keyword and assign a volume directly to those variants. There is no real interaction behind them at all. The number on the screen looks identical to one that comes from panel data, but as evidence it isn't in the same class.

Method What it actually sees Blind spot
Opt-in panel A sample of real conversation text Small sample share; mobile app and API usage may be out of scope; limited country coverage
Derivation from search volume Demand on the Google side Doesn't measure AI demand directly; the platform ratio rests on assumptions
Synthetic variation No real interaction The number has no verifiable basis

If a tool has no methodology page, you can't know which row it belongs to. A number from a provider that doesn't publish its methodology is, at best, a number that can't be classified.

Classic search volume was an estimate too

Some of the criticism aimed at AI volume data rests on the assumption that classic search volume is an exact measurement. That assumption is wrong. The average monthly searches in Google Ads Keyword Planner is an average calculated based on the selected date range, location and network settings. Google says in its own documentation that these statistics are rounded, and adds a warning that when you pull volume for more than one location, the numbers may not add up as expected. Historical data is also shown only for exact match and close variants. In other words, the reference number the industry has used for years is a band as well.

The difference between them still matters. The Keyword Planner figure is a rounded number derived from the records of the company that owns the data. The AI volume figure is a reconstruction made from indirect signals by a third party with no access to the records at all. Both are estimates, but they don't carry the same level of uncertainty. On the classic side you know the width of the band; on the AI side you don't.

In practice, this doesn't mean throwing classic volume away. On the contrary, the cheapest way to tell whether there is demand for a topic on the AI side is still to look at classic volume, because the derivation method draws on the same source anyway. Starting your assessment of a topic with a free keyword analysis and putting AI volume data on top as a second layer is a sturdier setup than giving the two numbers equal weight.

Why can't a prompt be counted like a keyword?

The problem isn't just data access. The object you want to measure isn't suited to counting in the first place. A query typed into a search box is short and repeatable; the same phrase is typed word for word by thousands of people. A prompt typed into a chat interface is long, carries personal constraints and gains context with the next message. When the same information need is expressed in hundreds of different sentences, the monthly repetition count of a single sentence stops being a meaningful quantity.

The second layer is even more decisive. Google writes in its own documentation that AI Overviews and AI Mode can use a query fan-out technique, meaning they can run multiple related searches across subtopics and different data sources for a single question. The queries the system generates in the background are not the sentence the user typed. So even what can be measured on the search side isn't the prompt itself.

Serious providers see this problem and aggregate at the topic level instead of the single prompt. Semrush explains that individual prompts are too specific and unique to be measured directly, so it groups prompts that move in the same semantic direction under a topic. Moving up to the topic level is an honest solution, but it has a cost: the number you have is no longer the volume of a particular sentence but of the cluster the provider built. You didn't make the clustering decision.

Which decisions is this data useful for?

The data being an estimate doesn't mean it's useless. Ranking information is far more durable than absolute numbers. Of two numbers produced with the same method, which one is larger is usually right; what either of them actually is usually isn't.

  • Topic prioritization. Ranking ten topics against each other within the same tool is enough to decide which one to write first.
  • Trend direction. A topic rising or falling over months is a meaningful signal as long as the methodology stays the same.
  • Intent separation. Prompts that include a brand name and open-ended category prompts behave differently and call for different content. Volume data helps you see this distinction.
  • Platform breakdown. Knowing which interface a topic comes up in most tells you which platform to test on.
  • Gap detection. Topics with low classic volume but visible demand on the AI side are gaps competitors usually skip.

When does it mislead?

Where the same data does harm is fairly clear.

  • Tying the budget to a single number. Multiplying monthly volume by a conversion rate to produce a revenue projection means treating a modeled estimate as if it were a measured baseline.
  • Comparing numbers from two different tools. Because their methods differ, the gap between them shows the difference between two models, not reality.
  • Adding it to Google volume. In a tool that uses the derivation method, AI volume is already calculated from Google volume. Adding the two means counting the same demand twice.
  • Expecting precision in the long tail. The error in data scaled from a sample grows for niche and branded prompts. Small numbers are the least reliable.
  • Presenting exact figures to clients. The sentence "this many people ask this every month" is a claim the data can't support.
  • Mistaking a methodology change for a trend. When a provider updates its model, past numbers may change too. A break in the chart may not be a change in behavior.

Five questions to ask the provider

When evaluating a tool, every question that goes unanswered reduces how much weight that number should carry in your decision.

  1. What source is the number produced from: an opt-in panel, derivation from classic volume, or algorithmic variation?
  2. If there is a panel, which devices, apps and countries does it cover, and which does it not?
  3. Is the number the volume of a single prompt or of a topic cluster the provider built?
  4. How often is the data updated, and what period does it cover?
  5. When the methodology changes, is historical data recalculated, or does the series break?

Cross-checking against measured data

The antidote to estimated data is not a better estimate but data that is actually measured. You have three sources that complement each other.

The first is official data on the search side. The generative AI performance report in Google Search Console shows impressions and clicks from the AI Overviews and AI Mode surfaces. The report supports page, country, date and device breakdowns. There is no query breakdown, so you can't see which prompt triggered it. Even so, these numbers are measured, not modeled, and once a Search Console data tracking discipline is in place, they form a natural reference for testing volume estimates.

The second is referral traffic in analytics. Sessions coming to the site from chat interfaces are few, but they are real. If the estimated volume for a topic looks high and not a single session has come from that surface in months, there is a gap between estimate and reality that needs explaining.

The third is the prompt tests you run yourself. You need to repeat the same question in different wordings, at different times and across more than one engine, because these systems' output is not deterministic. Drawing conclusions from a single run is misleading. Regularly repeated AI visibility measurement shows which topics you are actually mentioned in far more accurately than a volume estimate.

The decision rule follows from this. AI volume data goes into prioritization; measured data goes into decisions. A topic can be added to the content plan because it ranks high in estimated volume. If you're going to put extra budget behind that topic, the justification should be Search Console impressions, referral traffic or a repeated mention test.

Frequently Asked Questions

Is AI search volume data completely worthless?

No. It is useful for relative ranking, trend direction and intent separation. It becomes worthless when the absolute number is used as an input to a calculation.

If two tools give different numbers for the same topic, which one is right?

Probably neither is exactly right. The difference usually comes from methodology. The right approach is to pick one and stick with it consistently, not mix the two and not add their numbers together.

Does Google Search Console show which prompt was used?

No. The generative AI report supports page, country, date and device dimensions, but not the query dimension. AI Mode data is also aggregated within the web search type in the performance report and can't be filtered separately.

Which metric should you track instead of prompt volume?

Mentions and citations should be tracked separately. A source may be cited without the brand name appearing in the answer, or the brand name may appear without a link being given. Reducing these to a single metric is the same mistake as trusting a volume estimate.

Was this article helpful?
Add Seobaz as a preferred source on Google to see us more often in your search results and AI answers.
Add as preferred source
Share this article
Turan Doğan
Founder · SEO & GEO Specialist
Publishing up-to-date guides on SEO, GEO and AEO since 2014, helping brands get seen on both Google and AI engines.
WhatsApp Online · Quick reply
Gift Wheel A discount on every spin
View Cart