
Content Signals are a usage declaration added to robots.txt
Content Signals is a robots.txt extension that declares not whether a page may be crawled but how it may be used once it has been crawled. Cloudflare published it under a CC0 license, so anyone can implement it freely. It adds two pieces to the file: a comment block starting with # that defines the meaning of the signals in plain language, and a Content-Signal line that carries the preferences in machine-readable form.
The reason for this split is a limitation of robots.txt itself. The Robots Exclusion Protocol governs which paths a crawler may request. Whether content, once downloaded, goes into a search index, gets added to a model's training set or is used as input for a live AI answer is not a question robots.txt answers. Permission to access and permission to use are two different decisions, and classic robots.txt only expresses the first.
Cloudflare automatically serves this policy on domains with the managed robots.txt feature turned on. On a site where Cloudflare setup is complete, you turn the setting on from the dashboard, and no file is written to the server. A site that does not use Cloudflare can add the same text to its own robots.txt file manually, because the policy is freely licensed.
What each of the three signals means
The policy defines three categories, each corresponding to a different technical process. The policy text itself draws the boundaries explicitly, especially the one between search and AI summaries.
| Signal | What it covers | What it doesn't cover |
|---|---|---|
| search | Building a search index, returning links and short excerpts on the results page | AI-generated search summaries |
| ai-input | Feeding content to a model in real time, as in RAG and grounding | The training process that changes a model's weights |
| ai-train | Training a model or fine-tuning an existing model | Real-time content fetching while an answer is being generated |
The critical point is that the search definition contains a built-in exception. The search signal covers returning links and short excerpts, but not AI-generated summaries. So if a site owner wants to say yes to search and no to AI summaries, they express that through ai-input, not search.
In addition to these three, Cloudflare is also testing a fourth, optional field called use. This field describes how long content may be stored and reused after access and takes three values: immediate for real-time interaction only with nothing stored, reference for indexing, quoting and linking back, and full for summarizing and reproducing. Cloudflare's managed file serves reference as the default.
How are the signals written in robots.txt?
The machine-readable part is a single line. The syntax consists of comma-separated key=value pairs, and the value can only be yes or no.
User-agent: *
Content-Signal: search=yes, ai-train=no
Allow: /
Omitting a signal entirely is a third state, and it does not mean granting permission. According to the policy's own wording, if no signal is specified for a use, the site owner is considered to have neither granted nor restricted permission for that use through this mechanism. In the example above, ai-input is absent, which means no preference has been declared on that point.
When Cloudflare's managed robots.txt feature is turned on, search=yes, ai-train=no, use=reference is served by default. ai-input is deliberately left out, because Cloudflare states that it does not know the customer's preference on that point and does not want to guess. The same managed block also adds Disallow: / lines for known training crawlers. This list includes Amazonbot, Applebot-Extended, Bytespider, CCBot, ClaudeBot, Google-Extended, GPTBot and meta-externalagent.
Note that although the signal line and the Disallow lines sit in the same file, they do different jobs. The signal is a declaration; Disallow is an access request. The effect of robots.txt rules on crawl budget depends entirely on the second group; the signal line has no direct effect on crawl behavior.
A signal is a statement of intent, not a technical block
Cloudflare says so explicitly: content signals express preferences, they are not a technical countermeasure against scraping, and some companies may simply ignore them. Compliance with robots.txt is voluntary to begin with, and Content Signals does not change that.
In practice there are three separate layers, and mixing them up creates false expectations:
- Declaration layer. The
Content-Signalline. It declares the purposes for which content may be used and stops no requests. - Request layer. The
AllowandDisallowlines. They guide compliant crawlers and do not stop non-compliant ones. - Enforcement layer. WAF rules, bot management and crawler controls. This is the only layer that actually rejects a request.
The legal side of the policy fits into this picture as well. At the end of the text, it states that restrictions expressed through the signals constitute an express reservation of rights under Article 4 of the European Union's Directive on Copyright in the Digital Single Market. This is not a court ruling but a unilateral record made by the site owner. The practical effect of such a declaration varies by country, the nature of the dispute and the relevant jurisdiction, so it would be wrong to read it as a legal safeguard.
Who reads the signals today?
This is the most overlooked part of the topic. Just because a signal is defined does not mean crawlers process it.
Google's open-source robots.txt parser contains a list of tags that are recognized but not used. content-signal and content-usage are on this list. The code's own comment describes the situation clearly: these are tags commonly found in robots.txt files, Google does not use them for anything, but other search engines may. The only fields Google officially supports are user-agent, allow, disallow and sitemap, and unrecognized lines are silently ignored.
Google's John Mueller made a similar point, saying that as far as he knew no crawler or language model uses these directives, and that the line only adds bloat and maintenance overhead to the robots.txt file. Content Signals is not an accepted internet standard but a proposal published by one company and submitted to the IETF as an individual draft. Work on standardizing AI usage preferences is progressing on its own track in a separate working group within the IETF.
None of this makes the signal worthless, but it does set realistic expectations. Visibility in AI engines today is determined not by a signal line but by the accessibility of content, source quality and actual crawler permissions.
Search crawling and training crawling are separate processes
This is the real confusion Content Signals is trying to solve. Most major AI providers do not use a single bot. They have separate user agents for training, for search and answer generation, and for on-demand fetching triggered by a user request, and the robots.txt permissions for these are independent of each other.
| Provider | Training crawl | Search and answer crawl | User-triggered fetch |
|---|---|---|---|
| Google-Extended | Googlebot | No separate bot | |
| OpenAI | GPTBot | OAI-SearchBot | ChatGPT-User |
| Anthropic | ClaudeBot | Claude-SearchBot | Claude-User |
OpenAI draws the distinction explicitly in its own documentation: each setting is independent of the other, and a site can allow OAI-SearchBot while blocking GPTBot and continue to appear in search results. A site that blocks OAI-SearchBot, on the other hand, is not shown in ChatGPT's search answers. On Anthropic's side as well, blocking the training bot does not block the search bot.
On Google's side, the distinction works a little differently. Google-Extended is a separate token for training Gemini models and for grounding in Gemini apps. Google writes that Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal in search. AI Overviews and AI Mode, however, are part of search and are crawled by Googlebot. So what separates these two surfaces is not a separate bot but snippet controls.
How does misconfiguration damage search visibility?
The signal line itself does no harm because it is ignored. The harm comes from the access rules written next to the signal. Four scenarios come up again and again in practice.
Closing the site to all bots and assuming it stays open to search
This is the most expensive mistake. Writing Disallow: / under User-agent: * and adding Content-Signal: search=yes right below it does not allow search. Googlebot does not read the signal line; it reads the Disallow line. The result is that the site stops being crawled and over time disappears from search results. Even though the declaration and the request are in the same file, the crawler does not resolve the contradiction; it applies the request.
Blocking Googlebot to avoid AI summaries
A site owner who does not want to appear in AI Overviews and blocks Googlebot drops out not of AI Overviews but of search entirely. Google's own documentation says AI features are embedded in search and use the same crawling infrastructure. The right control is not a separate bot block but preview controls such as nosnippet, max-snippet and data-nosnippet. These controls restrict use in summaries while allowing the page to stay in the index.
Blocking every bot with "AI" in its name
Ready-made block lists often lump training bots and search bots together. Blocking OAI-SearchBot, Claude-SearchBot and similar search-side bots means the site drops out of the source pool for AI search answers. What you gain in return is limited, because scrapers that don't follow robots.txt anyway are not affected by these lines. The compliant operator's visibility channel closes, while the non-compliant scraper keeps pulling the content.
Assuming the file on the server and the published file are the same
Cloudflare's managed robots.txt feature does not delete the site's existing robots.txt file. It prepends its own managed block to the existing content and merges the two into a single response. As a result, the file can end up with two separate User-agent: * groups, and the file on the server and the file served to visitors are no longer the same. Always audit the published version. You can check which AI bots are allowed and which are blocked in your site's published robots.txt file with our free AI crawl audit tool.
Who should choose which signal?
The decision can be reduced to two questions. First, what do you gain from your content being training data for a model? Second, what do you gain from your content being used as a source in AI answers?
The visibility cost of saying ai-train=no is low. Training crawls embed content in a model's weights and produce neither clicks nor citations in return. That is why Cloudflare's default points this way too. For publishers producing original research, datasets, course content or archives, this choice is usually clear.
The ai-input side, however, involves a real trade-off. Turning it off means not wanting your content used in live answers. But being cited as a source in AI answers is now a new channel of traffic and brand visibility for many sites. For a business that wants its product and service pages, documentation or support content to appear in AI answers, turning off ai-input would be shooting itself in the foot. For publishers with paid archives, licensed data or subscription content, the opposite may be true.
There is no single correct combination. But you need to be realistic about what the signal does: writing no does not protect content; it only puts the preference on record. If protection is what you want, the job falls to the enforcement layer.
Frequently Asked Questions
Will adding Content Signals affect my Google rankings?
It has no direct effect. Google's robots.txt parser counts the content-signal line among tags that are recognized but not used, and does not process it. What affects rankings is not the line itself but the Allow and Disallow rules added to the same file alongside it.
If I block Google-Extended, will I drop out of search?
No. Google states that Google-Extended does not affect a site's inclusion in Google Search and is not used as a ranking signal in search. This token is for training Gemini models and for grounding in Gemini apps. What causes a site to drop out of search is blocking Googlebot.
Can I add the signals if I don't use Cloudflare?
Yes. The policy was published under a CC0 license, meaning the text can be used freely on any site. Adding the comment block and the Content-Signal line to your own robots.txt file manually is enough. Cloudflare's managed feature only automates this work; it is not required for the signal to be valid.
Are Content Signals and llms.txt the same thing?
They are not the same. Content Signals is a line written inside the existing robots.txt file that declares usage preferences. llms.txt is a separate file that aims to offer language models a curated summary of a site's content. Neither is an official standard supported by the major search engines today, so neither can be presented as a required condition for search visibility.



