AI Crawler Audit: can GPTBot and ClaudeBot see your site?
Enter your domain: the tool reads your robots.txt live and reports the access status of 12 AI bots on a single screen. Free, no sign-up required.
Enter a domain
We analyze your site's robots.txt file and show which AI bots are allowed in.
What is an AI crawler audit?
An AI crawler audit is the check that covers the technical access side of AI search visibility work (GEO, generative engine optimization). The audit has two components: the site's robots.txt file and the official bot identities of AI companies. Results are measured per bot; the tool on this page reads 12 bot identities, including GPTBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot, and reports whether each one is allowed or blocked. The need for this check is recent: OpenAI announced GPTBot in August 2023, Google introduced the Google-Extended token in September of the same year, and site owners have managed what they tell each bot separately ever since.
What is robots.txt and how does it relate to AI bots?
The robots.txt file sits at the root of a domain and tells crawlers which parts of the site they may access. The rules originated in 1994 with a proposal by Martijn Koster and became an official internet standard as RFC 9309 in September 2022; the syntax is explained in detail in Google's robots.txt documentation. AI bots read the same file, because every bot identifies itself with an identity line called the User-agent and looks for the rules written under that name.
The rules work in blocks: the User-agent line sets which bot you are targeting, and the Allow and Disallow lines below it set where that bot may go. A bot without its own block follows the rules of the catch-all wildcard group (User-agent: *). The example below opens the whole site to ChatGPT's training bot while blocking Gemini training:
User-agent: GPTBot
Allow: /
User-agent: Google-Extended
Disallow: /
What do GPTBot, ClaudeBot and Google-Extended do?
Bot identities fall into two job classes: training bots collect content for model training, while search bots crawl to find sources when generating live answers. OpenAI defines three separate identities in its official bot documentation. GPTBot crawls for training and OAI-SearchBot collects sources for ChatGPT search, while ChatGPT-User steps in when a user asks the chat to open a page. Anthropic and Perplexity document ClaudeBot and PerplexityBot on their own bot pages in the same way.
| Bot | Owner | Job class |
|---|---|---|
| GPTBot | OpenAI | Crawling for model training |
| OAI-SearchBot | OpenAI | ChatGPT search sources |
| ChatGPT-User | OpenAI | Page visits on user request |
| ClaudeBot | Anthropic | Crawling for model training |
| PerplexityBot | Perplexity | Answer engine index |
| Google-Extended | Gemini training permission token | |
| CCBot | Common Crawl | Open web archive |
Google-Extended is a special case: according to Google's crawler documentation, it is not a separate crawler but a permission token that controls whether content collected by regular Googlebot crawling is used for Gemini training. Blocking it does not change classic Google rankings. CCBot is the bot of Common Crawl, a non-profit founded in 2007; its open archive has fed the training sets of many models. The tool also checks the Claude-Web, Bytespider (ByteDance), Amazonbot, Applebot-Extended (Apple) and meta-externalagent (Meta) identities.
Why does being open to AI bots mean visibility?
Without access, no answer engine can cite your site; engines only quote the pages they can see. According to Google's announcement, AI Overviews (the AI summaries in Google Search) have been running in Turkish in Türkiye since February 18, 2026, so the summary layer is no longer the exception on Turkish results pages. The same logic applies to ChatGPT: its search mode is fed by pages OAI-SearchBot can crawl, and if the bot is blocked, your competitors get mentioned in the answers instead.
Bot access opens the door but does not bring citations: content quality, freshness and off-site brand signals decide which brand gets mentioned. Access is your ticket into the race, and the audit is not a one-off approval but the baseline measurement of GEO work. The reverse is clear-cut: a closed door produces zero visibility in that engine, however strong the content is.
Which bots might make sense to block?
The right answer depends on your content model: blocking saves training data and server load on one side, but shuts off your chance of being mentioned in answer engines on the other. The decision comes down to three questions and leads to three different robots.txt setups: does your site win customers, does it sell the content itself, or is crawl load the real problem?
- For service and product sites that win customers, openness pays: an AI answer is a new demand channel, there is no licensing revenue for a block to protect, and allowing every bot on the list is the standard choice in this segment.
- For publishers that sell the content itself, a mixed setup that blocks training bots (GPTBot, ClaudeBot, Google-Extended, CCBot) and keeps search bots (OAI-SearchBot, PerplexityBot) open is defensible: licensing value is protected and citation traffic continues.
- For large sites struggling with crawl load, the issue is capacity, not visibility preference: the fix is server-side rate limiting and a bot management layer, before any robots.txt line.
In all of these decisions, robots.txt is a request, not a lock: the file cannot physically stop a bot, and compliance depends on the bot's own word. OpenAI, Google and Anthropic state that they follow the rules, and server logs generally back this up. On the other hand, in August 2025 Cloudflare reported that Perplexity had bypassed robots.txt blocks using undeclared crawlers. A site that wants a hard block therefore moves the job to the server layer: bot verification and a 403 response via a WAF (web application firewall) or CDN (content delivery network) turn the request into a rule.
What is the difference between robots.txt and llms.txt?
The two files answer different questions: robots.txt controls who may enter, while llms.txt proposes giving those who enter a tidy summary of the content. llms.txt is a September 2024 proposal by Jeremy Howard, founder of Answer.AI; it is written in the plain-text Markdown format and lists a site's key pages for models. Their weight differs too: robots.txt is a standard under RFC 9309 and every major bot reads it, whereas Google's John Mueller said in 2025 that no major AI system uses llms.txt. Trying it costs little: the llms.txt generator tool produces the file in a few minutes, but the bulk of your visibility effort should stay on access and content.
Which robots.txt mistakes lock a site out without anyone noticing?
The same mistakes keep showing up in audit results, and most come from a single careless line. The most expensive one is the silent one: the site stays closed to AI engines for months and nobody notices, because classic Google traffic carries on as usual and no dashboard warns about the lockout. Six mistakes are especially common:
- Unintended global block: a broad Disallow under User-agent: * also covers every AI bot without its own block, and the site unintentionally drops out of ChatGPT and Perplexity sources.
- Wrong bot name: matching works on the exact identity name; hyphenated spellings such as GPT-Bot or Claude-Bot bind no bot at all and the rule becomes a dead line.
- Empty Disallow confusion: a Disallow line with an empty value allows everything, while Disallow: / closes the whole site; a single character produces two opposite results.
- Conflict precedence: when an Allow and a Disallow target the same path, Google's implementation lets the longest, most specific rule win; nested rules written without knowing this precedence produce surprises.
- Forgotten subdomains: robots.txt works per host; blog.site.com does not take rules from the root domain's file but looks at its own.
- Layer confusion: robots.txt may allow a bot while a WAF or CDN rule on the server turns it away with a 403; the audit on this page measures the robots.txt layer and does not show server-level blocks.
A roadmap for AI bot access
The roadmap has three steps. First, measure: enter your domain in the tool above, see the status of all 12 bots and compare the table with your intent; the riskiest case is not a deliberate block but one nobody has noticed. Next, match: a site that wants customers stays open to training and search bots, a publisher selling its content sets up a mixed configuration that blocks training bots and keeps search bots open, and a site with capacity issues looks for the fix at the server layer. The final step is continuity: the bot list keeps growing and new identities keep appearing, so rerun the audit after every robots.txt change and check newly added bots in the same pass. Access does not guarantee citations, but a closed door guarantees invisibility.
Common Questions
Clear answers about AI bots, robots.txt permissions and AI visibility.
Let us track your AI visibility
From bot access to brand tracking in ChatGPT, Gemini and Perplexity answers, we manage it end to end.