
The server access log is the only first-hand record that shows whether an AI engine has actually visited your site. To appear as a source in a ChatGPT, Claude or Perplexity answer, these systems first need to be able to reach the page, and whether they did is not something to guess at: it is written line by line in the server's own log.
Log analysis boils this record down to three questions: who came, what did they request, and what response did they get. The hard part is the first question, because a request can introduce itself under any name it likes, and not every GPTBot line in your log is really OpenAI. The sequence below starts with getting access to the log file and ends with a table of verified bots.
Where is the access log stored?
The access log is a plain text file in which the web server software writes one line per request. Its location depends on the server type.
- Nginx: the default path is
/var/log/nginx/access.log - Apache:
/var/log/apache2/access.logon Debian and Ubuntu-based systems,/var/log/httpd/access_logon RHEL and CentOS-based systems - cPanel hosting: the
logsfolder under the account, or/usr/local/apache/domlogs/server-wide - Plesk:
/var/www/vhosts/example.com/logs/access_log
On shared hosting panels, raw logs are usually downloaded as compressed archives and listed under a name like raw access logs. For sites without SSH access, this is the only way.
If the site sits behind a CDN or firewall, two more things come into play. First, the IP address in the origin server's log may be the CDN's own address. The real client address is carried in the X-Forwarded-For header, and IP verification is impossible until the log format is configured to record this header. Second, a request blocked at the CDN layer never reaches the origin server. That block won't show up in the origin log; you need to check the CDN's own logs or analytics screen.
Which fields do you read in a log line?
In the widely used Combined Log Format, a single line looks like this:
203.0.113.10 - - [12/Mar/2026:09:14:22 +0300] "GET /services/example HTTP/1.1" 200 18432 "-" "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot"
Five fields in this line are enough for analysis.
| Field | What it tells you |
|---|---|
| Client IP | The only reliable field verification can rest on. A user agent is easy to fake; an IP is not. |
| Timestamp | Crawl frequency, peak periods and the date of the last visit. |
| Request path | Which page the bot requested. This is the input for coverage analysis. |
| Status code | Whether the bot got the page or was turned away at the door. |
| User agent | The name the bot introduces itself with. It is a claim of identity, not proof of identity. |
AI bots do three different jobs
Recognizing a user agent isn't enough; you need to know what that bot does. Different bots from the same company work independently of each other, and their entries in the log mean different things. The operators' official documentation spells out this distinction.
| Operator | Search and index bot | Model training bot | User-triggered fetcher |
|---|---|---|---|
| OpenAI | OAI-SearchBot | GPTBot | ChatGPT-User |
| Anthropic | Claude-SearchBot | ClaudeBot | Claude-User |
| Perplexity | PerplexityBot | No separate bot defined | Perplexity-User |
| Googlebot | No separate user agent | Google-GeminiNotebook and similar |
When you see a search and index bot, the page becomes a candidate for the pool of sources that system draws on when it builds answers. This is the class most directly relevant to visibility. Perplexity states in its own documentation that PerplexityBot does not collect content to train AI foundation models but serves to surface sites in search results.
A model training bot makes content a candidate for a training corpus. This visit makes no direct contribution to visibility in today's answers, and blocking it does not automatically shut out the crawler on the search side. GPTBot and OAI-SearchBot are separate user agents and are managed separately in robots.txt.
The user-triggered fetcher is the most misread class in log analysis. A ChatGPT-User, Claude-User or Perplexity-User line means a real user asked a question at that moment and the system opened the page live. This is not a crawl signal but a demand signal, and it has high strategic value because it shows which page was fetched in response to which question. Google writes in its own documentation that user-triggered fetchers generally ignore robots.txt rules, because the request comes from a user. Other operators' fetchers may not behave the same way, so each bot's stance on robots.txt should be checked in its own documentation.
Two points on the Google side are especially easy to confuse. First, Google-Extended never appears in the log. Google states clearly that this name is not a separate HTTP user agent, that crawling is done with existing Google user agents, and that the name is used only for control purposes in robots.txt. Searching for Google-Extended in your log and finding nothing does not mean Google isn't using your content for AI. Second, there is no separate crawler for AI Overviews and AI Mode. Because these surfaces rely on Google's search crawling and indexing infrastructure, you won't find a separate AI Overviews visit in your logs.
How do you verify a fake user agent?
The user agent field is free text written by the client itself. A script can declare itself GPTBot, crawl the site and leave a line in the log that is indistinguishable from the real thing. That is why bot counting starts with the user agent and ends with the IP. The verification method varies by operator, and if you don't know the difference, the entire count is thrown off.
Reverse DNS verification for Google bots
Google supports the reverse DNS method, and verification takes three steps.
- Run a reverse DNS lookup on the IP from the log line, for example
host 66.249.66.1. - Check that the returned domain name ends in
googlebot.com,google.comorgoogleusercontent.com. - Run a forward DNS lookup on the same domain name and confirm that the returned address matches the IP in the log.
The third step can't be skipped. If you check only the reverse record, any server that points its PTR record at a domain it controls can pass as Google. Google also publishes its crawler IP ranges as JSON files in CIDR format and keeps separate files for common crawlers, special-case crawlers and user-triggered fetchers.
Published IP lists for AI bots
OpenAI, Anthropic and Perplexity use published address lists instead of reverse DNS. That's why a reverse DNS script written for Googlebot doesn't work for these bots and flags all of them as fake.
- OpenAI publishes a separate list for each bot. The address blocks for OAI-SearchBot, GPTBot, ChatGPT-User and OAI-AdsBot, which checks the safety of ad landing pages, are kept in independent files.
- Anthropic groups the addresses of its three Claude bots under a single list and states that a request whose source address is on this list comes from Anthropic.
- Perplexity publishes two separate lists for PerplexityBot and Perplexity-User.
The files contain CIDR blocks, and verification simply means checking whether the address in the log line falls within one of these blocks. Because the lists change, downloading them once and using them for a long time gives wrong results. For the same reason, these lists are for identity verification and are not a suitable source for writing permanent blocking rules.
How to run a log analysis step by step
The sequence below can be run on a single server with standard command-line tools. The file name is assumed to be access.log.
- Define the period and collect the logs. Work with at least thirty days of records. A single day of data isn't enough to draw conclusions about crawl behavior. Compressed archives can be combined with
zcat. - Filter the bot lines.
grep -Ei "GPTBot|OAI-SearchBot|ChatGPT-User|OAI-AdsBot|ClaudeBot|Claude-User|Claude-SearchBot|PerplexityBot|Perplexity-User" access.log > ai-bot.log - Count by bot. Parse the user agent names to see how often each engine visits.
grep -oEi "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-User|Claude-SearchBot|PerplexityBot|Perplexity-User" ai-bot.log | sort | uniq -c | sort -rn - Extract the list of unique IPs.
awk '{print $1}' ai-bot.log | sort -u - Compare each address with the CIDR list published by the operator. Collect non-matching lines in a separate file and leave them out of the bot count. If you skip this step, every number after it is inflated.
- Verify Googlebot lines separately with reverse DNS. Include the forward lookup step too, otherwise verification is only half done.
- Extract the most requested paths from the verified lines.
awk '{print $7}' ai-bot-verified.log | sort | uniq -c | sort -rn | head -30 - Extract the status code distribution.
awk '{print $9}' ai-bot-verified.log | sort | uniq -c | sort -rn - Compare with the set of URLs in the sitemap. Sort both lists and take the difference; what remains are the pages that haven't been crawled.
comm -23 sitemap-urls.txt crawled-urls.txt - Repeat the measurement at fixed intervals. A single pull gives you a snapshot; decisions need a trend.
What conclusions can you draw from logs?
Which pages are being crawled?
Put the paths most requested by verified bots side by side with the pages that actually carry value for the site. A common picture: bots concentrate on tag pages, pagination sequences or archive URLs while barely visiting service and product pages. This is not a content problem but a routing problem, and it is solved at the level of internal linking and the sitemap. The framework for managing where crawl resources flow is covered in detail in crawl budget and robots.txt optimization.
Is there a section that isn't being crawled?
The difference between the sitemap and the set of crawled URLs is the output that pays off fastest. If a section isn't crawled at all, the cause is usually one of five things: a disallow line in robots.txt, orphan pages with no internal links, a canonical tag pointing to another URL, a security layer rejecting that directory, or response times slowing down crawling.
Are bots getting error codes?
On most sites, the status code distribution produces the most valuable table.
- 403 is the most common and the quietest problem. A firewall, bot management plugin or hosting protection is rejecting the bot. Because the site opens fine in a browser, the problem goes unnoticed for a long time.
- 404 points to a broken internal link or an outdated sitemap.
- 301 chains cost one more request at every hop. Connect redirects to their destination in a single step.
- 429 and 5xx mean rate limiting or server load. If they repeat, bots lower their crawl rate on their own.
- 200 with a very small response size needs attention. If the page body is built in the browser with JavaScript, the returned HTML may be empty, and the status code alone doesn't prove the content was read.
If you can't access your log files yet, there is a simpler starting point for checking whether your site is open to AI bots. Seobaz's free AI crawl audit tool shows the basic access-side blocks without downloading any logs and clarifies which question to start your log analysis with.
Three situations that lead to misreading log data
Counting unverified lines. When fake user agents are included, the bot count climbs to several times the real figure and produces a visibility picture that doesn't exist. Verification is not an optional extra step; it is a precondition for counting.
Mistaking training bot visits for visibility. Heavy GPTBot or ClaudeBot activity does not mean your content is being cited as a source in answers. The training corpus and the search source pool are different pipelines, and they are measured separately.
Making a sweeping decision with a single line. A broad block written into robots.txt can shut out the search-side bot along with the training bot you meant to block. The practical consequences of the PerplexityBot and Perplexity-User distinction are covered in our Perplexity GEO article, and the OAI-SearchBot and GPTBot distinction on the OpenAI side in our ChatGPT GEO article.
Log analysis is not a one-off audit. Bot lists, user agent versions and published address blocks change, so the same set of commands can give different results months later. Measuring at a fixed interval tells you more than a single deep analysis.
Frequently Asked Questions
Can you see AI bots in Google Analytics?
No. Analytics tools only record a visit when the tracking code on the page runs, and they filter out known bot traffic as well. The server log, on the other hand, records every incoming request whether or not any code runs. The log is the only complete source for bot behavior.
How many times does a bot need to crawl a page?
There is no published threshold, and even if there were, its meaning would vary with site size. Instead of an absolute number, look at two things: whether your important pages are covered, and whether crawling has stopped at any point. A long stretch with no requests at all is a stronger signal than frequent requests.
If ClaudeBot is blocked, does Claude visibility end?
Anthropic defines separate user agents for ClaudeBot, Claude-SearchBot and Claude-User, and they are managed independently in robots.txt. Blocking the training bot doesn't change the permissions given to the search bot and the user fetcher. The same logic applies to OpenAI's distinction between GPTBot and OAI-SearchBot.
What if the server doesn't keep logs?
The access log is enabled in the web server settings, and most hosting panels already have a raw log option. The real issue is retention: in many setups, old logs are deleted at short intervals. If you plan to run an analysis, set up an archiving rule first, then wait for enough data to accumulate.



