Home/ Blog /SEO

Crawl Budget and Robots.txt Optimization: Whose Problem Is It?

Turan Doğan
Turan Doğan
SEO & GEO Specialist
SEO April 13, 2026 14 min read
Crawl Budget and Robots.txt Optimization: Whose Problem Is It?
SUMMARY
Crawl budget is the set of URLs Google can and wants to crawl on a site; it has two components, the crawl capacity limit and crawl demand. Google describes the topic as advanced and aimed at sites with more than a million pages whose content changes weekly, or more than ten thousand pages whose content changes daily, not at small and medium-sized sites. Robots.txt manages crawling and does not prevent indexing; a blocked URL can still appear in search results without a description if it is linked from other sites.

Googlebot does not come to a site with unlimited patience. Every request costs your server something, and the crawl resources Google can spread across the web are not infinite either. The point where these two limits meet is called crawl budget.

The concept itself is real. However, Google's guide on the subject says in its very first lines that if your site does not have a large number of rapidly changing pages, or if your pages are crawled on the day they are published, you do not need to read it. For most small and medium-sized sites, crawl budget is not a bottleneck that needs solving; it is a source of unnecessary work.

What Crawl Budget Is and How Google Defines It

Crawl budget is the set of URLs Google can and wants to crawl on a site. It is the intersection of two separate quantities: the load the server can handle (the crawl capacity limit) and how much Google wants to crawl the site (crawl demand).

A technical detail that is often overlooked is that "site" here means the hostname, not the domain. From the standpoint of Google's crawling infrastructure, www.example.com and code.example.com are separate sites with separate crawl budgets. Moving a heavy archive to a subdomain eases the main site's budget, but that archive starts from scratch with its own budget; every site opens with the same default, conservative capacity limit.

The second distinction is more important: being crawled does not mean being indexed. Google does not add every page it crawls to the index. After crawling, each page is evaluated, consolidated with similar pages and separately assessed for whether it is eligible for indexing. This two-stage relationship between crawling and indexing is part of how SEO fundamentally works, and almost every misunderstanding on this topic comes from confusing the two stages.

Who This Topic Really Applies To

Google describes its guide as an advanced resource and lists three site profiles:

  • Large sites whose content changes frequently (around once a week), meaning more than one million unique pages
  • Medium or large sites whose content changes daily, meaning more than ten thousand unique pages
  • Sites with a large share of their URLs classified in Search Console as "Discovered - currently not indexed"

Google also notes that these numbers are not exact thresholds but rough estimates to help you classify your site. So reading it as "my page count is under a million, so this doesn't concern me" is not correct either. What matters is the load produced by page count and rate of change together.

Search Console draws a similar line. The documentation for the Crawl Stats report itself says that for a site with fewer than a thousand pages, there is no need to use the report or to concern yourself with crawling at this level of detail.

Of the three criteria, the third is the most useful. The first two describe the size of the site; the third describes the symptom: Google knows about the URLs but has not gotten around to crawling them. Even if your page count is far below the threshold, if this backlog exists, there is something to investigate. The reverse is also true: if page count is high but there is no such backlog, the bottleneck is not on the crawling side.

The Two Components of the Budget

Crawl capacity limit

The capacity limit is the upper bound Google calculates so it can crawl without overloading the server. It takes into account both the number of parallel connections and the time between fetches. Every site starts with the same default, conservative limit; if demand rises and the site stays healthy, Google's systems adjust this limit automatically over time.

What raises the limit is consistency. If the site responds reliably and response times (including latency and time to first byte) stay stable or improve, the limit goes up and more connections become available for crawling. What lowers it is slowdowns, 5xx server errors and rate-limiting signals such as 429. The practical consequence: on the crawling side, server performance is not a matter of comfort but a matter of capacity.

Crawl demand

Demand is determined by the size of the site, how often it is updated, page quality and relevance compared with other sites. Google singles out three factors among these that site owners can influence:

  • Perceived inventory: without guidance, Google tries to crawl all or most of the URLs it knows about on your site. If many of those URLs are duplicates of each other or should not be crawled for some other reason, crawling is wasted. Google describes this as the factor you can control most positively.
  • Popularity: URLs that are more popular on the internet tend to be crawled more often so they stay fresher in Google's systems.
  • Staleness: Google recrawls documents at intervals to pick up changes.

The relationship between capacity and demand is not one-way. If demand is low, Google crawls the site less even if the capacity limit is never approached. Upgrading the server does not make Google want to crawl a site it is not interested in crawling.

One more detail: the capacity limit is shared among all of Google's crawlers. High demand from one crawler, for example AdsBot on a site that serves dynamic ad targets, can reduce the capacity available to the others.

URL Patterns That Eat Up Crawl Budget

On sites with budget problems, the problem is usually not the number of pages. It is how many URL variations those pages multiply into.

Faceted navigation and parameter combinations. Faceted navigation is useful for users, but in its most common implementation, which relies on URL parameters, it can create infinite URL spaces. A crawler cannot tell whether a URL is useful without crawling it; so it first fetches a large number of filter combinations and only afterwards realizes they are worthless. The damage comes in two layers: overcrawling and slower discovery of new, useful URLs. This pattern is most common on e-commerce sites, because color, size, brand and sort parameters multiply each other.

For sites that do not need filtered URLs indexed, Google recommends robots.txt directly and gives this example rule:

user-agent: Googlebot
disallow: /*?*products=
disallow: /*?*color=
disallow: /*?*size=
allow: /*?products=all$

If filtered URLs do need to be crawled, the rule set changes. Use the industry-standard & character as the parameter separator, because characters such as commas, semicolons and square brackets are hard for crawlers to recognize as separators. If filters are encoded in the URL path, keep their logical order fixed and do not allow duplicate filters. A filter combination that returns no results should not redirect to a generic error page; return a 404 on the URL itself.

Soft 404 pages. Pages that present non-existent content with a "not found" message but a successful response code keep getting crawled and spend budget. They appear as a separate row in Search Console's Page indexing report.

Duplicate content. Differently sorted versions of the same list, infinite scroll pages that repeat information already on the linked page, and near-duplicates generated from templates all inflate the number of unique URLs. Google's order of operations is clear: try to consolidate first, and block only if you cannot.

Long redirect chains. Every step in a chain is a separate request. In a setup where page 1 redirects to page 2 and page 2 redirects to page 3, Google spends three requests to reach a single destination.

Permanently removed pages. Google does not forget a URL it knows about, but a 404 or 410 is a strong signal not to crawl that URL again.

What Robots.txt Does and Does Not Do

Robots.txt tells search engine crawlers which URLs on your site they can access. Google's definition is that narrow: the file is mainly used to keep requests from overloading your site, and it is not a mechanism for keeping a web page out of Google.

The most concrete consequence of this distinction is this: if a page blocked by robots.txt is linked from other sites, Google can index the URL without ever visiting the page. The URL appears in search results, but because the page cannot be read, it appears without a description. The way to genuinely remove a page from search results is not blocking but noindex or password protection.

The second side effect of blocking concerns resources. Files embedded in a blocked page, such as images, videos and PDFs, stay out of crawling unless they are referenced from other pages that are allowed to be crawled. Blocking the script and style files needed to understand a page is risky for the same reason: Google cannot properly analyze a page that depends on those resources.

What the file supports and what it doesn't

Robots.txt supports less than people think. Google recognizes only four fields: user-agent, allow, disallow and sitemap. Widely used fields such as crawl-delay are not supported by Google. Other technical limits worth knowing:

  • The file must be UTF-8 encoded plain text. Invalid lines are ignored and only valid lines are used.
  • Google enforces a size limit of 500 kibibytes; content beyond that size is ignored. Grouping the content you want to exclude into a separate directory, rather than writing long lists of rules, keeps the file small.
  • Path values are case-sensitive and must start with a slash to denote the root directory. Field names, on the other hand, are not case-sensitive; User-agent and user-agent are treated the same way.
  • Google generally caches the contents of the file for up to 24 hours. This means a fix does not always take effect immediately.

An unreachable file triggers its own behavior too. If robots.txt returns a 5xx, Google stops crawling the site for the first 12 hours and keeps trying to fetch the file; if it cannot get a new version, it uses the last good version for 30 days. A maintenance error on the server side can therefore turn into a crawl outage on its own.

Three Mistakes When Using Robots.txt for Crawl Budget

The budget from a blocked page is not transferred to other pages. The most common advice is "block unnecessary pages so the budget goes to the important ones". Google rejects this directly: robots.txt should not be used to temporarily reallocate crawl budget to other pages, because unless the site is hitting its crawl capacity limit, Google does not shift the freed-up budget to other pages. The file's job is to block pages and resources you do not want crawled.

Using noindex to save budget does not reduce crawling. Noindex is not a crawl-saving tool. Google still sends the crawl request, sees the tag or header in the response and stops crawling the page there. The request has already been made; no crawl time has been saved.

Blocking is not forgetting. URLs blocked by robots.txt remain part of the crawl queue for a long time and are crawled again once the block is lifted. For permanently removed content, the right signal is not blocking but a 404 or 410 status code.

Read together, these three make the file's role clear. Robots.txt is a good tool for closing off areas you never want crawled from the start. It does not work as expected for fixing an existing crawl bottleneck after the fact.

Blocking the Search Crawler Shuts Down Visibility

When deciding which bots to block, the search crawler and the tokens meant for AI products are often confused. They are not the same thing.

Googlebot is the crawler for search. Blocking Googlebot stops the page from being crawled and therefore from entering the index; search visibility shuts down. Google-Extended, by contrast, is a standalone product token that controls whether content crawled from your site can be used to train Gemini models and for grounding in Gemini apps. Google's own documentation states that this token does not affect a site's inclusion in Google Search and is not used as a ranking signal.

The practical takeaway: if you want to limit use on the AI side, the search crawler is the wrong thing to block. In the other direction, there is an asymmetry. A rule that shuts off search also narrows AI visibility, because part of the content supplied to the model during grounding comes from the Google Search index. A page that is not in the search index cannot be fed from there either.

Diagnosing with the Search Console Crawl Stats Report

The way to move from guesswork to data is the Crawl Stats report in Search Console. The report sits under property settings and works only for root-level properties: a Domain property or a root-level URL-prefix property. It does not appear for properties defined at subfolder level.

The report breaks down total crawl requests, total download size, average response time, host status, crawl responses, file type, crawl purpose and Googlebot type. For diagnosis, three readings stand out:

  • Host status and average response time: these show which direction the capacity limit is heading. If response time is deteriorating or server errors are showing up, the problem is the server, not the URL inventory.
  • Crawl purpose breakdown: the split between discovery and refresh tells you whether crawling is being spent on finding new URLs or refreshing known ones.
  • Crawl responses: the share of redirect and error responses in the total shows where waste is accumulating.

Two technical details often cause the report to be misread. First, with server-side redirects each step in the chain is counted as a separate request, so redirect chains inflate the total request count. Second, crawls that were considered but not performed because of robots.txt are also included in the crawl totals. The example URLs in the report are a representative sample, not a complete list; you cannot conclude that a URL was not requested just because it does not appear there. Reading this report as part of a workflow that tracks Search Console data regularly yields far more meaningful results than a one-off look.

Where a Small or Medium-Sized Site Should Spend This Time

On a site with a few hundred to a few thousand pages whose content does not change daily, crawl budget optimization produces no measurable gain. Google names only two tasks for sites like these: keep the sitemap up to date and check the Page indexing report regularly.

If you are already doing both, putting any more energy into crawling brings no return. If pages are crawled on the day they are published, crawling is not your bottleneck. The bottleneck most likely lies in the information needs the content meets, the internal linking structure or how well the page matches the query.

Still, a few practices are right at any scale. They are worth doing on their own merits, not for crawl budget: keep the sitemap up to date and use the lastmod tag for changed content, avoid long redirect chains, make sure pages load quickly, and take advantage of HTTP caching by supporting the 304 response for pages that have not changed since the last crawl.

Frequently Asked Questions

Is a site without a robots.txt file penalized?

No. On a site without robots.txt, Google assumes there are no crawl restrictions and keeps crawling. The file is not a requirement but a tool for sites that want to manage crawl traffic. Returning responses such as 401 or 403 for the file in order to limit crawl rate is not the right approach either.

The page I disallowed still shows up in Google. Why?

Because disallow blocks crawling, not indexing. If the page is linked from other sites, Google can index the URL without visiting it, and the result appears without a description. To remove the page from results, you need to lift the block and use noindex, because Google cannot see the noindex tag on a blocked page anyway. Crawling and indexing rules stacked on the same page can cancel each other out like this.

Does the crawl-delay directive slow down Googlebot?

No, because Google does not support the crawl-delay field. The fields Google recognizes are limited to user-agent, allow, disallow and sitemap. Since other search engines may interpret this field, having it in the file is harmless, but it does not change Googlebot's behavior.

How do I increase my crawl budget?

Google describes two ways. The first is adding server resources, which only makes sense if the site cannot be crawled because of its own server capacity; a host load error in the URL Inspection tool is the indicator of this. The second is improving content quality for the Google product being targeted: on the Google Search side, popularity, overall user value, content uniqueness and serving capacity factor into this. There is no shortcut; the answer to the second point is the overall quality of the site.

Was this article helpful?
Add Seobaz as a preferred source on Google to see us more often in your search results and AI answers.
Add as preferred source
Share this article
Turan Doğan
Founder · SEO & GEO Specialist
Publishing up-to-date guides on SEO, GEO and AEO since 2014, helping brands get seen on both Google and AI engines.
WhatsApp Online · Quick reply
Gift Wheel A discount on every spin
View Cart