Home/ Blog /SEO

What Is the Googlebot Page Size Limit?

Turan Doğan
Turan Doğan
SEO & GEO Specialist
SEO September 13, 2026 10 min read
What Is the Googlebot Page Size Limit?
SUMMARY
The Googlebot page size limit is the upper threshold on the number of bytes Google downloads from a single URL. For Google Search, this threshold is 2 MB for supported file types and 64 MB for PDF files. Bytes beyond the threshold are content that Google does not download, render or index.

How many MB does Googlebot download from a page?

When crawling for Google Search, Googlebot downloads the first 2 MB of a supported file type. For PDF files, the limit is 64 MB. As soon as the limit is reached, the download stops, and the portion downloaded up to that point is passed to the indexing systems and the rendering service as if it were the entire file. Bytes beyond the limit are not downloaded, not rendered and not indexed.

The most important detail here is that the page is not rejected. When Googlebot encounters a 2.4 MB HTML document, it does not skip the page; it truncates it. That is why the problem usually shows up not as the page missing from the index entirely, but as the content in the lower half of the page never matching any query. Search Console produces no error or warning for this, so the truncation is silent. This is also one reason the limit rarely appears in classic search engine optimization checklists.

The limit only counts the HTML document itself

2 MB is not the total weight of a page. The limit applies per URL, and the counter measures the byte stream the server sends for that URL. Images, external CSS files and external JavaScript files referenced in the HTML are fetched with separate requests, and each is subject to its own counter. So a page carrying a 6 MB image gallery does not come near the limit as long as the HTML document stays small.

The same logic works in reverse. Everything embedded in the HTML counts toward the total. Inline style blocks, inline script blocks, base64-encoded images and JSON blocks printed for hydration are part of the document the server sends, so they eat directly into the 2 MB budget.

Google states that the limit also covers HTTP headers. Since headers take up only a few kilobytes in practice, this detail rarely changes a decision on its own, but it shows that the limit works on response bytes, not on some abstract text length.

The same limit applies during rendering. When Google's rendering service fetches JavaScript and CSS files and processes XHR requests, it applies the 2 MB limit to each resource separately, and it does not request images and videos for rendering. This has a less discussed consequence: when a single JavaScript bundle over 2 MB is cut off in the middle, what remains is not a syntactically valid script, and rendering can break at that point.

Compression won't save you from this limit

The limit applies to uncompressed data. When the server responds with gzip or Brotli, the browser's network panel may show 180 KB, but the number Google counts is the decompressed size. Because repetitive markup compresses very well, the decompressed size is several times the size that travels over the wire. It is common for a category page that transfers 250 KB to exceed a megabyte once decompressed.

This is exactly where most measurement errors come from. The size quoted for a page almost always describes the compressed size, and that is the wrong number as far as the limit is concerned.

Why does the 15 MB figure still circulate?

For a long time, the industry circulated the information that Googlebot crawls the first 15 MB of an HTML file, and this figure still appears in many sources. 15 MB is a real number, but it is no longer the number that applies to Google Search.

Every client that uses Google's crawling infrastructure sets its own byte limit. For a crawler that does not specify its own limit, the platform default is 15 MB. Googlebot, the client for Google Search, has set its own limit at 2 MB. The difference is not small: relying on the 15 MB figure for a web page means treating more than seven times the real limit as a safe zone. Other Google crawlers such as Googlebot Image and Googlebot Video also have their own separate limits, so there is no single byte count for Google.

Which pages hit the limit?

Plain text cannot exceed this limit on its own. A rough calculation makes this clear. 2 MB is room for about two million characters, and the plain text of a 1,500-word article stays around 12 KB, even accounting for non-ASCII characters (such as Turkish letters) taking two bytes in UTF-8. The limit leaves room for more than a hundred times the text of a normal article. The problem comes not from the text but from the load around it.

In practice, these page types are at risk:

  • Unpaginated lists. On category pages that print thousands of products or listings on a single URL, each row carries its own markup, data attributes and annotations. 1 KB of HTML per product comes to 2 MB at two thousand products.
  • Base64-embedded images. Because base64 turns every three bytes into four characters, an embedded image is about a third larger than the file itself. A 400 KB photo takes up close to 530 KB inside the HTML, and four of them fill the limit on their own.
  • Huge inline JSON. Hydration data printed by modern frameworks, data layer objects and API responses embedded as-is can reach many times the size of the visible text. Printing the same product data once as HTML and once as JSON is a common pattern.
  • Inline CSS and JavaScript. Inlining critical CSS for performance is a reasonable choice; inlining the entire stylesheet is not.
  • Unpaginated archives and comment threads. A forum thread with a thousand comments, or an archive that collects an entire year on one page, piles up template markup over and over along with the user content.

The second and more insidious effect of truncation is on link discovery. Because internal links at the bottom of the page are never downloaded, Google does not see them from that page. In an unpaginated archive, this means one of the discovery paths for thousands of URLs silently closes, and the effect is felt in crawl budget management.

How do you measure HTML size?

The right number to measure is the uncompressed response body. Working through the following order is enough in most cases.

  1. Download a single URL with compression turned off. The command curl -s -H "Accept-Encoding: identity" -o /dev/null -w "%{size_download}\n" https://example.com/category gives the size Google counts, in bytes.
  2. Open the network tab in your browser's developer tools and look at the document request. The transferred size and the resource size are shown separately; the second is the one that matters for the limit.
  3. View the crawled page with the URL Inspection tool in Search Console. This shows the HTML Google actually received, so you can see directly whether the end of the document is missing.
  4. If the suspicion is about a template rather than a single page, crawl the site and list document sizes in bulk. Template-driven bloat always clusters in the same page type. The on-page SEO analysis tool can be used as a starting point for this crawl.

The fastest way to confirm truncation is to place a short, unique string at the very end of the HTML. If this string, sitting just before the closing tag, does not appear in the crawled page view, the document is being truncated. The same method also shows in one step whether a fix has worked.

How do you reduce page size?

Start by looking at where the bytes go, because the gain is almost always concentrated in a single item. The order usually goes like this.

Moving inline assets out of the document gives the biggest and cheapest gain. Google's own advice points the same way: heavy CSS and JavaScript go into external files, because external resources are fetched separately and are subject to their own limits. Convert base64-embedded images into real files as well.

Next, clean up duplicated data. Trim the JSON printed into the page down to the fields the page actually needs. Carrying full product descriptions as both HTML and JSON doubles the document and gains nothing on the search side.

Paginate long lists. Pagination both keeps the document under the limit and spreads link discovery across crawlable URLs. Content loaded with infinite scroll does not do this job if it leaves no crawlable address behind.

Finally, order matters. Google recommends keeping meta tags, the title tag, link tags, the canonical and essential structured data at the top of the document. If these elements fall below the truncation point, information about the page's identity is lost entirely, which is a heavier consequence than losing content.

Which limits is this one confused with?

The page size limit is not the only byte limit Google publishes. Confusion usually arises among the following four items.

Limit What it applies to Value What happens if exceeded
Googlebot page size A single URL fetched for Google Search 2 MB, 64 MB for PDFs The excess is not downloaded, rendered or indexed
Crawler default Google crawlers that do not specify their own limit 15 MB Content beyond the limit is ignored
Sitemap file A single sitemap file 50 MB uncompressed or 50,000 URLs The file must be split and an index file used
robots.txt The site's robots.txt file 500 KiB Rules after that point are ignored

The one most often confused is the sitemap limit, because both are expressed in megabytes and both apply to uncompressed data. But one determines the downloaded portion of a web page, while the other determines how many URLs a file can carry. A page exceeding 2 MB and a sitemap exceeding 50,000 URLs are completely different problems with different solutions. We cover the splitting logic on the sitemap side separately in our article on XML sitemap index structure.

Crawl budget, on the other hand, is not a size limit. Crawl budget is about how often and how many URLs Google fetches from a site, while the page size limit is about where a single fetch gets cut off. The two are only indirectly related: when heavy pages strain the server, Google lowers its crawl rate on its own.

Frequently Asked Questions

If my page exceeds the limit, will Search Console tell me?

No. Truncation is not a crawl error but a successfully completed request that ends early. The page appears as valid in the coverage report, and its indexing status stays normal. The way to see truncation is to look at the end of the crawled HTML or to measure the uncompressed document size.

How does a truncated JavaScript file affect the page?

Because every resource fetched during rendering is subject to the same limit, a bundle over 2 MB is cut off in the middle. Since a half-finished script is not a valid program, rendering can break, and content generated only with JavaScript may never appear. Splitting large bundles is therefore not just a matter of speed.

Is the limit the same for every Google crawler?

No. Each client sets its limit according to its own needs. The limits of the image and video crawlers vary across a wide range, while much lower values are used for tasks such as fetching favicons.

Can this limit change in the future?

Google states explicitly that the limit is not fixed and may change as the web evolves. That is why keeping page size well away from the limit is a more resilient choice than trying to keep it just under it.

Was this article helpful?
Add Seobaz as a preferred source on Google to see us more often in your search results and AI answers.
Add as preferred source
Share this article
Turan Doğan
Founder · SEO & GEO Specialist
Publishing up-to-date guides on SEO, GEO and AEO since 2014, helping brands get seen on both Google and AI engines.
WhatsApp Online · Quick reply
Gift Wheel A discount on every spin
View Cart