
An XML sitemap is a structured index file that gives search engine bots an inventory of a site's crawlable URLs. As of 2026, a single sitemap file is capped at 50,000 URLs and 50 MB uncompressed. Managing these limits correctly is a technical prerequisite for using crawl budget efficiently and keeping indexing fast.
The Technical Basics of the XML Sitemap Protocol
The XML sitemap is a standard protocol defined by sitemaps.org and supported by every major search engine. The file holds each URL inside a <url> tag under the <urlset> root element. Each URL entry can define <loc> (address), <lastmod> (last modified), <changefreq> (change frequency) and <priority> (priority) fields.
Not all of these fields are required. <loc> is the only mandatory tag. Google has officially stated that it ignores <changefreq> and <priority>, so spending time filling them in is pointless. The <lastmod> tag, on the other hand, directly affects how the bot prioritizes crawling when it is kept accurate and up to date. Its value must reflect the page's real last modification date; fake dates that update automatically on every build wipe out the bot's trust in this signal.
Why the 50,000 URL Limit Exists
The 50,000 URL limit per sitemap file is an upper bound set by the protocol's designers. It keeps the memory and processing power the bot needs to parse the file under control. The 50 MB size limit follows the same logic, balancing transfer time over the network against server load.
In practice, performance problems can start long before you hit 50,000 URLs. Depending on URL length, a sitemap with 40,000 URLs can approach the 50 MB uncompressed limit. That means you need to monitor the limit in terms of file size, not just URL count. To check, run ls -lh sitemap.xml for the uncompressed size and gzip -k sitemap.xml && ls -lh sitemap.xml.gz for the compressed size.
How a Sitemap Index File Is Structured
Sites that exceed 50,000 URLs have to group multiple sitemap files under a sitemap index file. The index file lists the location of each sitemap with a <sitemap> tag inside the <sitemapindex> root element. A single sitemap index can itself reference up to 50,000 child sitemaps, which in theory allows you to manage 2.5 billion URLs.
The index file looks like this:
<?xml version="1.0" encoding="UTF-8"?>
<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<sitemap>
<loc>https://www.example.com/sitemap-posts.xml</loc>
<lastmod>2026-04-10T08:00:00+03:00</lastmod>
</sitemap>
<sitemap>
<loc>https://www.example.com/sitemap-products.xml</loc>
<lastmod>2026-04-12T14:30:00+03:00</lastmod>
</sitemap>
</sitemapindex>
Here the <lastmod> tag reports when each child sitemap was last updated. The bot uses this date to decide which sitemap it needs to reprocess. A sitemap that has not changed should keep the same <lastmod> value.
Segmenting Sitemaps by Content Type
Splitting sitemap files by content type gives you a clear advantage for both management and crawl behavior analysis. Blog posts, product pages, category pages, images and videos should live in separate sitemap files. This separation lets you track the indexing rate of each content type independently in Search Console.
A recommended segmentation looks like this: sitemap-posts.xml for blog posts and articles, sitemap-products.xml for product pages, sitemap-categories.xml for category and taxonomy pages, and sitemap-images.xml for images. On e-commerce sites where product variants have their own URLs, moving them out of the main product sitemap into sitemap-product-variants.xml makes their indexing status much easier to follow.
Using Lastmod Correctly (and How It Gets Misused)
The <lastmod> tag is one of the strongest sitemap signals the bot uses to prioritize crawling. It should be updated only when the page content actually changes. A change to the navigation menu or footer does not change the page content and does not call for a <lastmod> update.
The most common misuse is updating the <lastmod> value of every URL automatically on each build or deployment. This erodes the bot's trust in the <lastmod> signal. When the bot sees every date change at once, it treats the data as meaningless and falls back on its own crawl prioritization. The lastmod value should be updated only when the content genuinely changes. The right way to enforce this in a CMS is to define a "content modified date" field that is separate from the publish date and to tie the sitemap generator to that field.
Where Changefreq and Priority Stand Today
Google has announced that it officially ignores the <changefreq> and <priority> tags. Although they were defined in the first version of the protocol, in practice they do not influence crawl behavior. The bot learns how often a page changes from its own crawl history; it does not trust what the sitemap declares.
So you do not have to keep these tags in your sitemap files. Removing them shrinks the file and shortens parse time. In a 50,000 URL sitemap, dropping <changefreq> and <priority> from every URL reduces file size by about 15-20% on average. That saving can be decisive for large sitemap files approaching the 50 MB limit.
Include Only 200 OK URLs in Your Sitemap
Every URL in a sitemap tells the crawler "crawl and index this page." Including pages that return a 301 redirect, throw a 404, carry a noindex tag or have a canonical pointing to another URL sends conflicting signals. The bot reports these conflicts as warnings in Search Console.
The core rule of sitemap hygiene is simple: the file should contain only URLs that return a 200 HTTP status code and are indexable and canonical. To enforce this, add a validation layer to your sitemap generation process. The script or plugin that builds the sitemap should check the HTTP status code, the meta robots tag and the canonical URL before adding a URL to the list. This check should be automated, especially in CMS setups that generate sitemaps dynamically.
Declaring the Sitemap in Robots.txt
The location of the sitemap is declared in robots.txt with the Sitemap: directive. This speeds up the bot's discovery of the sitemap. Submitting the sitemap through Search Console is not an alternative to this declaration but a complement to it. Using both methods together is recommended.
The declaration, added at the bottom of robots.txt, should follow this format:
Sitemap: https://www.example.com/sitemap-index.xml
If you have more than one sitemap index file, declare each on its own line. The sitemap URL must be a full (absolute) URL; relative paths are not supported by the protocol. The sitemap URL must also be on the same host as robots.txt; you cannot declare a sitemap on a different subdomain or domain from robots.txt.
Dynamic vs. Static Sitemap Generation
Sitemaps can be served as static XML files or generated dynamically on every request. Dynamic generation is acceptable for small and mid-sized sites (under 10,000 URLs). On large sites, however, running a database query to build the sitemap on every bot request hurts server performance.
For large sites, the recommended approach is to generate the sitemaps as static files periodically (once or twice a day) and cache them on the server. In WordPress, Yoast SEO and Rank Math handle this automatically. In a custom CMS, the most efficient route is to write the sitemap to disk as an XML file with a cron job and have Nginx serve it directly as a static file. This way, sitemap requests never touch the database.
Keeping Sitemap Size in Check with Gzip
XML sitemaps can be served gzip-compressed. Compressed files use the .xml.gz extension and are supported by all major search engines. A typical XML sitemap shrinks by 60-80% with gzip.
Compression is an important tool for files approaching the 50 MB limit. One thing to keep in mind: the 50 MB limit applies to the uncompressed size. Compression only shortens transfer time; it does not let you exceed the limit. To enable gzip for sitemap files in Nginx, adding the line gzip_types application/xml text/xml; to the configuration is enough. Alternatively, you can pre-compress the sitemap and serve it as an .xml.gz file.
Getting Namespace Declarations Right (and Common Mistakes)
The namespace declaration in an XML sitemap states which protocol version the file conforms to. The standard sitemap namespace is http://www.sitemaps.org/schemas/sitemap/0.9. Image sitemaps need the image namespace, video sitemaps the video namespace and news content the news namespace.
The most common mistake is using HTTPS in the namespace URL. The URL in the protocol definition must stay HTTP; this is not a security hole but part of the XML schema definition. The second common mistake is leaving out the namespaces required for multiple content types. For example, if a sitemap containing image URLs does not declare the xmlns:image namespace, the bot cannot discover those images. A correct namespace declaration:
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9"
xmlns:image="http://www.google.com/schemas/sitemap-image/1.1">
Image Sitemap Integration and Image Discovery
Images can be declared inside the main sitemap with the <image:image> tag or in a separate image sitemap. Up to 1,000 images can be listed under each <url> tag. This integration directly improves indexing speed and coverage in Google Images.
An image sitemap looks like this:
<url>
<loc>https://www.example.com/article-title</loc>
<image:image>
<image:loc>https://www.example.com/images/image-name.webp</image:loc>
</image:image>
</url>
For images served through a CDN, the <image:loc> tag should point to the CDN URL. However, the CDN domain must be verified in Search Console or share the same origin as the main domain. CDN images on a different domain may not get indexed even if you list them in the sitemap.
Video Sitemap Configuration and Required Fields
Creating a separate video sitemap for video content improves visibility in video search results. A video sitemap is built either by adding the xmlns:video namespace to a standard sitemap or as a standalone file. Each video entry requires <video:thumbnail_loc>, <video:title>, <video:description> and either <video:content_loc> or <video:player_loc>.
A video sitemap tells Google clearly where a video sits on the page and what it covers; it makes discovery noticeably easier for videos that load late through JavaScript or are hard to spot on the page. Declaring the video length in seconds with the <video:duration> tag helps the bot categorize the video more accurately. Sites that embed YouTube videos put the YouTube player URL in <video:player_loc>; sites that host their own videos put the direct video file URL in <video:content_loc>.
News Sitemaps and Google News Integration
For news publishers, a Google News sitemap requires a structure separate from the standard sitemap. A news sitemap should contain articles published in the last 48 hours and be declared with the xmlns:news namespace. Every entry must include the <news:publication>, <news:publication_date> and <news:title> fields.
The key difference between a news sitemap and a standard one is timeliness. Content older than 48 hours should be removed from the file. The bot crawls news sitemaps far more often than standard ones, so the file has to be generated dynamically and updated the moment a new story is published. Static generation is not enough for this use case; the CMS needs to update the news sitemap automatically at publish time.
Managing Multilingual URLs with Hreflang Sitemaps
On multilingual sites, hreflang annotations can be placed in the HTML <head>, in HTTP headers or inside the sitemap. The sitemap method is the most scalable approach, especially for multilingual sites with thousands of pages, because you do not have to edit each page's HTML.
An hreflang sitemap is declared with the xmlns:xhtml namespace:
<url>
<loc>https://www.example.com/tr/sayfa</loc>
<xhtml:link rel="alternate" hreflang="tr" href="https://www.example.com/tr/sayfa"/>
<xhtml:link rel="alternate" hreflang="en" href="https://www.example.com/en/page"/>
<xhtml:link rel="alternate" hreflang="de" href="https://www.example.com/de/seite"/>
</url>
Each language version's <url> block must list all alternates and itself (self-referencing). If this cross-referencing is incomplete, the bot cannot validate the hreflang relationship and ignores the signal. A site in 10 languages needs 10 hreflang lines per page, which quickly bloats the sitemap and increases the need for segmentation.
Sitemap Splitting Strategy for Large Sites
For sites with more than 100,000 URLs, sitemap management is about more than just splitting at the 50K limit. The splitting strategy should be shaped by crawl prioritization and index tracking needs. Splitting by content type (posts, products, categories) is the base layer. As scale grows, time-based splitting becomes necessary too.
An effective approach on e-commerce sites is to split product sitemaps into child files by category or by date added: sitemap-products-electronics.xml, sitemap-products-fashion.xml, or sitemap-products-2026-q1.xml, sitemap-products-2026-q2.xml. This lets you track the indexing rate of each segment independently in Search Console. If one category has a low indexing rate, you investigate that segment specifically instead of wasting time across the whole site.
Reading the Search Console Sitemaps Report
The Sitemaps report in Search Console shows the status of submitted sitemaps, the number of discovered URLs and the number of indexed URLs. The gap between "Submitted" and "Indexed" is the most direct indicator of sitemap hygiene problems.
You can reach the report via Search Console > Sitemaps. The "Status" column next to each sitemap should read "Success." "Has errors" indicates an XML syntax error, an unreachable URL or an unsupported format in the file. When the gap between submitted and indexed URLs exceeds 20%, there is a systematic problem. In that case, cross-reference the unindexed URLs with the "Not indexed" reasons in the Page indexing report to find out why they were left out (this report used to be called "Coverage").
Notifying Search Engines: What Replaces Ping?
Google announced in June 2023 that it was deprecating the old sitemap ping endpoint (google.com/ping); requests sent to that address no longer do anything. Ping examples found in older guides should therefore not be used.
There are two current ways to tell Google about a sitemap: add a Sitemap: https://www.example.com/sitemap-index.xml line to robots.txt, and submit the sitemap in the Sitemaps screen of Search Console. Once Google knows a sitemap, it rereads it on its own schedule; keeping the <lastmod> value accurate is enough for changed pages to be noticed. If you manage many sitemaps, submission can be automated with the sitemaps endpoint of the Search Console API. For Bing and Yandex, the IndexNow protocol provides instant notification.
WordPress Sitemap Setup and Plugin Choice
Since version 5.5, WordPress has offered a built-in XML sitemap. This built-in sitemap, however, is limited in segmentation flexibility and advanced filtering. Yoast SEO and Rank Math offer advanced features such as automatic segmentation by content type, automatic filtering of noindex pages and control over the number of sitemap files.
Our hands-on experience shows that while the built-in WordPress sitemap is enough for small blogs, switching to the Yoast SEO or Rank Math sitemap module is far more efficient operationally on sites with more than 5,000 URLs. In the plugin settings, check one by one whether each content type should be included in the sitemap. Tag pages, author archives and date-based archives usually contain thin content; removing them from the sitemap steers the bot's attention toward your valuable pages.
Headless CMS and Custom Sitemap Generation
In headless CMS architectures (frameworks such as Next.js, Nuxt.js and Gatsby), sitemap generation works differently from traditional CMSs. In these setups, the sitemap is either generated statically at build time or served dynamically through an API. With build-time generation, the sitemap is rebuilt on every deployment; this approach is ideal for static sites.
The next-sitemap package is widely used for sitemap generation in Next.js. The siteUrl, generateRobotsTxt and exclude parameters are defined in next-sitemap.config.js. This configuration automatically creates segmented sitemap files and robots.txt after the build. On sites using ISR (Incremental Static Regeneration), generating the sitemap only at build time is not enough; you also need a server-side dynamic sitemap endpoint so newly added pages make it into the sitemap.
Validating and Debugging Sitemap XML
A single XML syntax error in a sitemap causes the bot to reject the entire file. An unclosed tag, an incorrectly encoded special character or a missing namespace declaration invalidates the whole file. That is why validating the XML after every sitemap update is a must.
The fastest way to validate is the xmllint --noout sitemap.xml command, which reports XML syntax errors with line and column numbers. The ampersand (&) in URLs is the most frequent source of errors; in XML it must be encoded as &. Turkish characters (ş, ç, ğ, ü, ö, ı) work fine with UTF-8 encoding, but the file must include the XML declaration <?xml version="1.0" encoding="UTF-8"?>.
URL Types to Exclude from Your Sitemap
Keeping certain URL types out of your sitemap is the most direct way to protect crawl budget. These include:
- Pages with a noindex tag: Adding URLs with a meta robots noindex directive to the sitemap puts the bot in a "crawl but don't index" contradiction and wastes crawl requests.
- Pages whose canonical points to another URL: Non-canonical URLs should not be in the sitemap; even if the bot crawls them, it will end up at the canonical target.
- Parameter-based duplicate URLs: URLs with sorting, filtering and session parameters are copies of the main page and bloat the sitemap for no reason.
- Soft 404s and thin content pages: Low-value pages with only a few sentences, or empty pages, should be removed from the sitemap; indexing them lowers the site-wide quality score.
- Deep pagination pages: /page/2/ and beyond are usually unnecessary in the sitemap; the content should already be listed with its own URLs in the main sitemap.
Sitemap Access Rules for CDNs and Other Domains
A sitemap must be served over the same protocol and host as robots.txt. A sitemap declared in https://www.example.com/robots.txt can only contain URLs under https://www.example.com/. URLs on a different subdomain (for example https://blog.example.com/) cannot be in this sitemap; that subdomain needs its own sitemap and robots.txt.
The sitemap file itself can be served through a CDN, but the CDN URL must be configured as an alias of the main domain. A sitemap at https://cdn.example.com/sitemap.xml is valid only if that address serves the same content as https://www.example.com/sitemap.xml and the declaration in robots.txt is correct. Otherwise, the bot may reject a sitemap served from a different host.
How Sitemap Update Frequency Affects Crawl Behavior
How often a sitemap is updated indirectly affects how often the bot checks it. Frequently updated sitemaps get higher priority in the bot's crawl queue. That prioritization does not hold, however, for updates that contain no real changes. If the bot finds nothing new each time it checks the sitemap, it checks less often.
The right strategy is to update the sitemap only when a URL is actually added or removed. The most consistent way to do this is to hook the update into the CMS's publish and delete events. In WordPress, a function attached to the save_post and delete_post actions can rebuild the sitemap automatically. After each update, changing the <lastmod> value only for the child sitemaps that changed prevents the bot from reprocessing unchanged files.
Making Sitemap Audits Part of Your Monthly Technical Checks
Sitemap management is not a one-time setup but an ongoing maintenance process. As content is added and deleted and URLs change, the sitemap has to reflect those changes accurately. Your monthly technical audit cycle should include three basic sitemap checks.
The first step is comparing the number of URLs in the sitemap with the number of indexed URLs in Search Console. The second is running a bulk HTTP status check to verify that every URL in the sitemap returns a 200 status code. You can automate this with wget --spider --force-html -i url-list.txt -o results.log 2>&1. The third is checking whether URL types that should not be in the sitemap (noindex, non-canonical, parameter-based) have slipped in. These three steps keep sitemap hygiene sustainable and help steer crawl budget toward your valuable pages.



