Skip to main content
Back to Features
Sitemap monitoring

Sitemap monitoring for complete coverage visibility

WebPixie discovers your sitemap tree, catalogs declared URLs, and cross-checks it against crawled internal links to surface coverage gaps.

XML, text, and gzip sitemaps
Cross-checked against Link Crawler
1 site · sitemap tree discovered
Sitemap fileURLsLast checkedStatus
/sitemap-index.xml4 files1hDiscovered
/sitemap-pages.xml861hOK
/sitemap-blog.xml2141hOK
/sitemap-products.xml.gz1,1406dNo validator
3 crawled links not found in sitemap
Sitemap Monitoring is available on paid plans.Compare all plans

Sitemap tree discovered from robots.txt

WebPixie starts from the Sitemap: directive in your robots.txt and walks supported indexes down to leaf files. Nested or recursive indexes are flagged as unsupported.

/sitemap-index.xmlindex
/sitemap-pages.xml86 URLs
/sitemap-blog.xml214 URLs
/sitemap-products.xml.gz1,140 URLs

Each leaf sitemap is cataloged with its declared URLs, lastmod, changefreq, and priority when present.

Crawled links cross-checked against your sitemap

Once Sitemap Monitoring successfully completes its first full discovery cycle, crawled internal links missing from your sitemap get flagged.

URLLastmodPriorityStatus
/pricing2026-06-020.8In sitemap
/blog/launch2026-05-140.5In sitemap
/new-landing-pageNot found in sitemap
/old-campaign-2024Not found in sitemap

A coverage signal, not a broken-link check: a missing row means the page is not declared in your sitemap, not that it failed to load.

How WebPixie tracks your sitemap

WebPixie discovers your sitemap tree automatically from the `Sitemap:` directive in your robots.txt, then walks supported sitemap indexes down to leaf files. It supports XML and text sitemap files, including gzip-compressed responses, and catalogs declared URLs per leaf sitemap along with lastmod, changefreq, and priority when present. Nested or recursive sitemap indexes are not supported and are flagged as unsupported.

The discovery cycle re-checks daily as part of your site’s daily snapshot. An individual sitemap file is re-fetched daily when it supports cache validators such as ETag or Last-Modified, and otherwise can go up to about a week between fetches.

Once Sitemap Monitoring has successfully completed its first full discovery cycle, WebPixie cross-checks it against the Link Crawler: crawled internal links missing from your sitemap get flagged, so a page engines may never discover through your sitemap does not go unnoticed.

01

Discover your whole sitemap tree automatically

From robots.txt Sitemap: directive to supported leaf files

There is nothing to configure. WebPixie reads the `Sitemap:` directive in your robots.txt, then walks supported sitemap indexes down to leaf files, cataloging declared URLs along with lastmod, changefreq, and priority when present. XML, text, and gzip-compressed sitemap files are all supported. Nested or recursive sitemap indexes are not supported yet and are flagged as unsupported so you know coverage is incomplete there.

02

Find pages missing from your sitemap

Cross-checked against crawled internal links

A sitemap generator can miss a new page, or drop an old one that is still linked internally. Once Sitemap Monitoring successfully completes its first full discovery cycle, WebPixie compares its catalog against the internal links the Link Crawler finds, and flags any crawled internal link that is not present in your sitemap. That surfaces a page search engines may never discover through your sitemap, before it costs you traffic.

03

Stay current without re-fetching everything daily

Conditional GET on ETag and Last-Modified

The discovery cycle re-checks your sitemap tree daily as part of your site’s daily snapshot. When an individual sitemap file supports cache validators such as ETag or Last-Modified, WebPixie re-fetches it daily using a conditional GET. When it does not, that file is left as-is for up to about a week between fetches, so large sitemap trees stay efficient to track.

04

See lastmod, changefreq, and priority per URL

The metadata your sitemap already declares

Sitemap files can declare a lastmod date, a changefreq hint, and a priority value per URL. WebPixie catalogs these fields when present, alongside the URL itself, giving you a structured view of what your sitemap says about each page without opening the raw XML.

Add sitemap coverage to your monitoring

Available on paid plans. Compare plans on the pricing page.

Everything you need to monitor a website. In one workspace.

A quick look at other WebPixie features.

Why teams choose WebPixie for sitemap coverage

No sitemap upload needed

WebPixie discovers your sitemap tree from the Sitemap: directive in your robots.txt and walks supported index files down to leaf sitemaps automatically.

Your whole site in one workspace

Sitemap coverage sits next to the Link Crawler, indexability signals, uptime, SSL, DNS, and domain monitoring, in one dashboard.

Surface coverage gaps, not just broken links

The Link Crawler finds broken links; Sitemap Monitoring surfaces pages missing from your sitemap, so search engines have a better chance of discovering them.

Frequently Asked Questions

Common questions about Sitemap Monitoring.

Sitemap Monitoring discovers your sitemap tree automatically from the Sitemap: directive in your robots.txt, then walks supported sitemap indexes down to leaf files. It supports XML and text sitemap files, including gzip-compressed responses, and catalogs declared URLs per leaf sitemap along with lastmod, changefreq, and priority when present. Nested or recursive sitemap indexes are not supported and are flagged as unsupported. The discovery cycle re-checks daily as part of your site’s daily snapshot; an individual sitemap file is re-fetched daily when it supports cache validators such as ETag or Last-Modified, otherwise it can go up to about a week between fetches. Sitemap Monitoring is available on Starter and above; compare plans on the pricing page.

Once Sitemap Monitoring has successfully completed its first full discovery cycle, WebPixie cross-checks it against the Link Crawler: crawled internal links that are not present in your declared sitemap are labeled "not found in your sitemap". This surfaces pages engines may never discover through your sitemap, such as a new page you forgot to add or an old URL your sitemap generator dropped. It is a coverage signal, not a health check: the label only compares sitemap membership against crawled internal links, it does not mean the page is broken or blocked. Sitemap Monitoring is available on Starter and above; see the pricing page for details.

The Indexability Checker only reports whether a Sitemap: directive is present in your robots.txt, a presence check. Sitemap Monitoring goes further: it actually fetches the declared sitemap tree, walks supported sitemap indexes down to leaf files, and catalogs declared URLs with lastmod, changefreq, and priority when present. It also cross-checks sitemap coverage against the Link Crawler, flagging crawled internal links that are missing from your sitemap. Sitemap Monitoring is available on Starter and above, while the robots.txt presence check is included on every plan; compare plans on the pricing page.

No. Sitemap Monitoring surfaces its findings in the dashboard for review; it does not open incidents or send notifications. WebPixie's automated incidents and alerts come only from uptime, SSL, and domain events, such as a site going down or a certificate failing its check. Sitemap Monitoring instead catalogs your declared URLs and flags things worth reviewing, like a sitemap index it could not fully process or a crawled internal link that is missing from your sitemap. Those coverage signals sit next to your other site data so you can act on them during review, rather than as a paged alert. If you want events that notify you, incident management tracks the availability and certificate events above. Sitemap Monitoring is available on Starter and above; compare plans on the pricing page.

Sitemap Monitoring supports standard XML and plain-text sitemap files, including gzip-compressed responses. During discovery it reads your sitemap index and walks the sitemaps it references down to the leaf files, then catalogs every declared URL along with its lastmod, changefreq, and priority when present. The sitemap protocol allows up to 50,000 URLs and 50 MB per file; on top of that, WebPixie applies its own processing caps, reading up to 10,000 URLs per sitemap file, up to 100 sitemap files per site each discovery cycle, and up to 50,000 sitemap URLs per site in total, while a file over 50 MB is skipped as too large. Most sites stay well within these limits, and splitting a large sitemap into multiple files under one index keeps every page covered. WebPixie discovers your sitemaps automatically from the Sitemap: directive in robots.txt, and that reference is also validated by the Indexability Checker. For more on discovery, see Sitemap Monitoring, available on Starter and above on the pricing page.

No, and neither do search engines. A sitemap index file may list individual sitemap files, but it cannot list other sitemap index files, so a nested or recursive index (an index that points to another index) is outside the sitemap protocol. Google states this directly: a sitemap index file can't list other sitemap index files, only sitemap files (opens in new tab). Sitemap Monitoring follows the same rule: it walks a normal sitemap index down to its leaf files, but when it meets a nested sitemap index it flags that entry as unsupported instead of following it further. If you see that flag, the fix is to flatten the structure so your top-level index points directly at sitemap files, which also keeps your sitemaps discoverable by search engines. Sitemap Monitoring is available on Starter and above; compare plans on the pricing page.

Ready to review your sitemap coverage?

Discover your sitemap tree and surface links missing from it, cross-checked against the Link Crawler.