Skip to main content
Back to Features
Indexability checker

Indexability checker for search engines and AI tools

Review the signals that affect whether search engines and AI tools can discover, crawl, and index your important pages, with robots.txt, llms.txt, canonical, hreflang, and link coverage checks.

Search engines and AI tools
Robots, canonical, hreflang
Core signals on every plan
Indexability signals · TXT files and page checks
Page / fileCheckResultChecked
robots.txtSyntaxValid1d
/pricingMeta robotsnoindex found1d
/blog/launchCanonicalMismatch1d
/old-productSitemap coverageNot found in sitemap1d
noindex on /pricing · canonical mismatch on /blog/launch
Core indexability signals are included on every plan. Sitemap coverage is available on paid plans.Compare all plans

Indexability signals across WebPixie

Review the live signals surfaced by TXT Files, Main Page, Link Crawler, and Sitemap Monitoring.

SourceSignalStatus
TXT Filesrobots.txtValid
TXT Filesllms.txtAccessible
Main PageCanonical / alternate signalsReviewed
Link Crawler/admin · robots.txt disallowedSignal found
Sitemap Monitoring/old-product · not found in sitemapCoverage gap

These are indexability-related signals from live feature views, not a single overall indexable or blocked verdict.

How WebPixie checks your indexability

WebPixie surfaces indexability-related signals across the relevant feature views. TXT Files validates robots.txt and llms.txt line by line, Main Page surfaces canonical, alternate, and header signals, Link Crawler labels crawled URLs that robots.txt disallows, and Sitemap Monitoring can flag crawled internal links missing from your sitemap.

A single accidental noindex tag or one restrictive robots.txt line can quietly remove a page from search results with no error to warn you. WebPixie surfaces these signals in the relevant views, so you can review a noindex directive, a Disallow rule, or a conflicting canonical or hreflang before it turns into a long debugging session.

Coverage goes past classic SEO. WebPixie also validates llms.txt, so you can review the guidance you publish for AI tools as that traffic grows.

01

Review accidental noindex and robots blocks

One wrong directive can drop a page from search

A noindex meta tag left on after a redesign, or a Disallow line added to robots.txt during a deploy, can remove a page from search results without any error. WebPixie surfaces indexability signals across the relevant views, so you can review the directive where that signal is collected instead of treating a single label as the diagnosis.

02

See conflicting canonical and hreflang signals

Canonical and hreflang tell engines which version wins

When a canonical tag points to the wrong URL, or hreflang alternates do not reciprocate, engines may favor a version you did not intend or split signals across near-identical URLs. WebPixie surfaces canonical and hreflang signals from the relevant page and link analyses, so you can review which URL engines are being asked to treat as the original.

03

Review guidance for AI tools, not only search

llms.txt availability and published guidance

As AI assistants send more referral traffic, the guidance you publish for them matters alongside classic indexing. WebPixie checks whether llms.txt is accessible and validates its contents next to robots.txt, so you can review the rules and instructions you publish for AI tools.

04

Confirm your robots.txt, sitemap, and link rel signals

robots.txt sitemap reference, sitemap coverage, rel nofollow, ugc, sponsored

WebPixie checks whether a sitemap is referenced in your robots.txt and reads the rel attributes on your links, including nofollow, ugc, and sponsored. On paid plans, Sitemap Monitoring crawls the declared sitemap tree and flags crawled internal links missing from your sitemap, so sitemap coverage gaps are visible alongside link-level signals.

Check your indexability in 60 seconds

Free plan, no credit card. Sitemap coverage is available on paid plans.

Everything you need to monitor a website. In one workspace.

A quick look at other WebPixie features.

Why teams choose WebPixie for indexability

Set up in 60 seconds

No agent to install and no access to your server needed. Enter a domain and WebPixie checks indexability from its own servers.

Your whole site in one workspace

Indexability sits next to uptime, SSL, DNS, domain, and link health, with one dashboard across all of them.

Search and AI coverage together

Search indexing and AI-engine access signals appear in the same workspace, so you can review both without switching tools.

Frequently Asked Questions

Common questions about the indexability checker.

The Indexability Checker covers signals that affect whether search engines and AI tools can discover, crawl, and index your important pages. TXT Files validates robots.txt and llms.txt, Main Page surfaces canonical and alternate signals, Link Crawler contributes page-level robots access signals, and Sitemap Monitoring contributes sitemap coverage signals. Review each signal in the feature view where it is collected, alongside the Link Crawler and Main Page Analyzer details that provide its context.

Uptime affects SEO by keeping your pages reachable for search engine crawlers and real users. A brief outage may not cause an immediate ranking change, but repeated downtime can interrupt crawling, reduce trust in page availability, waste crawl opportunities, and create poor user signals when visitors land on errors. Uptime monitoring helps you detect access problems quickly, including timeouts, unexpected status codes, and missing expected content. For marketers, this matters most on high-value pages such as campaign landing pages, product pages, checkout flows, and content that receives frequent organic traffic. WebPixie pairs uptime checks with multiple check attempts and cross-location verification on paid plans to reduce false positives before alerting your team. For search visibility, combine uptime checks with the Indexability Checker and Link Crawler to catch blocked pages, broken links, and crawlability issues.

Yes. The Indexability Checker validates and analyzes robots.txt and llms.txt line by line, giving each file a valid or invalid result and flagging syntax it cannot parse. This lives in the TXT Files view, where you can see how a directive is interpreted instead of guessing whether a rule is written correctly. It also checks whether a sitemap is referenced in robots.txt, so a missing reference is visible during review. On the Main Page view, WebPixie surfaces the canonical URL and alternates as supporting visibility signals rather than a standalone indexing verdict. Indexability signals contribute to the WebPixie Site Score without requiring you to inspect each technical finding manually. For link-level crawl health, pair it with the Link Crawler, and compare plan limits on the pricing page.

WebPixie alerts you when a monitored page becomes unreachable or fails the conditions you configured. No monitoring tool can guarantee that a search engine has not attempted to crawl during an outage, but fast detection helps reduce the duration and SEO risk of availability problems. Uptime monitoring checks for timeouts, connection failures, unexpected status codes, redirects, and missing expected content, then uses multiple check attempts and cross-location verification on paid plans to reduce false positives before opening an incident. Alerts are sent by email on every plan, with Slack and webhooks available on eligible plans, so the right team can act quickly. Confirmed outages can be tracked through incident management, including severity, affected resource, and closure timing. For search-specific risk, combine outage alerts with the Indexability Checker to catch crawl-blocking configuration problems.

They answer different questions. The Link Crawler is about link health: starting from your homepage it follows discovered public internal links and reports each checked link's status code, response time, redirect status, and whether it loads successfully, while also listing the outbound links each page points to without checking them. The Indexability Checker is about crawl signals: it validates robots.txt and llms.txt, reports robots.txt sitemap-reference presence, uses canonical and alternate signals from the Main Page view, and includes crawler-derived robots-disallowed tags. A link can be healthy yet blocked from search engines, and a page can be visible to crawlers yet link to broken destinations, so the two signals do not replace each other. The crawler deliberately ignores robots.txt while following links, because it must fetch a page to confirm it works; search eligibility is reviewed through indexability signals instead. Many teams run both and compare plan limits on the pricing page.

The Indexability Checker only reports whether a Sitemap: directive is present in your robots.txt, a presence check. Sitemap Monitoring goes further: it actually fetches the declared sitemap tree, walks supported sitemap indexes down to leaf files, and catalogs declared URLs with lastmod, changefreq, and priority when present. It also cross-checks sitemap coverage against the Link Crawler, flagging crawled internal links that are missing from your sitemap. Sitemap Monitoring is available on Starter and above, while the robots.txt presence check is included on every plan; compare plans on the pricing page.

Ready to review your indexability signals?

Free plan, no credit card. Core signals on every plan, with sitemap coverage on paid plans.