Broken link checker for discovered pages on your site
Starts from your homepage, follows discovered public internal links up to your plan limit, and lists outbound links for review with anchor text, occurrence counts, and rel attributes.
Internal links checked, outbound links listed
The crawler health-checks discovered internal links with HTTP status, response time, and content size. Outbound links are listed for review but not checked.
Checked internal links also carry the page’s meta tags in the report.
Discovered automatically from your homepage
WebPixie starts at your homepage and follows links across your site, building the crawl from publicly accessible pages. No sitemap upload and no login.
The crawler visits only publicly accessible pages. It does not log in or send custom headers.
How WebPixie crawls your site
WebPixie starts from your homepage, follows discovered internal links across publicly accessible pages, and health-checks them up to your plan limit. It flags links that do not load successfully, such as a 404, a 5xx error, or a timeout. It also lists the outbound links it finds on each page, with their anchor text and occurrence count, so you can review where your site points out. You can also look up any URL to see which of your crawled pages link to it, whether that is one of your own pages or an external destination.
Beyond status codes, the Link Crawler records each internal link’s redirect status, response time, content size, title, headers, and meta tags. It shows when a link returns a redirect response, but redirect chains and targets are handled by the Main Page Analyzer for your homepage.
Crawl limits depend on your plan; compare current link limits on the pricing page. Each crawl run reports the internal links it checked and the outbound links it found. On paid plans, once Sitemap Monitoring successfully completes its first discovery cycle, crawled internal links missing from your sitemap are flagged too.
Find broken links before customers do
Crawl discovered public pages from your homepage
A broken link on a key page costs traffic, conversions, and trust. The Link Crawler starts from your homepage, follows discovered internal links, and flags links that do not return a successful response, naming 404s, 5xx errors, and timeouts. Broken internal URLs also waste search-engine crawl budget and bleed away the link equity that should flow through your pages. This catches the long tail of broken links that uptime monitoring, which checks configured endpoints, cannot see.
Get technical data beyond status codes
Redirect status, response time, and page meta tags
A successful response does not always mean healthy. A wrong canonical, a mismatched hreflang, or an unintended rel=nofollow can cause real problems. Each checked internal link records its status code, redirect status, response time, content size, and the page’s meta tags.
See your outbound link profile
External links listed for review
The crawler records outbound links it finds, with the anchor text and how many times each appears across your site. You also get the rel attributes on each link, such as nofollow, sponsored, and ugc, so you can audit how your pages point to other domains. Outbound links are listed for review, not health-checked, so this maps where your site points rather than testing those external destinations.
See which pages link to any URL
Reverse link lookup across your crawled site
Look up any URL to see which of your crawled pages link to it. Point it at one of your own pages to find its internal inbound links and map how link equity flows through your site, or point it at an external URL to find every page that links out to that destination. Each result carries the same per-link detail as the rest of the crawl.
Set up link crawling in 60 seconds
Free plan, no credit card. Compare crawl limits by plan.
Everything you need to monitor a website. In one workspace.
A quick look at other WebPixie features.
Why teams choose WebPixie for link health
Set up in 60 seconds
No agent to install. Enter a domain and WebPixie discovers links from your homepage and starts crawling on a schedule.
Your whole site in one workspace
Link health sits next to uptime, SSL, DNS, and domain monitoring, with one dashboard across all of them.
Internal health and outbound inventory
The crawler health-checks the internal links between your pages, so dead internal paths surface with their status codes, and it lists the outbound links each page points to for review in the same report.
Frequently Asked Questions
Common questions about the Link Crawler.
Yes, WebPixie checks for broken internal links with the Link Crawler. The Link Crawler starts from your homepage, follows discovered public internal links up to your plan limit, and records per-link health data such as HTTP status code, response time, content size, page title, favicon, redirect status, and timeout behavior. This helps identify 4xx errors, 5xx errors, missing pages, slow destinations, and redirect responses that can affect user experience or search visibility. The crawler also discovers the outbound links on each page and extracts their rel values (nofollow, sponsored, ugc), anchor text, and occurrence count, so you can review your external link profile, though outbound links are listed for review rather than health-checked. Canonical URLs, hreflang tags, and other meta data are recorded for the internal links WebPixie crawls. Re-crawls run on a schedule, typically about once a week, and crawl limits depend on your plan. You can compare limits on the pricing page, then pair results with the Indexability Checker for broader crawlability review.
The Link Crawler re-scans discovered links on a schedule, typically about once a week, with a minimum of 24 hours between runs. WebPixie starts by discovering links from your homepage, then schedules automatic re-crawls across the sites in your workspace. Crawl limits per site depend on your plan, and the current limits are listed on the pricing page. During each Link Crawler scan, WebPixie health-checks each internal link it crawls, recording status code, response time, content size, page title, favicon, rel attributes, canonical URLs, hreflang tags, and related meta data. Outbound links are listed for review rather than health-checked, but still carry their rel attributes, anchor text, and occurrence count. This helps you catch broken internal links, slow pages, redirect responses, and link-level issues, and review the outbound links each page points to. For broader crawlability and indexing signals, pair crawler results with the Indexability Checker.
The Link Crawler discovers and lists the external links your pages point to, but it does not health-check them. On every crawl run it follows and health-checks the internal links between pages on your monitored domain, recording status code, redirect outcome, and response time for each. For outbound links to other domains, it records the anchor text, how often each appears, and the rel attributes such as nofollow, sponsored, and ugc, so you can audit your external link profile without WebPixie testing those third-party destinations. This keeps the crawl focused on the pages you control while still giving you a full map of where your site points out. Pair it with the Indexability Checker for broader crawlability signals, and compare crawl limits on the pricing page.
Yes. The Link Crawler supports a reverse lookup: give it any URL and it returns which of your crawled pages link to that URL. Point it at one of your own pages to see that page's internal inbound links, which helps you understand your internal link structure and how link equity flows toward your most important pages. Point it at an external URL to find every page on your site that links out to that destination, which is useful when you need to update or remove a specific outbound link across your site. Each result carries the same per-link detail as the rest of the crawl, such as anchor text and the page's own status. This works from the links discovered during the normal homepage-based crawl, up to your plan limit. For broader crawlability signals, pair it with the Indexability Checker, and compare crawl limits on the pricing page.
No. The Link Crawler follows only publicly accessible links, starting from your homepage, so it does not log in or send credentials, and pages that require authentication are not crawled. If you need to check a private URL, uptime monitoring supports authentication, custom headers, and expected status codes, though it checks one endpoint at a time rather than crawling a whole site. You can review per-plan crawl limits on the pricing page.
No. The Link Crawler does not skip links that robots.txt would disallow, because it has to fetch each linked page to confirm the link works and to read the sublinks on it. Honoring robots.txt mid-crawl would leave gaps in your link health report, so the crawler checks discovered links up to your plan limit and labels robots-disallowed URLs for review. Whether a URL is eligible for search engines or AI crawlers is a separate question that belongs to indexability rather than link health. This keeps link health and crawl eligibility as two distinct signals. For the search-eligibility side, pair it with the Indexability Checker, and review per-plan crawl limits on the pricing page.
Ready to crawl your site?
Free plan, no credit card. Homepage-based crawl that checks discovered internal links.