Why your sitemap and your crawl disagree
A sitemap is a snapshot of what your site looked like when it was generated. A crawl reflects your site right now. The two drift apart continuously, not just after a big migration.
A sitemap is a snapshot of what your site looked like when it was generated. A crawl reflects your site right now. The two drift apart continuously, not just after a big migration.
A weekly link crawl catches most broken links before they cost you traffic. When you find them, internal 404s on your highest-traffic pages come first, then redirect chains, then outbound links last.
AI agents now crawl, read, and act on websites the way search engines used to. Here is a practical, standards-graded checklist for making your site agent-ready, and what to skip for now.
llms.txt still does almost nothing for AI search citations, and a 137,000-site study found 97% of files never get fetched. But in 2026 Google moved it into a new agentic browsing standard alongside WebMCP. Here is what changed, and when llms.txt is still worth adding.
GEO is SEO with one extra job: getting cited inside AI assistant responses. The 2023 paper that coined the term tested 9 content interventions. Four worked strongly, three gave moderate gains, and two flatlined or backfired, including classic keyword stuffing.
Most sites should allow AI training crawlers in 2026: invisibility in AI assistant answers now costs more than uncrawled-for-training saves. Here is the per-provider breakdown (GPTBot, ClaudeBot, Google-Extended) and the robots.txt for each realistic decision.
Three small files with three different jobs. robots.txt controls who crawls, sitemap.xml maps what to crawl, and llms.txt curates what AI assistants read first. Get the mental model right with minimum-viable examples and the mistakes that confuse them.