Skip to content
Guide contents

User guide4 min

Crawlability

Everything that decides what a crawler may fetch: robots.txt, your sitemaps, your redirects and the indexing rules on each page. They share a screen because they often contradict each other.

How to get there#

Explore → Crawlability.

  1. 1

    Open Crawlability in the sidebar

    After Structure.

  2. 2

    Start on Overview - the crawler journey

    Access, then discovery, then reach, then the redirect path - in that order.

app.pixyscan.com/w/…/s/…/crawlability

The Crawlability overview: the crawler journey - crawler access, discovery, reach, redirect path - above the site's crawl settings card, with six lens tabs and export actions in the header.
The Overview tab, shown as a journey. Each stage only matters once the one before it has passed.

The six tabs#

A summary, then one tab per thing a crawler reads.

FieldAnswersWhat it does
OverviewCan a crawler get round?The journey, plus the crawl settings this scan actually used.
robots.txtWho is allowed?Your robots.txt rules as served, then a grid of named AI crawlers marked Allowed or Blocked.
SitemapsWhat did you declare?Every sitemap file the crawler followed, with the URLs it lists, the URLs actually reached, and how many carry a lastmod date.
RedirectsWhere does it bounce?Redirect chains and loops, with the status code at each step. A 307 where a 301 belongs shows up here.
DirectivesWhat does each page say?The noindex, nofollow, canonical and robots meta values on each page.
InternationalWhich language, for whom?Your hreflang declarations, and whether each one points back.

robots.txt

The robots.txt lens: stanzas with their Allow and Disallow rules, then a grid of AI crawlers each marked Allowed or Blocked with its vendor and class.
The robots.txt tab: the file as served, then what it actually means for each named crawler.

What to look for#

Places where robots.txt and your sitemap disagree.

A URL listed in your sitemap but disallowed in robots.txt tells a search engine two opposite things. So does a sitemap full of URLs that redirect. The Sitemaps tab shows URLs listed against URLs reached so that gap is a number you can see.

Decide about AI crawlers deliberately

The robots.txt tab names each AI crawler separately because most sites inherited their robots.txt and never chose. Whichever way you decide, AI readiness shows the effect.

If the product does not match this page, the page is wrong and we would like to know. Tell us