Skip to content
SolutionsAI search visibility

Whether answer engines can reach, parse and quote you

Three things have to be true before a model can cite you: it has to be allowed in, what it fetches has to be parseable, and the passage it wants has to be liftable. All three are observable. Whether it then cites you is not.

app.pixyscan.com/w/…/s/…/ai-readiness
The AI readiness screen: three weighted factors scored together with their contributions shown.
Sound familiar?

If any two of these are true, this page is for you

None of them is unusual. They are what happens to a site that ships regularly and is only ever looked at deliberately.

  • Traffic from assistants is showing up in the numbers and nobody knows what drives it.
  • Somebody added an llms.txt and it is not clear whether that did anything.
  • The robots.txt has AI bot rules in it that predate the current position on AI.
  • Every vendor in this space is promising citations, which is not a thing anyone can promise.

What PixyScan does

3 things, specifically

Each one is a mechanism in the running product, not a positioning statement.

  1. 01

    Bot access as a matrix

    Allow, disallow or not-set for each AI crawler, classed by whether it is there to retrieve an answer now, to fetch a link a user pasted, or to train a model. Blocking one of those and not the others is a coherent position; it should be a deliberate one.

  2. 02

    Three weighted factors, formula published

    Answer signals (70) - liftable answers, AI crawler access, llms.txt and extractable prose together - machine-readable meaning (20) and document structure (10). Every weight is on the screen next to the group it applies to.

  3. 03

    Schema validated per type

    22 checks parse and validate JSON-LD rather than detecting its presence. A model reading the page finds structure instead of inferring it.

How it fits your week

Day one, then every week after

When you would actually open it. A tool you have to remember is a tool that gets forgotten, so most of this runs without you.

  1. Step 1

    Day one: access, then structure

    Open the AI readiness screen and read AI engine access first - if the retrieval crawlers are blocked, nothing else matters yet. Then read the Crawlability matrix bot by bot and decide, deliberately, which classes you are admitting.

  2. Step 2

    The first month: make the content liftable

    Work down the answer signals and machine-readable meaning groups: reading level against audience, comparison tables where a page promises a comparison, JSON-LD that validates per type. Publish llms.txt and llms-full.txt and re-scan; they are graded on content, so a first draft tells you what to fix.

  3. Step 3

    Every week after: the score as a trend

    On a schedule, the readiness score becomes a line on the overview, and any template change that breaks structured data or drops a landmark is a New finding that week. The signals under your control stay under your control.

Worth knowing

  • Reading level measured against the audience the page is written for, not in the abstract.
  • llms.txt and llms-full.txt graded on content - format and whether their links resolve - not existence.
  • Semantic landmarks checked, because structure is how an extractor tells answer from furniture.
  • Searchable robots.txt: type a user-agent and read the exact stanza that applies to it.

What it will not do for you

  • It does not predict whether a model will cite you. Nothing can, and a tool that says otherwise is selling something.
  • No measurement of your appearances in AI answers - this reports the signals under your control.
  • No per-page “how much of this can AI see” score. The HTTP engine reads what the server sent and the Playwright engine audits the rendered DOM, but neither predicts what a model would lift.

Above the signup button on purpose

A month in

What is different four weeks later

Each of these is something you could check, not something you would have to take on trust.

  • A deliberate, written-down position on which AI crawlers are admitted, and evidence that robots.txt implements it.
  • Structured data that validates per type on every page that declares it.
  • llms.txt and llms-full.txt published, well-formed, and pointing at pages that resolve.
  • A readiness score with its three weights visible - answer signals 70, machine-readable meaning 20, document structure 10 - tracked over time rather than measured once.

Before you ask

Things people in your position ask first

Plans, edges, and the honest answer to the question this page is most often found by.

Will this get us cited by ChatGPT or Perplexity?

Nobody can promise that, and we do not. What this measures is whether you are reachable, parseable and quotable - the preconditions you control. Whether a model then chooses to cite you is not observable from your side.

What is the difference between GEO and AEO here?

Two different questions, and the checks sit in different places. Access and publication - AI bot rules in robots.txt, llms.txt and llms-full.txt - are read on the Crawlability screen. The content's shape - reading level for the audience, comparison tables where one is promised, prose an extractor can lift - is the Content & Answers discipline. Both feed the readiness score, which weights answer signals 70, machine-readable meaning 20 and document structure 10.

Should we block training crawlers?

That is a policy decision and the product does not make it for you. What it does is show allow, disallow or not-set per bot, grouped by whether the bot retrieves, fetches on demand or trains, so that whatever you decide is what your robots.txt actually says.

See what is actually on your site

Point PixyScan at a URL and read the first report in a few minutes. The free plan covers one site and 500 URLs a month - enough to find out whether any of this is true.

No card required · 500 URLs a month on the free plan