IOCSITE CHECK

Methodology

How IOC Site Check measures AI accessibility

This page explains what a Site Check result means, what we observe during a scan, and where the limits of that observation begin.

We test technical accessibility and machine readability.
We do not measure whether AI platforms index, cite, or recommend your website.

Plain-language explanation

What does this answer?

Before an AI system can do anything useful with your web page, it first has to be able to reach the page and extract its content. This tool checks whether that basic step works.

What does this not answer?

This tool does not tell you whether ChatGPT, Claude, Perplexity or any other AI system will show your website to a user. That depends on many factors inside those companies that are not publicly measurable.

The three layers

Accessible

Can a crawler reach the page at all. We look at what your site allows in robots.txt and what actually happens when we request the page.

Readable

Is the content there in a form a machine can extract. We check whether the main content is present in the page source.

Understandable

Are there clear signals about what your company is and what it offers. We look for structured information that helps machines understand your organisation and offerings.

About the score

The score summarises our observations. It is not a grade. A low score is not a failure and a high score is not a guarantee. It shows how easy the page is for an AI system to reach, read and understand at the time of the scan.

Want to run your own scan? Run a scan.

Technical details

Can a crawler reach the page?

We look at two things and compare them: what your site declares in robots.txt, and what we actually observe when we request the page. These can disagree. A site can allow a crawler in robots.txt and still block it at the firewall, CDN, or server.

Within declared access, we separate two groups that matter most:

  • Training data crawlers gather content used to train AI models.
  • Search and retrieval crawlers fetch pages to answer user questions in real time.

Is the content machine-readable?

Content can be present in the raw HTML, or it can appear only after JavaScript runs. Some AI crawlers execute JavaScript, some do not. If your content depends entirely on it, you are making a bet about which systems will be able to read your page.

We use a browser only when the raw HTML is inconclusive.

Does the page make sense to machines?

We look for structured data that identifies your organisation and what you offer. We also check for an llms.txt file. It is treated as experimental, not a standard. Its absence does not count against a site. Its presence is treated as a weak, experimental signal.

What this scan does not cover

  • One page per scan, not the whole site.
  • A single browser render at most, with a time limit.
  • We do not impersonate other crawlers. What a specific AI crawler receives from a CDN or firewall may differ from what we observed.
  • We follow our own robots.txt token.
  • Results are observations at one moment in time.

How we identify ourselves

Our crawler identifies itself with a unique user agent and publishes its IP addresses. You can see the full details on the bot page and block it if you choose.

Read more about our crawler at /bot.