IOCSITE CHECK

Common findings

Common findings and typical fixes

A Site Check report tells you what we observed. This page holds the general answers to the findings it raises most often — openly and for free. What applies to your specific site depends on its stack, its CDN and its templates; that part is specific work, not a general answer.

These changes improve technical accessibility for machines. None of them guarantees that any AI platform will index, cite or recommend a site — that is decided inside those platforms and is not measurable from outside. This page also stays inside what the scan actually measures: performance, security headers and accessibility are out of scope here.

Search and retrieval crawlers blocked in robots.txt

robots.txt rules for AI crawlers fall into two groups that are easy to conflate. Training crawlers (such as GPTBot, Google-Extended or ClaudeBot) gather content used to train AI models. Search and retrieval crawlers (such as OAI-SearchBot or PerplexityBot) fetch pages to answer user questions in real time. A broad rule written for one group also catches the other — worth checking which of the two the rule was meant for.

If the intention is to stay out of training data while remaining reachable for answer generation, the two groups can be addressed separately:

# Deliberately closed to training
User-agent: GPTBot
Disallow: /

User-agent: Google-Extended
Disallow: /

# Open to search / retrieval
User-agent: OAI-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

robots.txt is a declaration, not an enforcement mechanism — a CDN or firewall can still answer differently. Both sides appear separately in the report for that reason.

The crawler is challenged by a firewall or CDN

A report showing CHALLENGE_DETECTED or a 403 means the request was answered by a bot-management layer before any content was served — regardless of what robots.txt declares. Site owners are often unaware of this: the rules ship as defaults with CDN and firewall products.

Typical options, depending on the intention:

  • Most CDN panels have a setting that admits verified or known bots as a category.
  • Custom rules can allow specific crawlers by user agent, and more reliably by the published IP ranges of the crawler.
  • Keeping everything challenged is also a legitimate choice — it just means AI systems will not see the content, and that is worth deciding deliberately rather than by default.

To be clear about our own scan: when a report shows a challenge, that is what happened to us. Our crawler identifies itself and does not attempt to bypass challenges, so in that situation the report describes our access attempt — not how the site behaves for other crawlers.

Content depends on JavaScript

This finding means the initial HTML response contains little readable content, and the text appears only after scripts run in a browser. Some AI crawlers execute JavaScript; many fetch only the raw HTML. Content that exists only after rendering is invisible to the second group.

The typical options, independent of any particular framework:

  • Server-side rendering — the server sends the finished HTML, and scripts enhance it afterwards.
  • Static generation — pages are produced as HTML at build time.
  • A narrower version of both: making sure the critical text — what the organisation is, what it offers, how to reach it — is present in the initial HTML even if the rest stays dynamic.

The scan separates prediction from measurement here: static signals produce a first classification, and when they suggest high dependency, a real browser render measures the actual difference. A short page that is fully server-rendered is recognised as such by the measurement.

No Organization structured data

Organization structured data is a small machine-readable block that states who a website belongs to. Without it, a machine can read every page and still be unable to say what the company is called or where its official site is. The scan looks for an Organization entity and these fields: name, url, logo, description and sameAs (links to official profiles).

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Example Ltd",
  "url": "https://example.com",
  "description": "What the company actually does, in one sentence.",
  "logo": "https://example.com/logo.png",
  "sameAs": ["https://www.linkedin.com/company/example"]
}
</script>

Only real values belong here. If there is no logo file or no official profile, the honest option is to leave those fields out — not to invent a URL. Structured data that misdescribes an organisation is worse than none. Our own site publishes Organization data without a logo field for exactly this reason: the brand is typographic and no logo file exists yet.

Redirect chains and HTTPS issues

A single redirect to the canonical address is normal. Chains of several hops slow every crawler down, and a redirect loop stops them entirely. Where possible, the first redirect can point directly at the final address.

An expired or mismatched certificate ends most automated visits before any content is exchanged. The report shows what our connection observed about HTTPS and the certificate.

llms.txt

llms.txt is an experimental convention: a plain-text file at the site root that offers a machine-oriented summary of the site. It is not a web standard, and its effect on AI systems is unproven. The scan treats it accordingly — its absence is not a deficiency and does not count against a site; its presence is recorded as a weak, experimental signal. For the curious: the idea is to give language models a curated starting point instead of leaving them to infer one.

From general to specific

These are the general answers, and they are free. What applies to your site — its stack, its CDN rules, its templates, the journeys your customers actually take — is specific work. That is what IOC Audit does: it examines a site the way a customer uses it, not only the way a crawler reads it.