BLOG · AI CRAWLER ACCESS

How to Check If AI Can Crawl Your Website

Published · Updated · 6 min read

The most common AI visibility failure is not missing schema markup or thin content — it is a robots.txt file that quietly blocks every AI crawler on the internet. The fix takes five minutes. The check takes thirty seconds. Here is how to do both.

Step 1: Fetch Your robots.txt

Navigate to yourdomain.com/robots.txt in a browser. You will see either a plain-text file with directives, or a 404 page.

  • 404: You have no robots.txt. Crawlers treat a missing file as allow-all, and Aura scores it as all 22 tracked crawlers allowed. Publish one anyway — even a permissive one — so you are explicitly in control.
  • File found: Proceed to Step 2.

Step 2: Search for AI Crawler User Agents

Use Ctrl+F (or Cmd+F) to search the page for each of the following strings:

User Agent StringOperatorUsed For
ChatGPT-UserOpenAIFetches pages during a live ChatGPT answer
OAI-SearchBotOpenAIChatGPT search index (live answers)
GPTBotOpenAIModel training (not live answers)
Claude-UserAnthropicFetches pages during a live Claude answer
Claude-SearchBotAnthropicClaude search index (live answers)
ClaudeBotAnthropicModel training (not live answers)
PerplexityBotPerplexityPerplexity search index + citations
Perplexity-UserPerplexityFetches pages during a live Perplexity answer
Google-ExtendedGoogleGemini training (separate from Googlebot)

If none of these strings appear in your robots.txt, all AI crawlers fall through to your wildcard group. Go to Step 3. (OpenAI documents which of its crawlers does what at platform.openai.com/docs/bots.)

Step 3: Check the Wildcard Rule

Look for a User-agent: * block. This catches every crawler that has no group of its own anywhere in the file. The two most common configurations are:

Blocking configuration (most common failure):

User-agent: *
Disallow: /

If no AI crawler has a group of its own in the file, every AI crawler — ChatGPT-User, OAI-SearchBot, GPTBot, ClaudeBot, PerplexityBot, all of them — is blocked from your entire site.

Permissive configuration (correct for most sites):

User-agent: *
Allow: /
Disallow: /admin/
Disallow: /private/

If your wildcard rule is permissive like this, AI crawlers are already allowed unless they have a group of their own that blocks them.

Step 4: Add Explicit AI Crawler Rules

Whether you want to allow or selectively block AI crawlers, being explicit is better than relying on wildcard inheritance. Here is a correct allow configuration for the major AI crawlers:

# Live-answer crawlers — allow (these drive citations)
User-agent: ChatGPT-User
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

# Training crawlers — allow (optional; blocking these is a warning, not a fail)
User-agent: GPTBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: Google-Extended
Allow: /

# Aggressive scrapers — block
User-agent: Bytespider
Disallow: /

# Catch-all
User-agent: *
Allow: /

Note: order does not matter. Under RFC 9309 a crawler uses the most specific user-agent group that matches its name, wherever that group sits in the file, and within that group the longest matching path rule wins. A named group therefore overrides the wildcard whether it appears above or below it. Also note the Bytespider block above: Aura scores this file as a warning, not a fail — Bytespider is a training-only crawler, so blocking it costs about 0.4 of the 30 crawler-access points (29.6/30). aaivisibility.com itself allows Bytespider.

What If You Want to Block AI Training But Allow Citations?

Some publishers draw a distinction between training data use (allowing a company to train a model on your content) and real-time citation (letting an AI engine retrieve and cite your content when answering user questions). This is a legitimate policy position.

If you want Perplexity to cite you but do not want OpenAI to use your content for training, you can implement this explicitly:

# Allow real-time citation crawlers
User-agent: PerplexityBot
Allow: /

User-agent: Perplexity-User
Allow: /

# Block training crawlers (your policy decision)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: *
Allow: /

The key is that this is a deliberate decision, not the accidental result of an unchanged legacy robots.txt. Make your policy explicit. Because ChatGPT-User, OAI-SearchBot, Claude-User and Claude-SearchBot have no group of their own here, they inherit the permissive wildcard and can still cite you; Aura scores this file as a warning (training crawlers blocked, live-answer crawlers allowed).

Frequently Asked Questions

What is GPTBot and why does it matter?
GPTBot is OpenAI’s training crawler — it collects content for model training, not for live answers. The crawlers that fetch your pages when ChatGPT answers a live question are ChatGPT-User and OAI-SearchBot. Blocking GPTBot alone costs only a warning; blocking ChatGPT-User or OAI-SearchBot means ChatGPT cannot retrieve your content when answering questions, and that is the failure that matters.

What is the difference between GPTBot and OAI-SearchBot?
GPTBot is OpenAI’s training crawler. OAI-SearchBot is OpenAI’s search crawler and ChatGPT-User is the agent that fetches pages during a live ChatGPT conversation. Allow OAI-SearchBot and ChatGPT-User if you want to be cited; allowing GPTBot is a separate, optional choice.

Does a wildcard Disallow block AI crawlers?
Yes. A User-agent: * block with Disallow: / blocks every crawler that has no group of its own — including every AI crawler. Add an explicit allow group for each AI crawler to override this; its position in the file does not matter.

Should I block AI training crawlers but allow search-augmented crawlers?
That is a legitimate business decision. Implement it explicitly — do not rely on a blanket block that catches both training and citation crawlers unintentionally.

Run the full 18-check AI visibility scan — free

How to Check If AI Can Crawl Your Website | Aura AI Visibility