Navid MoazzezNavid Moazzez

Free AI Crawler Checker (+ Guide)

An AI crawler checker that shows which AI crawlers your robots.txt lets in, and the exact line that decides each one.

5,0(1 rating)
Navid Moazzezby Navid Moazzez·Updated 3. okt. 2026·5 min read·

Paste your site to see which AI crawlers your robots.txt lets in: for AI search, for pages people ask about, and for training.

1

Check your site

The site is read under this site's own name, never as an AI crawler. Its robots.txt is kept here for 30 minutes.

Rate this tool

This free AI crawler checker shows which AI crawlers can read your site, and which ones your robots.txt blocks.

It checks 24 crawlers from OpenAI, Anthropic, Google, Perplexity, Apple, Meta, Amazon and more. And it quotes the exact line in your robots.txt that decides each one.

Here's what AI crawlers are, the 3 kinds, and how to stay in AI search while keeping out of training.

key_takeaways.mdTL;DR

Key takeaways

This free AI crawler checker shows which AI crawlers can read your site and which ones your robots.txt blocks.
It checks 24 crawlers and quotes the exact line in your robots.txt that decides each one.
User fetch crawlers from OpenAI, Perplexity and Meta may ignore robots.txt rules when a person asks for a page.
A firewall, a CDN like Cloudflare or your host can block AI crawlers before they ever read your robots.txt.

What is an AI crawler?

An AI crawler is a bot that an AI company sends to read web pages. Some read pages to train AI models, some to answer searches, and some fetch 1 page because a person asked.

Most follow the rules in your robots.txt, the file at the root of your site that says which bots may read what.

How to check if your site blocks AI crawlers

Check your site's AI crawlers0/4

Paste a page's full link to check that page's path instead of the home page.

The 3 kinds of AI crawlers

Each AI company runs different crawlers for different jobs. Blocking one does nothing to the others:

AI search
What it does
Reads pages to show and link them in AI answers
Examples
OAI-SearchBot, Claude-SearchBot, PerplexityBot
User fetches
What it does
Opens 1 page because someone asked the AI about it
Examples
ChatGPT-User, Claude-User, Perplexity-User
Training
What it does
Reads pages to train future AI models
Examples
GPTBot, ClaudeBot, CCBot

So you can block GPTBot to keep out of OpenAI's training, and still show up in ChatGPT search through OAI-SearchBot.

Tokens that aren't crawlers

Google-Extended and Applebot-Extended don't crawl anything. They're names in robots.txt that tell Google and Apple whether their normal crawlers' pages may train AI.

Blocking Google-Extended keeps your pages out of Gemini's training and grounding, according to Google's crawler docs. It doesn't touch Google Search.

Crawlers that may skip robots.txt

A fetch a person asked for is treated differently. OpenAI says robots.txt rules may not apply to ChatGPT-User, and Perplexity says Perplexity-User generally ignores them.

Meta says the same of Meta-ExternalFetcher. The checker marks each of these, so you don't count on a rule they may not read.

How to allow AI search and block AI training

Pick "AI search and user fetches yes, training no" in the checker, and copy the lines it writes. They look like this:

robots.txt lines
User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

Put them in your robots.txt, at the root of your site. Then check your site again to see the change.

What the checker can't see

robots.txt isn't the only gate. A firewall, a CDN like Cloudflare or your host can block AI crawlers before they ever read your robots.txt.

The checker reads your site under its own name, never pretending to be GPTBot. So if your robots.txt allows a crawler but your firewall blocks it, check your host's bot settings too.

Your crawler access score

The score is out of 100, and it has two halves.

The first half is the search engines, Googlebot first. The second is AI search, with the fetches people ask an AI for, like ChatGPT opening a page you linked.

I test both on your home page and up to nine pages from your sitemap, so a rule that only blocks your blog still shows up.

Training crawlers never count. Keeping GPTBot or ClaudeBot out of training is a choice, and the score shouldn't punish it.

Here's what three sites scored on October 1, 2026:

vercel.com
Score
100
Search engines
100
AI search
100
navid.me
Score
100
Search engines
100
AI search
100
nytimes.com
Score
66
Search engines
100
AI search
31

The New York Times keeps most AI search crawlers out, so its AI half drops to 31 while Google still reads everything.

Every check in the report

CheckWhat it looks at
robots.txt fileWhether it loads, and what a missing or broken file means
File sizeGoogle reads only the first 500 KiB
Search enginesThe search engines' crawlers on your home page
AI searchOAI-SearchBot, Claude-SearchBot, PerplexityBot and four more
AI assistantsThe fetches people ask ChatGPT, Claude and others for
AI trainingGPTBot, ClaudeBot, Google-Extended and five more, as information
Site-wide blockA Disallow: / that keeps everyone out
SitemapEach sitemap loaded and read, not just named
Sampled pagesUp to 10 of your pages, for every crawler that counts
Crawl-delayLines Google ignores and others follow
CSS and JavaScriptRules that stop Google drawing your pages
Rule syntaxLines no crawler can read
Content signalsHow you let your content be used
Named AI rulesHow many AI crawlers you name, and how many fall back to the rules for everyone

Failed checks and warnings also go into a fix list at the top, with what to change for each.

Content signals in robots.txt

Content signals are a newer, optional line in robots.txt. Cloudflare published the policy on September 24, 2025.

The policy has one signal per use:

SignalWhat it covers
searchA search index, with links and short excerpts
ai-inputLive AI answers that read your page
ai-trainTraining or fine-tuning AI models

Vercel's robots.txt allows search and AI answers but says no to training. That's a clear way to say "cite me, don't train on me."

A signal doesn't block anyone, though. It states a preference, and the Disallow lines still decide who may crawl.

Limits Google sets on robots.txt

Google doesn't support crawl-delay, so the line does nothing for Googlebot. Bing and some other crawlers slow down when they see it.

Google also reads only the first 500 KiB of a robots.txt. Anything after that is ignored, rules included.

And a noindex line in robots.txt has done nothing since September 1, 2019, when Google stopped reading it. Use a noindex meta tag on the page instead.

When my checker can't load your sitemap within 12 seconds, it says so. A slow answer is a limit on my side, so it never counts as a failed check.

llms.txt and your sitemap

The checker also looks for an llms.txt, a short guide to your site for AI, and for your sitemap.

If you don't have an llms.txt, make one in a minute with my free llms.txt generator.

Ready for AI search?

Make the files AI reads with my free tools

AI Crawler Checker FAQs

Questions about AI crawlers and robots.txt? Here's what to know.

An AI crawler checker is a tool that reads your robots.txt and shows which AI crawlers may read your site. This one checks 24 crawlers, splits them into AI search, user fetches and training, and quotes the line that decides each one.

Your site blocks ChatGPT's search if your robots.txt disallows OAI-SearchBot, and OpenAI's training if it disallows GPTBot. Paste your site into the AI crawler checker to see both, with the exact line.

GPTBot is OpenAI's crawler for training its AI models. Blocking it keeps your pages out of OpenAI's training but doesn't remove you from ChatGPT search, which uses OAI-SearchBot.

ClaudeBot is Anthropic's crawler for training its Claude models. Anthropic also runs Claude-SearchBot for search and Claude-User for pages a person asks about, and all 3 honor robots.txt.

Yes, block the training crawlers (GPTBot, ClaudeBot, Google-Extended and the rest) and allow the search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot). The AI crawler checker writes those lines for you.

No, blocking AI crawlers doesn't hurt your Google rankings as long as you leave Googlebot allowed. Google says blocking Google-Extended doesn't affect Google Search.

No, not all AI crawlers follow robots.txt every time. OpenAI, Perplexity and Meta say their fetchers for a person's request may skip it, and the checker marks those.

The checker reads your site under its own name and never pretends to be an AI crawler, so it can't see a firewall rule that only blocks GPTBot or ClaudeBot. Check your host's or CDN's bot settings for those.

llms.txt is a short Markdown file at the root of a site that guides AI to its most useful pages. The checker tells you if you have one, and the llms.txt generator makes one.

Yes, the AI crawler checker is free, with no sign-up. To keep it free for everyone, each person can check 30 sites every 10 minutes.

The AI crawler checker keeps a site's robots.txt on my site for 30 minutes, so a second check is instant. The address you type stays in your browser.

100 means every search engine and AI search crawler I test may read the pages I sampled.

Training crawlers never count, so blocking them doesn't lower it.

They're optional lines that say how your content may be used, from search to AI training.

Cloudflare published the policy in September 2025, and they state a preference without blocking anyone.

Google doesn't support crawl-delay, so Googlebot ignores the line.

Bing and some other crawlers slow down when they see it.

Google reads the first 500 KiB and ignores the rest.

The checker shows your file's size and warns when it's over.

The AI crawler checker is for site owners, creators and marketers who want to be found in AI answers, or to keep their work out of AI training.

For SEO, the title and meta description checker, the schema markup generator and the llms.txt generator cover the rest of a page's setup. See all my free tools.

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers