A free robots.txt tester that checks any page against every crawler and quotes the line that decided each result.
Read any site's robots.txt, or paste one, and see which crawlers may read a page, with the line that decided.
The robots.txt to test
A robots.txt tester shows whether a crawler may read a page, by the rules in a site's robots.txt file.
This free robots.txt tester reads any site's robots.txt, or one you paste, and checks any address against the 23 crawlers I track: search engines, AI search, AI training and assistants.
For each crawler, it says allowed or kept out, and quotes the exact line that decided.
I built it because one wrong line in robots.txt can keep Google out of every page on a site... and today, Googlebot is only one of 23 crawlers worth checking.
So here's how crawlers read the file, what a real one looks like, and the mistakes to look for.
Key takeaways
What is a robots.txt tester?
A robots.txt tester is a tool that takes a robots.txt file and an address, and tells you which crawlers may read that address.
This one does it for 23 crawlers at once, from 10 companies, including Google, OpenAI, Anthropic, Perplexity, Apple and Meta.
What the robots.txt tester shows
- What it shows
- The site's robots.txt as crawlers see it
- Why it matters
- Read it from the site, or paste one that isn't live yet
- What it shows
- Any path or full link on the site
- Why it matters
- Test the pages that matter most
- What it shows
- Allowed or kept out, for all 23
- Why it matters
- See search and AI crawlers side by side
- What it shows
- The exact rule, and whether it came from the crawler's own group
- Why it matters
- Find the line to fix in seconds
- What it shows
- Google kept out, blocked styles and scripts, no sitemap, crawl-delay, noindex
- Why it matters
- The costly mistakes first
How to test a robots.txt file
Pro tip Paste your new file before you upload it. Nothing you paste leaves your browser, so you can test changes on a live site without touching it.
How crawlers read robots.txt
Every crawler follows the same steps, set by the standard, RFC 9309, and by Google's spec:
Two wildcards help: * matches anything, and $ marks the end of an address.
And paths are case-sensitive, so a rule for /Blog/ doesn't cover /blog/.
A real example: three sites, one article each
I tested an article or repo address on three sites, against all 23 crawlers.
- Lines in its robots.txt
- 235
- Kept out
- 0 of 23
- Lines in its robots.txt
- 364
- Kept out
- 15 of 23
- Lines in its robots.txt
- 279
- Kept out
- 0 of 23
My own file lets every crawler read my articles. The tester shows 15 of them allowed through their own group, and the other eight through the group for every crawler.
The New York Times keeps 15 out, including the AI search crawlers from OpenAI, Anthropic and Perplexity. Googlebot and Applebot can still read the article.
And GitHub names several AI crawlers in its file, but all 23 may read a public repo page... its rules keep crawlers out of other parts of the site.
So the same article can be open to everyone, or closed to most AI tools, depending on a few lines of text.
Why a page is blocked when you allowed it
This one catches a lot of site owners, and it's almost always the same cause.
When a crawler has its own group, it ignores the * group completely. So if you added Allow: /blog/ to User-agent: *, but GPTBot has its own group with Disallow: /, GPTBot is still kept out.
The tester shows which group each result came from, "its own group" or "the group for every crawler", so you see it right away.
When the robots.txt doesn't load
What Google does depends on the answer your server gives:
| Answer | What Google does |
|---|---|
| 200 | Follows the file |
| 404 or another 4xx | Treats it as no file, so every page is open to crawl |
| 5xx, a 429 or no answer | Stops crawling the site for up to 12 hours, then uses the last good copy for 30 days |
So a robots.txt that errors can be worse than having none at all. The tester warns you when a site's file answers with a server error.
The mistakes to look for
- Disallow: / under User-agent: *, which keeps Google out of every page
- Blocking a folder of styles or scripts, so Google can't see the page as people do
- An Allow in the * group that a crawler never reads, because it has its own group
- A path that doesn't start with /, which crawlers ignore
- Crawl-delay, which Google ignores
- Noindex in robots.txt, which Google stopped reading in 2019
How I test before I change robots.txt
Who is the robots.txt tester for?
It's for anyone who changes a robots.txt, or wonders why a page isn't showing up:
- How to start a blog – bloggers checking a new site isn't blocking Google
- Robots.txt generator – anyone writing a new file from scratch
- AI crawler checker – anyone deciding what AI companies may read
- Redirect checker – anyone who moved a site and wants to check both ends
- SEO checker – site owners who want every check on a page
And if a page suddenly dropped out of Google, the robots.txt is one of the first things I'd test.
After you test your robots.txt
| Tool | Use it when |
|---|---|
| Robots.txt generator | You want a new file from a few choices |
| AI crawler checker | You want a site's AI crawlers at a glance |
| Sitemap generator | Your file has no Sitemap line yet |
| SEO checker | You want every other check on a page |
More on crawlers
Robots.txt Tester FAQs
Questions about testing robots.txt? Here's what to know.
It's a tool that checks whether crawlers may read an address, by a site's robots.txt.
This one checks 23 crawlers at once.
Type your site above and press Read robots.txt, then type an address to test.
Each crawler shows allowed or kept out, with the line that decided.
A crawler follows the group that names it, or the group for every crawler when none does.
Inside it, the longest matching path wins.
A crawler with its own group ignores the group for every crawler.
Check which group the result names.
A 404 means no file, so every page is open.
A server error stops Google from crawling the site for up to 12 hours.
Yes, /Blog/ and /blog/ are different paths.
Crawler names aren't, so GPTBot and gptbot are the same.
No, Google ignores Crawl-delay.
Bing and some other crawlers read it.
23 crawlers from 10 companies, including Googlebot, GPTBot, ClaudeBot and PerplexityBot.
They're grouped as search engines, AI search, AI training and assistants.
Yes, choose Paste it and paste the file.
Nothing you paste leaves your browser.
Google usually keeps a copy for up to 24 hours.
So a change can take a day to show.
Yes, it's free, with no sign-up.
I did. I'm Navid Moazzez, and I made the robots.txt tester as one of my free tools on navid.me.
You can read more about me.
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
More free tools
Related MCP servers & CLIs
Free AI newsletterThe most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.






