Navid MoazzezNavid Moazzez

Free Robots.txt Tester (+ Guide)

A free robots.txt tester that checks any page against every crawler and quotes the line that decided each result.

Navid Moazzezby Navid Moazzez·Updated Oct 2, 2026·5 min read·

Read any site's robots.txt, or paste one, and see which crawlers may read a page, with the line that decided.

1

The robots.txt to test

Try:
Rate this tool

A robots.txt tester shows whether a crawler may read a page, by the rules in a site's robots.txt file.

This free robots.txt tester reads any site's robots.txt, or one you paste, and checks any address against the 23 crawlers I track: search engines, AI search, AI training and assistants.

For each crawler, it says allowed or kept out, and quotes the exact line that decided.

I built it because one wrong line in robots.txt can keep Google out of every page on a site... and today, Googlebot is only one of 23 crawlers worth checking.

So here's how crawlers read the file, what a real one looks like, and the mistakes to look for.

key_takeaways.mdTL;DR

Key takeaways

A crawler follows the group that names it, and ignores the * group when one does.
Inside that group, the rule with the longest path wins, and Allow wins a tie.
A robots.txt that answers with a server error stops Google from crawling the site for up to 12 hours.
Always test your most important pages after you change the file.
The line that decided is the fastest way to find the rule you need to fix.

What is a robots.txt tester?

A robots.txt tester is a tool that takes a robots.txt file and an address, and tells you which crawlers may read that address.

This one does it for 23 crawlers at once, from 10 companies, including Google, OpenAI, Anthropic, Perplexity, Apple and Meta.

What the robots.txt tester shows

The file
What it shows
The site's robots.txt as crawlers see it
Why it matters
Read it from the site, or paste one that isn't live yet
Your address
What it shows
Any path or full link on the site
Why it matters
Test the pages that matter most
Every crawler
What it shows
Allowed or kept out, for all 23
Why it matters
See search and AI crawlers side by side
The line that decided
What it shows
The exact rule, and whether it came from the crawler's own group
Why it matters
Find the line to fix in seconds
Warnings
What it shows
Google kept out, blocked styles and scripts, no sitemap, crawl-delay, noindex
Why it matters
The costly mistakes first

How to test a robots.txt file

Test a robots.txt0/5

Pro tip Paste your new file before you upload it. Nothing you paste leaves your browser, so you can test changes on a live site without touching it.

How crawlers read robots.txt

Every crawler follows the same steps, set by the standard, RFC 9309, and by Google's spec:

How a crawler decides0/6

Two wildcards help: * matches anything, and $ marks the end of an address.

And paths are case-sensitive, so a rule for /Blog/ doesn't cover /blog/.

A real example: three sites, one article each

I tested an article or repo address on three sites, against all 23 crawlers.

My site, navid.me
Lines in its robots.txt
235
Kept out
0 of 23
nytimes.com
Lines in its robots.txt
364
Kept out
15 of 23
github.com
Lines in its robots.txt
279
Kept out
0 of 23

My own file lets every crawler read my articles. The tester shows 15 of them allowed through their own group, and the other eight through the group for every crawler.

The New York Times keeps 15 out, including the AI search crawlers from OpenAI, Anthropic and Perplexity. Googlebot and Applebot can still read the article.

And GitHub names several AI crawlers in its file, but all 23 may read a public repo page... its rules keep crawlers out of other parts of the site.

So the same article can be open to everyone, or closed to most AI tools, depending on a few lines of text.

Why a page is blocked when you allowed it

This one catches a lot of site owners, and it's almost always the same cause.

When a crawler has its own group, it ignores the * group completely. So if you added Allow: /blog/ to User-agent: *, but GPTBot has its own group with Disallow: /, GPTBot is still kept out.

The tester shows which group each result came from, "its own group" or "the group for every crawler", so you see it right away.

When the robots.txt doesn't load

What Google does depends on the answer your server gives:

AnswerWhat Google does
200Follows the file
404 or another 4xxTreats it as no file, so every page is open to crawl
5xx, a 429 or no answerStops crawling the site for up to 12 hours, then uses the last good copy for 30 days

So a robots.txt that errors can be worse than having none at all. The tester warns you when a site's file answers with a server error.

The mistakes to look for

  • Disallow: / under User-agent: *, which keeps Google out of every page
  • Blocking a folder of styles or scripts, so Google can't see the page as people do
  • An Allow in the * group that a crawler never reads, because it has its own group
  • A path that doesn't start with /, which crawlers ignore
  • Crawl-delay, which Google ignores
  • Noindex in robots.txt, which Google stopped reading in 2019

How I test before I change robots.txt

Change robots.txt safely0/5

Who is the robots.txt tester for?

It's for anyone who changes a robots.txt, or wonders why a page isn't showing up:

And if a page suddenly dropped out of Google, the robots.txt is one of the first things I'd test.

After you test your robots.txt

ToolUse it when
Robots.txt generatorYou want a new file from a few choices
AI crawler checkerYou want a site's AI crawlers at a glance
Sitemap generatorYour file has no Sitemap line yet
SEO checkerYou want every other check on a page

Robots.txt Tester FAQs

Questions about testing robots.txt? Here's what to know.

It's a tool that checks whether crawlers may read an address, by a site's robots.txt.

This one checks 23 crawlers at once.

Type your site above and press Read robots.txt, then type an address to test.

Each crawler shows allowed or kept out, with the line that decided.

A crawler follows the group that names it, or the group for every crawler when none does.

Inside it, the longest matching path wins.

A crawler with its own group ignores the group for every crawler.

Check which group the result names.

A 404 means no file, so every page is open.

A server error stops Google from crawling the site for up to 12 hours.

Yes, /Blog/ and /blog/ are different paths.

Crawler names aren't, so GPTBot and gptbot are the same.

No, Google ignores Crawl-delay.

Bing and some other crawlers read it.

23 crawlers from 10 companies, including Googlebot, GPTBot, ClaudeBot and PerplexityBot.

They're grouped as search engines, AI search, AI training and assistants.

Yes, choose Paste it and paste the file.

Nothing you paste leaves your browser.

Google usually keeps a copy for up to 24 hours.

So a change can take a day to show.

Yes, it's free, with no sign-up.

I did. I'm Navid Moazzez, and I made the robots.txt tester as one of my free tools on navid.me.

You can read more about me.

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers