Navid MoazzezNavid Moazzez

Free Robots.txt Validator (+ Guide)

A free robots.txt validator that checks every line the way Google reads it, and up to 500 of your pages against the file.

5,0(1 rating)
Navid Moazzezby Navid Moazzez·Updated 3 oct. 2026·6 min read·

Read any site's robots.txt, or paste one, and I check every line the way Google reads it, then check up to 500 of your pages against it.

1

The robots.txt to check

Try:
Rate this tool

A robots.txt validator reads a robots.txt file line by line, and flags every line crawlers can't read, or read differently than you meant.

This free robots.txt validator checks any site's file, or one you paste, against Google's own rules: 28 checks, each with the line it's on and how to fix it.

Then it checks up to 500 addresses against the file, for one crawler or all 24 I track, starting with the pages in your own sitemap.

I built it because a robots.txt can look fine and still fail... one typo in a field name, and Google skips the rule without a word.

key_takeaways.mdTL;DR

Key takeaways

Google reads only four fields: User-agent, Allow, Disallow and Sitemap.
A crawler with its own group skips every rule for *, so copy in the rules it should still follow.
Paths are case-sensitive and start with /, and a Sitemap line needs a full address.
Google reads only the first 500 KiB of the file.
A page in your sitemap should never be kept out for Googlebot.

What is a robots.txt validator?

A robots.txt validator is a tool that checks a robots.txt file for mistakes: lines crawlers can't read, rules that don't do what they seem to, and fields Google ignores.

A tester tells you what a file allows, and a validator tells you if the file is right. This one does both.

How to validate your robots.txt

Validate a robots.txt0/5

What the robots.txt validator checks

Every check follows Google's robots.txt rules, updated August 31, 2026, and the standard, RFC 9309.

Errors
What it checks
Typos in field names, lines with no colon, rules before any User-agent, paths without /, full links as paths, a Sitemap without https://, a file over 500 KiB, Google kept out of the home page
What Google does
Skips the line, or stops reading
Warnings
What it checks
A crawler whose own group skips the * rules, blocked styles and scripts, noindex, a $ in the middle of a path, a file served as something other than plain text
What Google does
Reads the file, but not the way you meant
Tips
What it checks
Crawl-delay, Host, fields Google doesn't read, a * at the end of a path, capitals in paths, a crawler named twice, no Sitemap line
What Google does
Nothing breaks, but the file could be clearer

It also lists every check the file passed, so you know what's already right.

A real example: a file with six mistakes

Press "Try a file with mistakes" on the tool to load this short file. I wrote it with six of the mistakes I see most often.

robots.txt
User-agent: *
Disallow: /cart/
Disallow: /Search
Dissallow: /checkout/
Crawl-delay: 10
Noindex: /old-page/

User-agent: Googlebot
Allow: /blog/

Sitemap: /sitemap.xml

Here's what the validator finds, worst first:

4
Level
Error
What it says
Dissallow is a typo for Disallow, so /checkout/ stays open
11
Level
Error
What it says
The Sitemap line needs a full address with https://
6
Level
Warning
What it says
Google stopped reading noindex in robots.txt in 2019
8
Level
Warning
What it says
Googlebot has its own group, so it skips both rules for *
3
Level
Tip
What it says
Paths are case-sensitive, so /Search doesn't cover /search
5
Level
Tip
What it says
Google ignores Crawl-delay

Line 8 is the one that hurts most. The file looks like it keeps every crawler out of the cart, but Googlebot follows only its own group, so it reads /cart/ anyway.

The most common robots.txt mistakes

These are the mistakes the validator catches most, and each one is a quick fix:

  • A typo in a field name, like Dissallow, so Google skips the rule
  • A crawler's own group without the * rules, so it reads what other crawlers can't
  • A path without / at the start, which crawlers ignore
  • A full link as a path, like Disallow: https://example.com/private/
  • A Sitemap line with only a path, like Sitemap: /sitemap.xml
  • Noindex or Crawl-delay, which Google doesn't read
  • Blocking styles and scripts, so Google can't see pages the way people do
  • The file in a folder, like /blog/robots.txt, where crawlers never look

Fix the errors first. A warning can wait a day, but an error means a rule isn't working right now.

Check your sitemap's pages against robots.txt

When you read a site, the validator also reads its sitemap, and fills in up to 500 of its pages as the addresses to check.

A sitemap lists the pages you want in Google, and robots.txt lists the ones crawlers should skip. So a page in both sends Google two opposite signals.

The validator warns you when Googlebot is kept out of any page in your sitemap. You can also pick any of the 24 crawlers, or all of them at once, and download the results as a .csv.

Real files from four big sites

I read four sites' robots.txt files with the validator on October 2, 2026:

My site, navid.me
Lines
235
Errors
0
Warnings
0
Tips
0
nytimes.com
Lines
364
Errors
0
Warnings
6
Tips
31
github.com
Lines
521
Errors
0
Warnings
3
Tips
294
en.wikipedia.org
Lines
712
Errors
0
Warnings
0
Tips
7

None of them has an error. The New York Times gets six warnings, though, because Googlebot and five other crawlers have their own groups that skip some rules for *.

So Googlebot may read checkout pages that the other crawlers can't.

GitHub's 294 tips are almost all a * at the end of a path, which changes nothing. So a long list of tips doesn't mean a bad file.

How Google reads a robots.txt file

Google's rules for the file itself are short, and each one can cost you:

RuleWhat happens
LocationOnly the file at the root counts, like example.com/robots.txt, and only for that host, protocol and port
FormatUTF-8 plain text, and Google skips any line it can't read
SizeGoogle reads the first 500 KiB and ignores the rest
RedirectsGoogle follows at least five redirects, then treats the file as missing
A 4xx answerNo file, so every page is open to crawl
A 5xx answerGoogle stops crawling for 12 hours, then uses its last good copy for up to 30 days
CacheGoogle keeps a copy for up to 24 hours, so a change can take a day to show

So a robots.txt that errors can be worse than having none at all.

Validator, tester or generator?

All three read robots.txt the same way, so pick the one for your job:

ToolUse it when
This validatorYou want to know if the file is right, line by line
Robots.txt testerYou want one page checked against every crawler
Robots.txt generatorYou want a new file from a few choices

The tabs at the top of the tool switch between them, and keep the site you typed.

How I keep my robots.txt right

Keep robots.txt right0/4

Who is the robots.txt validator for?

It's for anyone who edits a robots.txt, or wonders why Google skips a page:

And if a page suddenly dropped out of Google, robots.txt is one of the first files I'd validate.

Robots.txt Validator FAQs

Questions about validating robots.txt? Here's what to know.

It's a tool that checks a robots.txt file for mistakes, line by line.

This one runs 28 checks based on Google's own rules, and says how to fix each one.

Type your site above and press Read robots.txt, or paste the file.

The errors show first, each with its line number.

A validator checks that the file is right, and a tester checks what it allows.

This validator does both, for up to 500 addresses at once.

Google reads only four: User-agent, Allow, Disallow and Sitemap.

It skips every other field, like Crawl-delay and Host.

Googlebot has its own group, so it skips the rules for *.

Copy the rules it should follow into its own group.

Paths are, so /Blog/ and /blog/ are different.

Field names and crawler names aren't, so Disallow and disallow work the same.

Yes, Google needs the full address, starting with https://.

The sitemap can be on another host.

It has to be at the root of the host, like example.com/robots.txt.

A file in a folder is never read, and each subdomain needs its own.

Yes, read your site and the validator fills in up to 500 pages from its sitemap.

Tick Only the kept out ones to see just the blocked pages.

Yes, choose Paste it and paste the file.

Nothing you paste leaves your browser.

Yes, it's free, with no sign-up.

I did. I'm Navid Moazzez, and I made the robots.txt validator as one of my free tools on navid.me.

You can read more about me.

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers