Navid MoazzezNavid Moazzez

Free Robots.txt Generator (+ Guide)

A free robots.txt generator that writes the file from a few choices, keeps AI search and training apart, and tests it live.

Navid Moazzezby Navid Moazzez·Updated 2.10.2026·7 min read·

Choose who may read your site, and I write the robots.txt, with AI search and AI training crawlers apart, then test it live.

1

Search engines

Google, Bing and every other crawler can read your site, except the paths below.

2

Paths to keep out

Add:

One path per line, starting with /. Start a line with + to allow it inside a folder you keep out, like +/wp-admin/admin-ajax.php.

3

AI crawlers

AI searchOAI-SearchBot, Claude-SearchBot, PerplexityBot, Meta-WebIndexer, Amzn-SearchBot, MistralAI-Index, DuckAssistBot
AI trainingGPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, Amazonbot, MistralAI-Training, CCBot
Assistants acting for a personChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher, Amzn-User, MistralAI-User

AI search crawlers put your pages in answers with a link, and training crawlers collect pages to train models. Each company runs them apart, so you can keep training out and stay in AI search.

4

Your sitemap

5

Your robots.txt

User-agent: *
Disallow:

Put the file at the root of your site, like example.com/robots.txt. Crawlers look for it only there.

6

Test it

There's no Sitemap line. Add your sitemap's link so crawlers find every page.

For /: 0 of 24 crawlers are kept out.

CrawlerResultThe line that decided
OAI-SearchBotOpenAI · AI searchAllowedthe group for every crawler (*)
Claude-SearchBotAnthropic · AI searchAllowedthe group for every crawler (*)
PerplexityBotPerplexity · AI searchAllowedthe group for every crawler (*)
Meta-WebIndexerMeta · AI searchAllowedthe group for every crawler (*)
Amzn-SearchBotAmazon · AI searchAllowedthe group for every crawler (*)
MistralAI-IndexMistral · AI searchAllowedthe group for every crawler (*)
DuckAssistBotDuckDuckGo · AI searchAllowedthe group for every crawler (*)
ChatGPT-UserOpenAI · Assistants acting for a personAllowedthe group for every crawler (*)
Claude-UserAnthropic · Assistants acting for a personAllowedthe group for every crawler (*)
Perplexity-UserPerplexity · Assistants acting for a personAllowedthe group for every crawler (*)
Meta-ExternalFetcherMeta · Assistants acting for a personAllowedthe group for every crawler (*)
Amzn-UserAmazon · Assistants acting for a personAllowedthe group for every crawler (*)
MistralAI-UserMistral · Assistants acting for a personAllowedthe group for every crawler (*)
GPTBotOpenAI · AI trainingAllowedthe group for every crawler (*)
ClaudeBotAnthropic · AI trainingAllowedthe group for every crawler (*)
Google-ExtendedGoogle · AI trainingAllowedthe group for every crawler (*)
Applebot-ExtendedApple · AI trainingAllowedthe group for every crawler (*)
Meta-ExternalAgentMeta · AI trainingAllowedthe group for every crawler (*)
AmazonbotAmazon · AI trainingAllowedthe group for every crawler (*)
MistralAI-TrainingMistral · AI trainingAllowedthe group for every crawler (*)
CCBotCommon Crawl · AI trainingAllowedthe group for every crawler (*)
GooglebotGoogle · Search engineAllowedthe group for every crawler (*)
BingbotMicrosoft · Search engineAllowedthe group for every crawler (*)
ApplebotApple · Search engineAllowedthe group for every crawler (*)
Rate this tool

A robots.txt file is a short text file that tells crawlers which parts of your site they may read... and in 2026, it's also how you choose what AI companies do with your pages.

This free robots.txt generator writes the file from a few choices: search engines in or out, the paths to keep out, and every AI crawler by what it does.

Then it tests the file right away against the 23 crawlers I track, from Googlebot to GPTBot, so you see what each one may read before you upload anything.

I built it because AI search and AI training are different crawlers, often run by the same company... and you should be able to say yes to one and no to the other in a click.

So here's how robots.txt works, and how I'd set up yours.

key_takeaways.mdTL;DR

Key takeaways

A robots.txt file sits at the root of your site and tells crawlers which paths they may read.
Google reads four fields: user-agent, allow, disallow and sitemap. It ignores crawl-delay.
The most specific rule wins, and on a tie, Google picks the least restrictive one.
AI companies run separate crawlers for search and for training, so you can keep training out and stay in AI answers.
Blocking a page in robots.txt doesn't remove it from Google. Use a noindex tag for that.

What is a robots.txt file?

A robots.txt file is a plain text file at yoursite.com/robots.txt that gives crawlers rules.

Each group of rules names a crawler, then lists the paths it may or may not read.

Google's own introduction to robots.txt says it's mainly there to stop crawlers from overloading your site... and that it's not a way to keep a page out of Google.

It's also a request, not a lock. Googlebot and the big AI companies follow it, but a scraper can ignore it completely.

What the robots.txt generator does

Search engines
What you choose
Let them in, or keep them all out
What it writes
A group for every crawler, User-agent: *
Paths to keep out
What you choose
Presets for WordPress, shops, site search and admin, or your own
What it writes
Disallow lines, with Allow for exceptions
AI crawlers
What you choose
AI search, AI training and assistants, each same as above or kept out
What it writes
A group per crawler you keep out
Your sitemap
What you choose
Your sitemap's full address
What it writes
A Sitemap line
The test
What you choose
Any address on your site
What it writes
Allowed or kept out for all 23 crawlers, with the line that decided

It also warns you about the mistakes that cost the most, like keeping Google out of every page or blocking your styles and scripts.

How to make a robots.txt file

Make your robots.txt0/6

Pro tip Test your most important page before you upload the file, like your home page. If any crawler you want shows "Kept out", fix it now, not after Google stops visiting.

The four lines Google reads

Google's robots.txt spec supports four fields, and ignores the rest:

User-agent
What it does
Names the crawler the group is for, or * for every crawler
Example
User-agent: GPTBot
Disallow
What it does
A path the crawler may not read
Example
Disallow: /admin/
Allow
What it does
A path inside a blocked folder that it may read
Example
Allow: /wp-admin/admin-ajax.php
Sitemap
What it does
Your sitemap's full address, outside any group
Example
Sitemap: https://example.com/sitemap.xml

Crawl-delay isn't on the list, so Google ignores it (Bing reads it). And Google stopped reading noindex in robots.txt back in 2019.

One small detail catches people out: paths are case-sensitive, so /Blog/ and /blog/ are different.

Crawler names aren't, so GPTBot and gptbot are the same crawler.

How Google picks the rule

This is the part that trips people up, so here's how it works, step by step.

First, a crawler looks for the group that names it. If there's one, it follows only that group and ignores the * group completely.

Then, inside its group, the most specific rule wins, which means the one with the longest path. And if an Allow and a Disallow are equally long, Google picks the least restrictive one, the Allow.

You can also use two wildcards: * matches anything, and $ marks the end of the address. So Disallow: /*.pdf$ keeps crawlers out of every PDF.

The generator writes the file with these rules in mind, and the test reads it the same way, following the standard, RFC 9309.

What to keep out

These are the paths I'd keep crawlers out of on most sites:

  • Admin and login pages, where nothing helps a searcher
  • Cart and checkout pages, which are different for every visitor
  • Site search results, which can create endless pages with little value
  • Thank-you pages after a sign-up or a purchase
  • Your API, if your site has one

The presets add the common ones in one click, like /wp-admin/ with an Allow for admin-ajax.php, which WordPress themes need.

AI crawlers: search and training are different

This is where robots.txt changed the most in the last few years.

Most AI companies now run separate crawlers for separate jobs, and each one has its own name in robots.txt:

AI search
What it does
Puts your pages in AI answers, usually with a link
Crawlers I track
OAI-SearchBot, Claude-SearchBot, PerplexityBot and 4 more
AI training
What it does
Collects pages to train future models
Crawlers I track
GPTBot, ClaudeBot, Google-Extended, CCBot and 4 more
Assistants
What it does
Opens a page when a person asks an assistant to
Crawlers I track
ChatGPT-User, Claude-User, Perplexity-User and 3 more
Search engines
What it does
Classic search
Crawlers I track
Googlebot, Applebot

So you can keep GPTBot out of training and still let OAI-SearchBot put your pages in ChatGPT's answers.

Google-Extended is a special case. Google says it controls whether your pages train Gemini and ground its answers, but it doesn't affect your place in Google Search at all, and it isn't a ranking signal.

The AI crawler checker shows what each of these crawlers may read on any site today.

A real example: my robots.txt and the New York Times

I tested an article address on two very different sites, against the same 23 crawlers.

My site, navid.me
Groups in the file
26
Kept out of an article
0 of 23
nytimes.com
Groups in the file
69 user-agent lines
Kept out of an article
15 of 23

My own robots.txt lets every crawler read my articles, AI training included. It keeps all of them out of the same few paths, like my API and a couple of thank-you pages.

The New York Times goes the other way. It keeps out 15 of the 23, including OpenAI's, Anthropic's and Perplexity's AI search crawlers... while Googlebot can still read everything.

Neither one is wrong. It's a choice, and the generator makes it a few clicks instead of hundreds of lines of text.

Where the file goes

The file only works at the root of your site, like example.com/robots.txt. Crawlers don't look anywhere else.

It also only covers the host it sits on. So blog.example.com needs its own robots.txt, and so does a different protocol or port.

And two limits from Google: it reads up to 500 KiB of the file and ignores anything after that, and it usually keeps a copy for up to 24 hours. So a change can take a day to show.

robots.txt doesn't remove a page from Google

This one's easy to miss.

Google says a page blocked by robots.txt can still show in search results, just without a description, if other sites link to it.

So to keep a page out of search, add a noindex robots meta tag and let Google read the page. If Google can't read the page, it never sees the noindex.

Mistakes to avoid

  • Disallow: / for every crawler on a live site, which keeps Google out of everything
  • Blocking the folders with your styles and scripts, so Google can't see pages the way people do
  • Using robots.txt to hide private pages, when anyone can read the file
  • Blocking a page you want removed from Google, so it never sees the noindex
  • Keeping out every AI crawler when you only meant to stop training
  • Forgetting a subdomain, which needs its own file

Who is the robots.txt generator for?

It's for anyone who runs a site and wants to decide who reads it:

And if you've never looked at your robots.txt, test it once with the robots.txt tester. It takes a few seconds.

After you make your robots.txt

ToolUse it when
Robots.txt testerYou want to check a live site's file, or any address against it
AI crawler checkerYou want to see what AI companies read on a site
Sitemap generatorYou need the sitemap your Sitemap line points to
SEO checkerYou want every other check on a page

Robots.txt Generator FAQs

Questions about robots.txt? Here's what to know.

It's a text file at the root of a site that tells crawlers which paths they may read.

Search engines and AI companies read it before they crawl.

Choose your rules above, then download the file.

Upload it to the root of your site, as robots.txt.

At the root of the site, like example.com/robots.txt.

Crawlers don't look for it anywhere else.

Yes, the file only covers the host it's on.

So blog.example.com needs its own.

Google reads user-agent, allow, disallow and sitemap.

It ignores crawl-delay and anything else.

The one with the longest path wins.

On a tie, Google picks the least restrictive rule.

Yes, the big AI companies say their crawlers follow it.

Each one has its own name, and this generator lists them by what they do.

Yes, most AI companies run separate crawlers for each.

Keep the training crawlers out and leave the search crawlers in.

No, Google says it doesn't affect Search or rankings.

It controls whether your pages train Gemini and ground its answers.

No, a blocked page can still show if other sites link to it.

Use a noindex robots meta tag, and let Google read the page.

Yes, a Sitemap line helps every crawler find all your pages.

Add the full address, like https://example.com/sitemap.xml.

Google usually keeps a copy of your robots.txt for up to 24 hours.

So a change can take a day to show.

Yes, it's free, with no sign-up.

The file is written in your browser.

I did. I'm Navid Moazzez, and I made the robots.txt generator as one of my free tools on navid.me.

You can read more about me.

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers