A free robots.txt generator that writes the file from a few choices, keeps AI search and training apart, and tests it live.
Choose who may read your site, and I write the robots.txt, with AI search and AI training crawlers apart, then test it live.
Search engines
Google, Bing and every other crawler can read your site, except the paths below.
Paths to keep out
One path per line, starting with /. Start a line with + to allow it inside a folder you keep out, like +/wp-admin/admin-ajax.php.
AI crawlers
AI search crawlers put your pages in answers with a link, and training crawlers collect pages to train models. Each company runs them apart, so you can keep training out and stay in AI search.
Your sitemap
Your robots.txt
User-agent: * Disallow:
Put the file at the root of your site, like example.com/robots.txt. Crawlers look for it only there.
Test it
For /: 0 of 24 crawlers are kept out.
| Crawler | Result | The line that decided |
|---|---|---|
| OAI-SearchBotOpenAI · AI search | Allowed | the group for every crawler (*) |
| Claude-SearchBotAnthropic · AI search | Allowed | the group for every crawler (*) |
| PerplexityBotPerplexity · AI search | Allowed | the group for every crawler (*) |
| Meta-WebIndexerMeta · AI search | Allowed | the group for every crawler (*) |
| Amzn-SearchBotAmazon · AI search | Allowed | the group for every crawler (*) |
| MistralAI-IndexMistral · AI search | Allowed | the group for every crawler (*) |
| DuckAssistBotDuckDuckGo · AI search | Allowed | the group for every crawler (*) |
| ChatGPT-UserOpenAI · Assistants acting for a person | Allowed | the group for every crawler (*) |
| Claude-UserAnthropic · Assistants acting for a person | Allowed | the group for every crawler (*) |
| Perplexity-UserPerplexity · Assistants acting for a person | Allowed | the group for every crawler (*) |
| Meta-ExternalFetcherMeta · Assistants acting for a person | Allowed | the group for every crawler (*) |
| Amzn-UserAmazon · Assistants acting for a person | Allowed | the group for every crawler (*) |
| MistralAI-UserMistral · Assistants acting for a person | Allowed | the group for every crawler (*) |
| GPTBotOpenAI · AI training | Allowed | the group for every crawler (*) |
| ClaudeBotAnthropic · AI training | Allowed | the group for every crawler (*) |
| Google-ExtendedGoogle · AI training | Allowed | the group for every crawler (*) |
| Applebot-ExtendedApple · AI training | Allowed | the group for every crawler (*) |
| Meta-ExternalAgentMeta · AI training | Allowed | the group for every crawler (*) |
| AmazonbotAmazon · AI training | Allowed | the group for every crawler (*) |
| MistralAI-TrainingMistral · AI training | Allowed | the group for every crawler (*) |
| CCBotCommon Crawl · AI training | Allowed | the group for every crawler (*) |
| GooglebotGoogle · Search engine | Allowed | the group for every crawler (*) |
| BingbotMicrosoft · Search engine | Allowed | the group for every crawler (*) |
| ApplebotApple · Search engine | Allowed | the group for every crawler (*) |
A robots.txt file is a short text file that tells crawlers which parts of your site they may read... and in 2026, it's also how you choose what AI companies do with your pages.
This free robots.txt generator writes the file from a few choices: search engines in or out, the paths to keep out, and every AI crawler by what it does.
Then it tests the file right away against the 24 crawlers I track, from Googlebot to GPTBot, so you see what each one may read before you upload anything.
I built it because AI search and AI training are different crawlers, often run by the same company... and you should be able to say yes to one and no to the other in a click.
So here's how robots.txt works, and how I'd set up yours.
Key takeaways
What is a robots.txt file?
A robots.txt file is a plain text file at yoursite.com/robots.txt that gives crawlers rules.
Each group of rules names a crawler, then lists the paths it may or may not read.
Google's own introduction to robots.txt says it's mainly there to stop crawlers from overloading your site... and that it's not a way to keep a page out of Google.
It's also a request, not a lock. Googlebot and the big AI companies follow it, but a scraper can ignore it completely.
What the robots.txt generator does
- What you choose
- Let them in, or keep them all out
- What it writes
- A group for every crawler, User-agent: *
- What you choose
- Presets for WordPress, shops, site search and admin, or your own
- What it writes
- Disallow lines, with Allow for exceptions
- What you choose
- AI search, AI training and assistants, each same as above or kept out
- What it writes
- A group per crawler you keep out
- What you choose
- Your sitemap's full address
- What it writes
- A Sitemap line
- What you choose
- Any address on your site
- What it writes
- Allowed or kept out for all 24 crawlers, with the line that decided
It also warns you about the mistakes that cost the most, like keeping Google out of every page or blocking your styles and scripts.
How to make a robots.txt file
Pro tip Test your most important page before you upload the file, like your home page. If any crawler you want shows "Kept out", fix it now, not after Google stops visiting.
The four lines Google reads
Google's robots.txt spec supports four fields, and ignores the rest:
- What it does
- Names the crawler the group is for, or * for every crawler
- Example
- User-agent: GPTBot
- What it does
- A path the crawler may not read
- Example
- Disallow: /admin/
- What it does
- A path inside a blocked folder that it may read
- Example
- Allow: /wp-admin/admin-ajax.php
- What it does
- Your sitemap's full address, outside any group
- Example
- Sitemap: https://example.com/sitemap.xml
Crawl-delay isn't on the list, so Google ignores it (Bing reads it). And Google stopped reading noindex in robots.txt back in 2019.
One small detail catches people out: paths are case-sensitive, so /Blog/ and /blog/ are different.
Crawler names aren't, so GPTBot and gptbot are the same crawler.
How Google picks the rule
This is the part that trips people up, so here's how it works, step by step.
First, a crawler looks for the group that names it. If there's one, it follows only that group and ignores the * group completely.
Then, inside its group, the most specific rule wins, which means the one with the longest path. And if an Allow and a Disallow are equally long, Google picks the least restrictive one, the Allow.
You can also use two wildcards: * matches anything, and $ marks the end of the address. So Disallow: /*.pdf$ keeps crawlers out of every PDF.
The generator writes the file with these rules in mind, and the test reads it the same way, following the standard, RFC 9309.
What to keep out
These are the paths I'd keep crawlers out of on most sites:
- Admin and login pages, where nothing helps a searcher
- Cart and checkout pages, which are different for every visitor
- Site search results, which can create endless pages with little value
- Thank-you pages after a sign-up or a purchase
- Your API, if your site has one
The presets add the common ones in one click, like /wp-admin/ with an Allow for admin-ajax.php, which WordPress themes need.
AI crawlers: search and training are different
This is where robots.txt changed the most in the last few years.
Most AI companies now run separate crawlers for separate jobs, and each one has its own name in robots.txt:
- What it does
- Puts your pages in AI answers, usually with a link
- Crawlers I track
- OAI-SearchBot, Claude-SearchBot, PerplexityBot and 4 more
- What it does
- Collects pages to train future models
- Crawlers I track
- GPTBot, ClaudeBot, Google-Extended, CCBot and 4 more
- What it does
- Opens a page when a person asks an assistant to
- Crawlers I track
- ChatGPT-User, Claude-User, Perplexity-User and 3 more
- What it does
- Classic search
- Crawlers I track
- Googlebot, Applebot
So you can keep GPTBot out of training and still let OAI-SearchBot put your pages in ChatGPT's answers.
Google-Extended is a special case. Google says it controls whether your pages train Gemini and ground its answers, but it doesn't affect your place in Google Search at all, and it isn't a ranking signal.
The AI crawler checker shows what each of these crawlers may read on any site today.
A real example: my robots.txt and the New York Times
I tested an article address on two very different sites, against the same 24 crawlers.
- Groups in the file
- 26
- Kept out of an article
- 0 of 24
- Groups in the file
- 63
- Kept out of an article
- 15 of 24
My own robots.txt lets every crawler read my articles, AI training included. It keeps all of them out of the same few paths, like my API and a couple of thank-you pages.
The New York Times goes the other way. It keeps out 15 of the 24, including OpenAI's, Anthropic's and Perplexity's AI search crawlers... while Googlebot can still read everything.
Neither one is wrong. It's a choice, and the generator makes it a few clicks instead of hundreds of lines of text.
Where the file goes
The file only works at the root of your site, like example.com/robots.txt. Crawlers don't look anywhere else.
It also only covers the host it sits on. So blog.example.com needs its own robots.txt, and so does a different protocol or port.
And two limits from Google: it reads up to 500 KiB of the file and ignores anything after that, and it usually keeps a copy for up to 24 hours. So a change can take a day to show.
robots.txt doesn't remove a page from Google
This one's easy to miss.
Google says a page blocked by robots.txt can still show in search results, just without a description, if other sites link to it.
So to keep a page out of search, add a noindex robots meta tag and let Google read the page. If Google can't read the page, it never sees the noindex.
Mistakes to avoid
- Disallow: / for every crawler on a live site, which keeps Google out of everything
- Blocking the folders with your styles and scripts, so Google can't see pages the way people do
- Using robots.txt to hide private pages, when anyone can read the file
- Blocking a page you want removed from Google, so it never sees the noindex
- Keeping out every AI crawler when you only meant to stop training
- Forgetting a subdomain, which needs its own file
Who is the robots.txt generator for?
It's for anyone who runs a site and wants to decide who reads it:
- How to start a blog – bloggers setting up a new site
- How to create an online course – course creators keeping checkout and lesson pages out of search
- AI crawler checker – anyone deciding what AI companies may read
- Sitemap generator – site owners who want crawlers to find every page
- llms.txt generator – site owners who want AI tools to understand their site
And if you've never looked at your robots.txt, test it once with the robots.txt tester. It takes a few seconds.
After you make your robots.txt
| Tool | Use it when |
|---|---|
| Robots.txt tester | You want to check a live site's file, or any address against it |
| AI crawler checker | You want to see what AI companies read on a site |
| Sitemap generator | You need the sitemap your Sitemap line points to |
| SEO checker | You want every other check on a page |
More on crawlers
Robots.txt Generator FAQs
Questions about robots.txt? Here's what to know.
It's a text file at the root of a site that tells crawlers which paths they may read.
Search engines and AI companies read it before they crawl.
Choose your rules above, then download the file.
Upload it to the root of your site, as robots.txt.
At the root of the site, like example.com/robots.txt.
Crawlers don't look for it anywhere else.
Yes, the file only covers the host it's on.
So blog.example.com needs its own.
Google reads user-agent, allow, disallow and sitemap.
It ignores crawl-delay and anything else.
The one with the longest path wins.
On a tie, Google picks the least restrictive rule.
Yes, the big AI companies say their crawlers follow it.
Each one has its own name, and this generator lists them by what they do.
Yes, most AI companies run separate crawlers for each.
Keep the training crawlers out and leave the search crawlers in.
No, Google says it doesn't affect Search or rankings.
It controls whether your pages train Gemini and ground its answers.
No, a blocked page can still show if other sites link to it.
Use a noindex robots meta tag, and let Google read the page.
Yes, a Sitemap line helps every crawler find all your pages.
Add the full address, like https://example.com/sitemap.xml.
Google usually keeps a copy of your robots.txt for up to 24 hours.
So a change can take a day to show.
Yes, it's free, with no sign-up.
The file is written in your browser.
I did. I'm Navid Moazzez, and I made the robots.txt generator as one of my free tools on navid.me.
You can read more about me.
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
More free tools
Related MCP servers & CLIs
Free AI newsletterThe most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.






