Robots.txt Generator — Block AI Crawlers and Test Your Rules
/admin also blocks /administrator/.Test a URL against the file
Which rule wins for a given crawler and path? The tester applies Google's matching rules: the longest matching path decides, and Allow wins a tie. Paste your own robots.txt to test an existing file.
* group.What robots.txt Does and How Crawlers Read It
A robots.txt file is a plain-text file at the root of a site, such as https://example.com/robots.txt, that tells crawlers which URL paths they may fetch. Before a well-behaved crawler requests any page on a host, it fetches that file, finds the group of rules written for its user-agent token, and checks the path it wants against that group's Allow and Disallow lines. The format is defined in RFC 9309, the Robots Exclusion Protocol, and Google publishes its own detailed interpretation, which this robots.txt generator follows.
Three things about the file surprise people. First, it is per host and per protocol: https://example.com/robots.txt says nothing about https://blog.example.com/ or http://example.com/. Second, it controls crawling, not indexing. A URL that is disallowed can still appear in Google's results if other pages link to it; Google just shows it without a description. To keep a page out of the index, let it be crawled and put <meta name="robots" content="noindex"> on it instead. Third, it is advisory. Search engines and the major AI companies honour it, but there is no enforcement; a scraper that ignores it has to be stopped at the server or the CDN.
Google reads only the first 500 KiB of the file and recognises four fields: User-agent, Allow, Disallow and Sitemap. Field names are case-insensitive; path values are not, so /Admin/ and /admin/ are different rules. Anything after a # is a comment. Google caches the file, typically for up to a day, so an edit takes a while to change crawler behaviour.
Robots.txt Example Files
The presets at the top of the tool produce the four files most sites need. The simplest robots.txt allows everything and points at the sitemap; it is what this site uses:
User-agent: * Allow: / Sitemap: https://stackiox.com/sitemap.xml
An empty Disallow: line is the traditional way to write "allow everything", and it means the same thing: Google ignores a rule with no path, so the group ends up with no restrictions. The opposite file blocks every crawler from every path. Use it on staging and preview hosts, never on the production domain, because it removes the site from search:
User-agent: * Disallow: /
The WordPress preset reproduces the file WordPress serves by default: the admin area is blocked, but the one admin endpoint that themes and plugins call from public pages is allowed through, and the sitemap WordPress has generated since version 5.5 is listed:
User-agent: * Disallow: /wp-admin/ Allow: /wp-admin/admin-ajax.php Sitemap: https://example.com/wp-sitemap.xml
The worked example combines path rules, an exception, a wildcard, the AI-bot section and a sitemap. With the "Worked example" preset the generator outputs:
# robots.txt generated with https://stackiox.com/robots-txt-generator/ User-agent: * Disallow: /admin/ Disallow: /search Disallow: /*.pdf$ Allow: /admin/public/ # AI training crawlers User-agent: GPTBot User-agent: ClaudeBot User-agent: Google-Extended User-agent: CCBot User-agent: Applebot-Extended User-agent: meta-externalagent User-agent: Amazonbot User-agent: Bytespider Disallow: / Sitemap: https://example.com/sitemap.xml
Several User-agent lines in a row share the rules that follow them; RFC 9309 and Google both define it that way. If you have to support a very old parser that reads one agent per group, tick "Write a separate group for each bot" and the generator repeats the Disallow: / under every token.
Block AI Crawlers with robots.txt
Each AI company publishes the token its crawler looks for in robots.txt. Every token in the tool was checked against the vendor's own documentation on 6 October 2026; a wrong token is silently ignored, which is why copying a list from an old blog post is risky. The tool separates the bots into two groups because they do different jobs and blocking them has different consequences.
| Token | Vendor | What it does | Honours robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Crawls content that may be used to train OpenAI's foundation models. | Yes |
| ClaudeBot | Anthropic | Collects web content used to train Claude models. | Yes |
| Google-Extended | Control token only. Opts your pages out of Gemini training and grounding. Google states it has no effect on Search inclusion or ranking. | Yes | |
| CCBot | Common Crawl | Builds the open Common Crawl archive, the source of many public training datasets. | Yes |
| Applebot-Extended | Apple | Opt-out token for training Apple's foundation models. Pages that disallow it still appear in Siri and Spotlight search. | Yes |
| meta-externalagent | Meta | Crawls for training Meta's AI models and for direct indexing of content. | Yes |
| Amazonbot | Amazon | Improves Amazon products; content may be used to train Amazon AI models. Does not support Crawl-delay. | Yes |
| Bytespider | ByteDance | No vendor documentation. Widely reported as the crawler behind ByteDance's LLM training. | Reported, not documented |
| OAI-SearchBot | OpenAI | Indexes sites for ChatGPT search results. Not used for training. | Yes |
| ChatGPT-User | OpenAI | Fetches a page when a ChatGPT user asks about it. Not automatic crawling. | User-initiated |
| Claude-SearchBot | Anthropic | Crawls to improve Claude's search result quality. | Yes |
| Claude-User | Anthropic | Fetches a page when a Claude user asks about it. | Yes |
| PerplexityBot | Perplexity | Indexes sites for Perplexity search. Perplexity says it is not used to train models. | Yes |
| Perplexity-User | Perplexity | Fetches on a user's request. Perplexity says this agent generally ignores robots.txt. | No |
The "Block AI training bots" preset ticks the first group only. Training crawlers copy your content into datasets you get nothing back from, and blocking them costs no traffic. The second group is different: OAI-SearchBot, Claude-SearchBot and PerplexityBot decide whether your pages are cited in AI search answers, which send real visitors. Block them only if you have decided you do not want that traffic. The user-initiated fetchers are listed for completeness; when you tick one, the tool notes that the rule is best-effort.
Google-Extended deserves its own sentence because it is the token people misunderstand most. It does not have a user-agent string and never fetches anything. Googlebot crawls the page as usual, and the Google-Extended rule only tells Google whether that crawled copy may be used for Gemini. Disallowing it does not remove you from Google Search, and the tool says so in its notes when you tick it.
Robots.txt Directives Reference
| Directive | Meaning | Example |
|---|---|---|
| User-agent | Starts a group and names the crawler token it applies to. * matches any crawler that has no group of its own. | User-agent: GPTBot |
| Disallow | A path prefix the crawler must not fetch. An empty value permits everything. | Disallow: /admin/ |
| Allow | A path prefix the crawler may fetch, used to carve exceptions out of a Disallow. | Allow: /admin/public/ |
| Sitemap | Full URL of an XML sitemap. File-wide, not tied to a group; may repeat. | Sitemap: https://example.com/sitemap.xml |
| Crawl-delay | Seconds between requests. Not part of RFC 9309; Google ignores it. | Crawl-delay: 5 |
| * | Wildcard inside a path: matches any run of characters. | Disallow: /*?sort= |
| $ | Anchors the pattern to the end of the URL. | Disallow: /*.pdf$ |
| # | Comment to the end of the line. | # staging only |
Allow vs Disallow: Which Rule Wins
When more than one rule matches a URL, Google applies the rule with the longest path, and if an Allow and a Disallow of the same length both match, the Allow wins. The tester at the top of the page applies exactly that logic and names the deciding rule. Against the worked-example file:
/admin/public/report.pdffor*is allowed. Three rules match:Disallow: /admin/,Disallow: /*.pdf$andAllow: /admin/public/. The Allow is the longest, so it wins even though the file blocks PDFs./admin/report.pdffor*is blocked./admin/and/*.pdf$are both seven characters and both Disallow, so the first one listed is reported as the deciding rule./searchingfor*is blocked, becauseDisallow: /searchis a prefix, not a whole word. Write/search/or/search$if you mean the one page./blog/forGPTBotis blocked by theDisallow: /in the AI group, while the same path forGooglebotis allowed, because no group names Googlebot, the*group applies, and nothing in it matches/blog/./file.pdf?download=1for*is allowed. The$in/*.pdf$anchors the rule to the end of the URL, and the query string comes after.pdf. Drop the$to block PDFs with any query string.
A crawler only ever reads one group. If a file has a User-agent: * group and a User-agent: GPTBot group, GPTBot follows its own group and ignores the * rules entirely. That is why the generator writes Disallow: / under the AI tokens rather than relying on anything in the default group, and why giving a bot its own group with no rules accidentally allows it everything.
Crawl-delay: What It Actually Does
Crawl-delay asks a crawler to wait a number of seconds between requests. It is not in RFC 9309, and Google's documentation lists it among the fields Google does not support, so for Googlebot the line is noise. Anthropic documents that ClaudeBot honours it, and Amazon states that Amazonbot does not. Some other crawlers read it; most ignore it. If Googlebot is hitting your server too hard, the fix is Search Console's crawl-rate settings or returning HTTP 429 or 503 for a while, which Googlebot responds to by slowing down. The generator accepts a value because some crawlers use it, and flags it with a note so you do not expect it to do more than it does.
Common robots.txt Mistakes
- Blocking CSS and JavaScript. Google renders pages. If it cannot fetch your stylesheets and scripts, it sees a broken page and ranks it accordingly. Do not disallow
/assets/,/static/or/_next/. - Using robots.txt to hide a page. A disallowed URL can still be indexed from links. Use
noindexand let it be crawled. Blocking and addingnoindexis the worst of both: Google cannot see the noindex because it cannot fetch the page. - A path without a leading slash.
Disallow: adminis not a valid path. The tool flags it as an error, and the tester treats it as/adminso you can see what you probably meant. - Forgetting the prefix rule.
Disallow: /appalso blocks/apple-pie/and/application-form. End the path with/or$when you mean one thing. - A relative Sitemap line.
Sitemap: /sitemap.xmlis invalid; the value must be a full URL. The tool rejects anything that is nothttp://orhttps://. - Shipping the staging file. A
Disallow: /that was fine on a preview host and got deployed to production is the classic way a site vanishes from Google overnight. The tool warns whenever you choose the block-everything policy. - Expecting it to stop bad actors. Scrapers that do not identify themselves never read the file. Rate limiting and bot rules at your CDN are the tools for that.
- Giving a bot an empty group.
User-agent: Bytespiderfollowed by nothing allows Bytespider everything, because it no longer reads the*group.
How to Test robots.txt in Code
The tester on this page is handy, but when a deploy pipeline needs to check that a release has not shipped the wrong file, do it in code. Python's standard library includes a parser:
from urllib.robotparser import RobotFileParser
rp = RobotFileParser()
rp.set_url("https://example.com/robots.txt")
rp.read()
print(rp.can_fetch("GPTBot", "https://example.com/blog/")) # False with the worked example
print(rp.can_fetch("Googlebot", "https://example.com/blog/")) # True
In Node, the robots-parser package implements the same longest-match rules:
const robotsParser = require('robots-parser');
const txt = await (await fetch('https://example.com/robots.txt')).text();
const robots = robotsParser('https://example.com/robots.txt', txt);
console.log(robots.isAllowed('https://example.com/admin/public/report.pdf', '*')); // true
For a production-exact answer, Google open-sourced the C++ parser Googlebot uses as google/robotstxt; it ships a small command-line tool that takes a file, a user agent and a URL. And to see what you are actually serving, fetch the live file rather than reading the one in your repository, because a CDN rule, a framework route or a Cloudflare Pages _redirects entry can change what crawlers receive:
curl -s https://example.com/robots.txt
The Sitemap line is only as good as the file it points at. If you generate sitemaps by hand or need to inspect one a plugin produced, our XML to JSON converter turns the <urlset> into a structure you can read and diff. When you are checking a page's other crawler-facing tags at the same time, the Open Graph generator writes the <title>, meta description, canonical and social tags in one block; remember that link-preview crawlers must be allowed to fetch the page for a share card to appear. And a staging site full of placeholder copy from the lorem ipsum generator is exactly the kind of host the block-everything preset exists for.
Frequently Asked Questions
User-agent, Allow, Disallow, Sitemap and Crawl-delay lines from a form, adds AI crawler blocks using tokens verified against each vendor's documentation, checks the file for common errors, and tests any path to show which rule applies. It runs entirely in your browser.Disallow: /. For OpenAI that is User-agent: GPTBot, for Anthropic User-agent: ClaudeBot, for Google's Gemini training User-agent: Google-Extended, and for Common Crawl User-agent: CCBot. The "Block AI training bots" preset writes all eight training tokens at once; the AI search agents are listed separately because blocking them removes you from AI search results.https://example.com/robots.txt. Crawlers do not look in subdirectories, and each subdomain and protocol needs its own file. Download the generated file and deploy it alongside your site's other root files.noindex robots meta tag or X-Robots-Tag header instead.Disallow: /admin/ with Allow: /admin/public/ blocks the admin area except the public folder. The tester on this page reports the deciding rule for any path.User-agent lines share the rules that follow them under both RFC 9309 and Google's parser. The generator writes the AI bots that way by default; tick "Write a separate group for each bot" to repeat the rules under every token for parsers that only read one agent per group.