Robots.txt Generator — Block AI Crawlers and Test Your Rules

Everything runs in your browser. Nothing you type is uploaded, logged, or stored.
A path matches everything that starts with it. /admin also blocks /administrator/.
AI training crawlers
AI search & assistant agents
robots.txt — save at the root of your site

        

    Test a URL against the file

    Which rule wins for a given crawler and path? The tester applies Google's matching rules: the longest matching path decides, and Allow wins a tie. Paste your own robots.txt to test an existing file.

    The token a crawler matches in robots.txt, not its full User-Agent header. Matching is case-insensitive; a crawler with no group of its own uses the * group.

    What robots.txt Does and How Crawlers Read It

    A robots.txt file is a plain-text file at the root of a site, such as https://example.com/robots.txt, that tells crawlers which URL paths they may fetch. Before a well-behaved crawler requests any page on a host, it fetches that file, finds the group of rules written for its user-agent token, and checks the path it wants against that group's Allow and Disallow lines. The format is defined in RFC 9309, the Robots Exclusion Protocol, and Google publishes its own detailed interpretation, which this robots.txt generator follows.

    Three things about the file surprise people. First, it is per host and per protocol: https://example.com/robots.txt says nothing about https://blog.example.com/ or http://example.com/. Second, it controls crawling, not indexing. A URL that is disallowed can still appear in Google's results if other pages link to it; Google just shows it without a description. To keep a page out of the index, let it be crawled and put <meta name="robots" content="noindex"> on it instead. Third, it is advisory. Search engines and the major AI companies honour it, but there is no enforcement; a scraper that ignores it has to be stopped at the server or the CDN.

    Google reads only the first 500 KiB of the file and recognises four fields: User-agent, Allow, Disallow and Sitemap. Field names are case-insensitive; path values are not, so /Admin/ and /admin/ are different rules. Anything after a # is a comment. Google caches the file, typically for up to a day, so an edit takes a while to change crawler behaviour.

    Robots.txt Example Files

    The presets at the top of the tool produce the four files most sites need. The simplest robots.txt allows everything and points at the sitemap; it is what this site uses:

    User-agent: *
    Allow: /
    Sitemap: https://stackiox.com/sitemap.xml

    An empty Disallow: line is the traditional way to write "allow everything", and it means the same thing: Google ignores a rule with no path, so the group ends up with no restrictions. The opposite file blocks every crawler from every path. Use it on staging and preview hosts, never on the production domain, because it removes the site from search:

    User-agent: *
    Disallow: /

    The WordPress preset reproduces the file WordPress serves by default: the admin area is blocked, but the one admin endpoint that themes and plugins call from public pages is allowed through, and the sitemap WordPress has generated since version 5.5 is listed:

    User-agent: *
    Disallow: /wp-admin/
    Allow: /wp-admin/admin-ajax.php
    
    Sitemap: https://example.com/wp-sitemap.xml

    The worked example combines path rules, an exception, a wildcard, the AI-bot section and a sitemap. With the "Worked example" preset the generator outputs:

    # robots.txt generated with https://stackiox.com/robots-txt-generator/
    User-agent: *
    Disallow: /admin/
    Disallow: /search
    Disallow: /*.pdf$
    Allow: /admin/public/
    
    # AI training crawlers
    User-agent: GPTBot
    User-agent: ClaudeBot
    User-agent: Google-Extended
    User-agent: CCBot
    User-agent: Applebot-Extended
    User-agent: meta-externalagent
    User-agent: Amazonbot
    User-agent: Bytespider
    Disallow: /
    
    Sitemap: https://example.com/sitemap.xml

    Several User-agent lines in a row share the rules that follow them; RFC 9309 and Google both define it that way. If you have to support a very old parser that reads one agent per group, tick "Write a separate group for each bot" and the generator repeats the Disallow: / under every token.

    Block AI Crawlers with robots.txt

    Each AI company publishes the token its crawler looks for in robots.txt. Every token in the tool was checked against the vendor's own documentation on 6 October 2026; a wrong token is silently ignored, which is why copying a list from an old blog post is risky. The tool separates the bots into two groups because they do different jobs and blocking them has different consequences.

    TokenVendorWhat it doesHonours robots.txt
    GPTBotOpenAICrawls content that may be used to train OpenAI's foundation models.Yes
    ClaudeBotAnthropicCollects web content used to train Claude models.Yes
    Google-ExtendedGoogleControl token only. Opts your pages out of Gemini training and grounding. Google states it has no effect on Search inclusion or ranking.Yes
    CCBotCommon CrawlBuilds the open Common Crawl archive, the source of many public training datasets.Yes
    Applebot-ExtendedAppleOpt-out token for training Apple's foundation models. Pages that disallow it still appear in Siri and Spotlight search.Yes
    meta-externalagentMetaCrawls for training Meta's AI models and for direct indexing of content.Yes
    AmazonbotAmazonImproves Amazon products; content may be used to train Amazon AI models. Does not support Crawl-delay.Yes
    BytespiderByteDanceNo vendor documentation. Widely reported as the crawler behind ByteDance's LLM training.Reported, not documented
    OAI-SearchBotOpenAIIndexes sites for ChatGPT search results. Not used for training.Yes
    ChatGPT-UserOpenAIFetches a page when a ChatGPT user asks about it. Not automatic crawling.User-initiated
    Claude-SearchBotAnthropicCrawls to improve Claude's search result quality.Yes
    Claude-UserAnthropicFetches a page when a Claude user asks about it.Yes
    PerplexityBotPerplexityIndexes sites for Perplexity search. Perplexity says it is not used to train models.Yes
    Perplexity-UserPerplexityFetches on a user's request. Perplexity says this agent generally ignores robots.txt.No

    The "Block AI training bots" preset ticks the first group only. Training crawlers copy your content into datasets you get nothing back from, and blocking them costs no traffic. The second group is different: OAI-SearchBot, Claude-SearchBot and PerplexityBot decide whether your pages are cited in AI search answers, which send real visitors. Block them only if you have decided you do not want that traffic. The user-initiated fetchers are listed for completeness; when you tick one, the tool notes that the rule is best-effort.

    Google-Extended deserves its own sentence because it is the token people misunderstand most. It does not have a user-agent string and never fetches anything. Googlebot crawls the page as usual, and the Google-Extended rule only tells Google whether that crawled copy may be used for Gemini. Disallowing it does not remove you from Google Search, and the tool says so in its notes when you tick it.

    Robots.txt Directives Reference

    DirectiveMeaningExample
    User-agentStarts a group and names the crawler token it applies to. * matches any crawler that has no group of its own.User-agent: GPTBot
    DisallowA path prefix the crawler must not fetch. An empty value permits everything.Disallow: /admin/
    AllowA path prefix the crawler may fetch, used to carve exceptions out of a Disallow.Allow: /admin/public/
    SitemapFull URL of an XML sitemap. File-wide, not tied to a group; may repeat.Sitemap: https://example.com/sitemap.xml
    Crawl-delaySeconds between requests. Not part of RFC 9309; Google ignores it.Crawl-delay: 5
    *Wildcard inside a path: matches any run of characters.Disallow: /*?sort=
    $Anchors the pattern to the end of the URL.Disallow: /*.pdf$
    #Comment to the end of the line.# staging only

    Allow vs Disallow: Which Rule Wins

    When more than one rule matches a URL, Google applies the rule with the longest path, and if an Allow and a Disallow of the same length both match, the Allow wins. The tester at the top of the page applies exactly that logic and names the deciding rule. Against the worked-example file:

    A crawler only ever reads one group. If a file has a User-agent: * group and a User-agent: GPTBot group, GPTBot follows its own group and ignores the * rules entirely. That is why the generator writes Disallow: / under the AI tokens rather than relying on anything in the default group, and why giving a bot its own group with no rules accidentally allows it everything.

    Crawl-delay: What It Actually Does

    Crawl-delay asks a crawler to wait a number of seconds between requests. It is not in RFC 9309, and Google's documentation lists it among the fields Google does not support, so for Googlebot the line is noise. Anthropic documents that ClaudeBot honours it, and Amazon states that Amazonbot does not. Some other crawlers read it; most ignore it. If Googlebot is hitting your server too hard, the fix is Search Console's crawl-rate settings or returning HTTP 429 or 503 for a while, which Googlebot responds to by slowing down. The generator accepts a value because some crawlers use it, and flags it with a note so you do not expect it to do more than it does.

    Common robots.txt Mistakes

    How to Test robots.txt in Code

    The tester on this page is handy, but when a deploy pipeline needs to check that a release has not shipped the wrong file, do it in code. Python's standard library includes a parser:

    from urllib.robotparser import RobotFileParser
    
    rp = RobotFileParser()
    rp.set_url("https://example.com/robots.txt")
    rp.read()
    print(rp.can_fetch("GPTBot", "https://example.com/blog/"))   # False with the worked example
    print(rp.can_fetch("Googlebot", "https://example.com/blog/"))  # True

    In Node, the robots-parser package implements the same longest-match rules:

    const robotsParser = require('robots-parser');
    const txt = await (await fetch('https://example.com/robots.txt')).text();
    const robots = robotsParser('https://example.com/robots.txt', txt);
    console.log(robots.isAllowed('https://example.com/admin/public/report.pdf', '*'));  // true

    For a production-exact answer, Google open-sourced the C++ parser Googlebot uses as google/robotstxt; it ships a small command-line tool that takes a file, a user agent and a URL. And to see what you are actually serving, fetch the live file rather than reading the one in your repository, because a CDN rule, a framework route or a Cloudflare Pages _redirects entry can change what crawlers receive:

    curl -s https://example.com/robots.txt

    The Sitemap line is only as good as the file it points at. If you generate sitemaps by hand or need to inspect one a plugin produced, our XML to JSON converter turns the <urlset> into a structure you can read and diff. When you are checking a page's other crawler-facing tags at the same time, the Open Graph generator writes the <title>, meta description, canonical and social tags in one block; remember that link-preview crawlers must be allowed to fetch the page for a share card to appear. And a staging site full of placeholder copy from the lorem ipsum generator is exactly the kind of host the block-everything preset exists for.

    Frequently Asked Questions

    What is a robots.txt generator?
    A robots.txt generator builds the text file that tells search engines and other crawlers which paths on your site they may fetch. This one writes the User-agent, Allow, Disallow, Sitemap and Crawl-delay lines from a form, adds AI crawler blocks using tokens verified against each vendor's documentation, checks the file for common errors, and tests any path to show which rule applies. It runs entirely in your browser.
    How do I block AI crawlers like GPTBot and ClaudeBot?
    Add a group for each token with Disallow: /. For OpenAI that is User-agent: GPTBot, for Anthropic User-agent: ClaudeBot, for Google's Gemini training User-agent: Google-Extended, and for Common Crawl User-agent: CCBot. The "Block AI training bots" preset writes all eight training tokens at once; the AI search agents are listed separately because blocking them removes you from AI search results.
    Does blocking Google-Extended affect my Google rankings?
    No. Google states that Google-Extended does not affect a site's inclusion in Google Search and is not a ranking signal. It is a control token with no crawler of its own: Googlebot still fetches and indexes the page, and the rule only opts the content out of training and grounding Gemini models.
    Where do I put the robots.txt file?
    At the root of the host, so it is served at https://example.com/robots.txt. Crawlers do not look in subdirectories, and each subdomain and protocol needs its own file. Download the generated file and deploy it alongside your site's other root files.
    Does Google honour Crawl-delay?
    No. Google's documentation lists Crawl-delay among the fields it does not support, so the line has no effect on Googlebot. Some other crawlers read it: Anthropic documents that ClaudeBot honours it, while Amazon says Amazonbot does not. To slow Googlebot, use Search Console's crawl settings or respond with HTTP 429 or 503.
    Does Disallow in robots.txt remove a page from Google?
    Not reliably. Disallow stops Googlebot fetching the page, but if other pages link to the URL, Google can still index it without a description. To keep a page out of search results, allow crawling and add a noindex robots meta tag or X-Robots-Tag header instead.
    Which rule wins when Allow and Disallow both match?
    The rule with the longest path wins. If an Allow and a Disallow of equal length both match, Google uses the Allow. So Disallow: /admin/ with Allow: /admin/public/ blocks the admin area except the public folder. The tester on this page reports the deciding rule for any path.
    Can one User-agent group list several bots?
    Yes. Consecutive User-agent lines share the rules that follow them under both RFC 9309 and Google's parser. The generator writes the AI bots that way by default; tick "Write a separate group for each bot" to repeat the rules under every token for parsers that only read one agent per group.