Technical SEO and AI search

robots.txt generator with AI crawler controls

Choose which AI crawlers may read your site, add your own rules and sitemap, then copy or download the file. Paste an existing robots.txt to test any path.

Access rules

These apply to every crawler that does not have its own group.

Default rule

Disallow all stops every compliant crawler, including Google Search.

One per line. Paths start with /. Wildcards and $ are allowed.

Exceptions inside a disallowed area. The longer rule wins.

Full https:// addresses, one per line.

Honoured by Bing, ignored by Google.

AI crawlers

Search and answer crawlers surface your pages in AI results. Training crawlers collect content for model training. Each choice adds its own group only when it differs from the default rule.

  • OAI-SearchBot OpenAI Search index

    Indexes pages for ChatGPT search results. Not used for model training.

  • ChatGPT-User OpenAI Answers on demand

    Visits a page when a ChatGPT user asks a question that needs it.

  • GPTBot OpenAI AI training

    Collects content that may be used to train OpenAI models.

  • Claude-SearchBot Anthropic Search index

    Improves search result quality for Claude users. Not used for training.

  • Claude-User Anthropic Answers on demand

    Fetches a page when a Claude user asks about it.

  • ClaudeBot Anthropic AI training

    Collects web content that may contribute to model training.

  • PerplexityBot Perplexity Search index

    Surfaces and links your pages in Perplexity search results. Not used to crawl for foundation models.

  • Perplexity-User Perplexity Answers on demand

    Visits a page when a Perplexity user asks a question that needs it.

  • Google-Extended Google AI training

    Controls whether crawled content may train Gemini or ground its answers. Google Search is not affected.

  • CCBot Common Crawl AI training

    Crawls for the open Common Crawl archive, which many language models use for training.

  • Applebot-Extended Apple AI training

    Controls whether Applebot content may train Apple foundation models. Applebot search is not affected.

  • Bytespider ByteDance AI training

    ByteDance crawler associated with collecting data for AI model training.

  • meta-externalagent Meta AI training

    Indexes the web for uses that include training foundation AI models.

Validator: test an existing robots.txt

Paste a file, choose a product token and a path. Longest match wins, and Allow wins ties (RFC 9309).

For example Googlebot, GPTBot or *.

Query strings count. A full URL uses its path and query.

Paste a robots.txt file to test it.

    Runs entirely in your browser. Nothing is uploaded.

    Method

    How it works

    The generator writes one default group for User-agent: *. Rules inside a group apply to every crawler that has no group of its own.

    An AI crawler gets its own group only when its answer differs from the default. For example, with allow-all as the default, a blocked training bot gets a Disallow: / group. With disallow-all as the default, an allowed search bot gets an Allow: / group, which repeats your custom rules.

    The validator reads groups as RFC 9309 describes. Groups that name the same product are merged. A crawler with no exact match uses the wildcard group. Among matching rules, the longest path wins, an Allow wins a tie, and an empty value matches nothing.

    FAQ

    Questions people ask

    Does robots.txt stop AI crawlers?

    Only the ones that choose to follow it. robots.txt is a voluntary standard, not access control. The large AI companies publish their crawler names and say they honour robots.txt. Other scrapers may ignore it, so treat the file as a request, and use server-side blocking if you need enforcement.

    What is the difference between training and search crawlers?

    Training crawlers collect content that may be used to train AI models. Search and answer crawlers index pages so an assistant can link to them, or fetch a page when a user asks about it. Many sites allow the second group and block the first, which is what the quick preset does.

    Does blocking Google-Extended affect Google Search?

    No. Google-Extended is a control token for Gemini training and grounding. Google Search still crawls and ranks your pages through Googlebot. Blocking Googlebot itself would remove you from search results, so never disallow it by mistake.

    Why does the output repeat my custom rules?

    A crawler obeys only the group that names it. When an AI bot needs a different answer from the default group, it gets its own group, and that group repeats your custom Disallow and Allow lines. Without the repeat, your path rules would silently stop applying to that bot.

    How does the validator decide between Allow and Disallow?

    It follows RFC 9309. The longest matching rule wins, and an Allow wins an equal-length tie. The asterisk matches any run of characters and a trailing dollar sign anchors the end. Paths are compared after percent-encoding normalisation, and the group is chosen by product token, falling back to the wildcard group.

    Is Crawl-delay supported?

    Google ignores Crawl-delay. Bing and some other crawlers honour it. If you need a guaranteed request rate, enforce it on the server or at the CDN instead.

    Next step

    Want this done properly across your whole stack?

    Tracking, search, automation and reporting, engineered and operated by subimpact network. Start with a free audit.

    Get Free Audit