Technical SEO and AI search
robots.txt generator with AI crawler controls
Choose which AI crawlers may read your site, add your own rules and sitemap, then copy or download the file. Paste an existing robots.txt to test any path.
Access rules
These apply to every crawler that does not have its own group.
One per line. Paths start with /. Wildcards and $ are allowed.
Exceptions inside a disallowed area. The longer rule wins.
Full https:// addresses, one per line.
Honoured by Bing, ignored by Google.
AI crawlers
Search and answer crawlers surface your pages in AI results. Training crawlers collect content for model training. Each choice adds its own group only when it differs from the default rule.
-
OAI-SearchBot OpenAI Search index
Indexes pages for ChatGPT search results. Not used for model training.
-
ChatGPT-User OpenAI Answers on demand
Visits a page when a ChatGPT user asks a question that needs it.
-
GPTBot OpenAI AI training
Collects content that may be used to train OpenAI models.
-
Claude-SearchBot Anthropic Search index
Improves search result quality for Claude users. Not used for training.
-
Claude-User Anthropic Answers on demand
Fetches a page when a Claude user asks about it.
-
ClaudeBot Anthropic AI training
Collects web content that may contribute to model training.
-
PerplexityBot Perplexity Search index
Surfaces and links your pages in Perplexity search results. Not used to crawl for foundation models.
-
Perplexity-User Perplexity Answers on demand
Visits a page when a Perplexity user asks a question that needs it.
-
Google-Extended Google AI training
Controls whether crawled content may train Gemini or ground its answers. Google Search is not affected.
-
CCBot Common Crawl AI training
Crawls for the open Common Crawl archive, which many language models use for training.
-
Applebot-Extended Apple AI training
Controls whether Applebot content may train Apple foundation models. Applebot search is not affected.
-
Bytespider ByteDance AI training
ByteDance crawler associated with collecting data for AI model training.
-
meta-externalagent Meta AI training
Indexes the web for uses that include training foundation AI models.
Validator: test an existing robots.txt
Paste a file, choose a product token and a path. Longest match wins, and Allow wins ties (RFC 9309).
For example Googlebot, GPTBot or *.
Query strings count. A full URL uses its path and query.
Runs entirely in your browser. Nothing is uploaded.
Method
How it works
The generator writes one default group for User-agent: *. Rules inside a group apply to every crawler
that has no group of its own.
An AI crawler gets its own group only when its answer differs from the default. For example, with allow-all as the default, a
blocked training bot gets a Disallow: / group. With disallow-all as the default, an allowed
search bot gets an Allow: / group, which repeats your custom rules.
The validator reads groups as RFC 9309 describes. Groups that name the same product are merged. A crawler with no exact match uses the wildcard group. Among matching rules, the longest path wins, an Allow wins a tie, and an empty value matches nothing.
FAQ
Questions people ask
Does robots.txt stop AI crawlers?
Only the ones that choose to follow it. robots.txt is a voluntary standard, not access control. The large AI companies publish their crawler names and say they honour robots.txt. Other scrapers may ignore it, so treat the file as a request, and use server-side blocking if you need enforcement.
What is the difference between training and search crawlers?
Training crawlers collect content that may be used to train AI models. Search and answer crawlers index pages so an assistant can link to them, or fetch a page when a user asks about it. Many sites allow the second group and block the first, which is what the quick preset does.
Does blocking Google-Extended affect Google Search?
No. Google-Extended is a control token for Gemini training and grounding. Google Search still crawls and ranks your pages through Googlebot. Blocking Googlebot itself would remove you from search results, so never disallow it by mistake.
Why does the output repeat my custom rules?
A crawler obeys only the group that names it. When an AI bot needs a different answer from the default group, it gets its own group, and that group repeats your custom Disallow and Allow lines. Without the repeat, your path rules would silently stop applying to that bot.
How does the validator decide between Allow and Disallow?
It follows RFC 9309. The longest matching rule wins, and an Allow wins an equal-length tie. The asterisk matches any run of characters and a trailing dollar sign anchors the end. Paths are compared after percent-encoding normalisation, and the group is chosen by product token, falling back to the wildcard group.
Is Crawl-delay supported?
Google ignores Crawl-delay. Bing and some other crawlers honour it. If you need a guaranteed request rate, enforce it on the server or at the CDN instead.
Related tools
- Technical SEO & AI Search llms.txt Generator Write an llms.txt file in the llmstxt.org format, with a live preview and validation.
- Tracking & Analytics GA4 UTM Builder Build campaign URLs with bulk CSV import, a GA4 default channel check, local QR codes and team presets.
- Calculators Marketing Calculators ROAS, break-even ROAS, CPA, CPC, CPM, CTR, conversion rate, LTV:CAC and CAC payback in one place.
Next step
Want this done properly across your whole stack?
Tracking, search, automation and reporting, engineered and operated by subimpact network. Start with a free audit.
Get Free Audit