← All articles·technical·Pillar

The AI bot robots.txt complete guide for 2026

Complete reference for AI bot robots.txt configuration. Every crawler and control token that matters (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, etc.), the gold-standard explicit-allow pattern, Cloudflare and WAF traps, same-day verification procedure.

Data for AI Search Editorial Team··14 min read

The robots.txt file controls which crawlers, including AI bots, can read your site. As of mid-2026, at least thirteen AI bot names matter for citation: GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Google-Extended and Applebot-Extended (both robots.txt control tokens, not crawlers), Bytespider, cohere-ai, ImagesiftBot, and FacebookBot. Each crawler reads different parts of your site for different purposes (training data ingestion, real-time retrieval, image indexing), and blocking a platform's search crawler in robots.txt vetoes that platform in our 10-Point AI Citation Audit. As of September 2026, roughly a third of audited sites had a firewall that refused requests identifying as AI crawlers, and none blocked one in robots.txt. A brand can have flawless content, complete directory presence, and active brand mention engineering and still block OAI-SearchBot and miss ChatGPT search answers for months without knowing why. This guide is the reference document: every AI bot user agent that matters, the explicit-allow robots.txt pattern, the Cloudflare and WAF traps, and the same-day verification procedure.

What is the AI bot robots.txt?

robots.txt is a text file at the root of your domain (https://yourdomain.com/robots.txt) that tells crawlers which paths they can access. The format is simple: one or more rule groups, each specifying a User-agent directive followed by Allow and/or Disallow directives.

The traditional use case was managing search engine crawlers: telling Googlebot to skip admin pages or rate-limit Bingbot. AI bots arrived between 2023 and 2025 and adopted the same convention. OpenAI ships GPTBot. Anthropic ships ClaudeBot. Perplexity ships PerplexityBot. Google added Google-Extended, which as of September 2026 is a robots.txt control token for Gemini training and grounding and not a crawler (distinct from Googlebot for search). Each respects standard robots.txt directives when configured correctly.

The complication: many brands have robots.txt files configured for traditional search engines without explicit consideration of AI bots. The default User-agent: * block applies to every crawler including AI bots, which works fine. But infrastructure layers above robots.txt (Cloudflare's AI Crawl Control, WAF rules, Vercel firewalls) can block AI bots independently of what robots.txt says. As of September 2026, roughly a third of the sites in our stored 10-Point AI Citation Audit results had a firewall that refused a request identifying as an AI crawler, and none of them had a robots.txt block on an AI crawler. From outside, we cannot tell whether a firewall that refuses our test request also refuses the real crawler.

The robots.txt file is necessary but not sufficient for AI bot accessibility. Both layers (robots.txt AND infrastructure) must be configured correctly.

Which AI bots should you allow?

The complete current list of AI bots that materially affect AEO citation in 2026:

User-AgentOwnerPurposeAffects
GPTBotOpenAITraining data ingestionWhat OpenAI's models learn, not ChatGPT search answers
ChatGPT-UserOpenAIUser-action triggered retrievalChatGPT real-time
OAI-SearchBotOpenAIChatGPT Search retrievalWhether a site appears in ChatGPT search answers
ClaudeBotAnthropicTraining data ingestionWhat Claude models learn, not Claude search answers
Claude-SearchBotAnthropicSearch result quality for Claude usersClaude search
Claude-UserAnthropicUser-action triggered retrievalClaude real-time
PerplexityBotPerplexitySearch result retrieval, not model trainingPerplexity citation
Google-ExtendedGooglerobots.txt control token for Gemini training and groundingGemini, not AI Overviews
Applebot-ExtendedApplerobots.txt control token for Apple model trainingApple Intelligence
BytespiderByteDanceTikTok / Doubao trainingDoubao + TikTok AI
cohere-aiCohereCohere trainingCohere model citation
ImagesiftBotImagesiftImage AI trainingImage generation models
FacebookBotMetaMeta AI trainingMeta AI citation

Source: compiled by Data for AI Search, July 2026, with each bot detailed in the AI bot user-agent reference.

For most brands optimizing for the major AI assistants (ChatGPT, Perplexity, Claude, Gemini), allowing the first eight is essential. A robots.txt block on OAI-SearchBot, PerplexityBot, Google-Extended or Googlebot, or on both Claude-SearchBot and Claude-User, vetoes the corresponding platform per the 10-Point AI Citation Audit Check 1 as of September 2026, and a block on a training crawler (GPTBot, ClaudeBot) or on any other crawler in the table lowers the crawler accessibility score without a veto.

The remaining five (Applebot-Extended, Bytespider, cohere-ai, ImagesiftBot, FacebookBot) matter more selectively. Allowing Bytespider is useful for international brands targeting Chinese-speaking markets via Doubao. Applebot-Extended for brands optimizing for Apple Intelligence (still emerging as of mid-2026). The others depend on specific business contexts.

What's wrong with robots.txt by default?

Default robots.txt configurations have three problems: they rely on the wildcard rule, they can hide implicit blocks, and they leave out AI-specific declarations.

Reliance on User-agent: *. A blanket allow with User-agent: * / Allow: / does work: most AI bots respect it. But it's silent. There's no explicit signal that AI bots are welcome, and infrastructure layers above can independently block bots regardless of what robots.txt says. Brands that explicitly enumerate each AI bot in robots.txt send a clearer signal to both the bots and to internal/external auditors.

Implicit blocks via overly aggressive Disallow rules. A Disallow: /api/ or Disallow: /admin/ is fine. A Disallow: / under any named user-agent is a hard block on that bot. A brand can carry Disallow: / on user-agents it doesn't recognize, and if those are AI bots they stay blocked silently for months.

Missing AI-specific bot declarations entirely. Older robots.txt files written before 2024 may not mention any AI bots. They work by default because of the wildcard rule, but they leave the brand unaware of the AI bot crawler layer. When something breaks at the infrastructure level, there's no audit trail in robots.txt to investigate.

What is the Cloudflare GPTBot trap?

Cloudflare's AI Crawl Control panel (introduced 2024, evolved through 2025-2026) provides a UI for managing AI bot access at the CDN layer, before requests reach the origin server. The panel includes toggles for GPTBot, PerplexityBot, ClaudeBot, Bytespider, and others. Google-Extended is not in Cloudflare's crawler list as of September 2026 because it is a control token set in robots.txt.

The trap: a crawler toggled to "Block" in the panel is refused at the edge, whatever robots.txt says. A blocked GPTBot keeps the site's content out of OpenAI's model training, and a blocked OAI-SearchBot means ChatGPT search cannot show the site. The issue is invisible from the brand's perspective (robots.txt looks fine, server logs show no errors, content is being published) but Cloudflare is returning 403s to that crawler's requests at the edge. Cloudflare also has a managed robots.txt setting which, when it is switched on, writes rules into robots.txt that disallow AI training crawlers, as of September 2026. That setting leaves ChatGPT search alone, but its Google-Extended rule takes the site out of Gemini training and grounding.

The fix is same-day:

  1. Sign in to Cloudflare → select the domain
  2. Navigate to Security → Bots → AI Crawl Control
  3. Verify OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, and PerplexityBot are all toggled to Allow. Google-Extended is not a toggle here because it is set in robots.txt. Check the state of each toggle on your own account.
  4. Save changes
  5. Test using curl with the OAI-SearchBot user agent (see verification procedure below)

The Cloudflare panel takes precedence over robots.txt. A site with permissive robots.txt and a blocking Cloudflare toggle is effectively blocked. A firewall that refuses requests identifying as AI crawlers is a common finding in our audits as of September 2026, though from outside we cannot tell whether it also refuses the real crawler. Per the Two-Track Law, Track-1 content investment is wasted when the crawlers can't read the site.

What is the gold-standard explicit-allow pattern?

The gold-standard pattern names each AI bot in robots.txt and gives it an explicit allow, and it is the pattern we recommend in every audit we produce:

User-Agent: *
Allow: /

User-Agent: GPTBot
Allow: /

User-Agent: ChatGPT-User
Allow: /

User-Agent: OAI-SearchBot
Allow: /

User-Agent: ClaudeBot
Allow: /

User-Agent: Claude-SearchBot
Allow: /

User-Agent: Claude-User
Allow: /

User-Agent: PerplexityBot
Allow: /

User-Agent: Google-Extended
Allow: /

User-Agent: Applebot-Extended
Allow: /

User-Agent: CCBot
Allow: /

User-Agent: Bytespider
Allow: /

User-Agent: Meta-ExternalAgent
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

Why explicit allow rather than reliance on User-agent: *:

  • Audit trail. Anyone reading the file (internal team, external consultant, future auditor) sees an explicit signal that AI bots are welcome.
  • Defense against overly-restrictive blanket rules. If a security review adds Disallow: / under User-agent: *, the explicit AI bot allows remain in effect.
  • Industry convention. Brands publishing the explicit allow pattern (Stripe, Mintlify, Linear, Stack Overflow, and others) have established it as the standard for AI-friendly sites in 2026.

As of July 2026, we score the explicit-allow pattern at 10/10 on Check 1 of the 10-Point AI Citation Audit. The case study we cite most frequently is a San Diego painting contractor whose robots.txt explicitly allows every major AI bot by name. They scored 10/10 on crawler accessibility despite being a small local-services business with limited technical resources. The pattern is replicable.

How do you test crawler accessibility?

You test crawler accessibility by requesting a page with each bot's user agent and reading the response status. The procedure works for any domain:

# Test OAI-SearchBot
curl -A "OAI-SearchBot" -I https://yourdomain.com/

# Expected: HTTP/2 200 with response headers

# Test ChatGPT-User
curl -A "ChatGPT-User" -I https://yourdomain.com/

# Test GPTBot
curl -A "GPTBot" -I https://yourdomain.com/

# Test PerplexityBot
curl -A "PerplexityBot" -I https://yourdomain.com/

# Test Claude-SearchBot
curl -A "Claude-SearchBot" -I https://yourdomain.com/

# Test Claude-User
curl -A "Claude-User" -I https://yourdomain.com/

# Test ClaudeBot
curl -A "ClaudeBot" -I https://yourdomain.com/

# Check Google-Extended (a robots.txt control token, not a crawler)
curl -s https://yourdomain.com/robots.txt
# Look for a Google-Extended group with Disallow: /

Any response other than 200 (especially 403, 429, or 503) indicates a block somewhere in the request pipeline, which is what the crawler accessibility check in the 10-Point AI Citation Audit tests as of July 2026. Investigate by:

  1. Checking robots.txt for explicit Disallow rules under that user-agent.
  2. Checking Cloudflare AI Crawl Control panel for that bot's toggle state.
  3. Checking WAF rules for blanket AI-bot blocks.
  4. Checking Vercel firewall configuration (if deployed on Vercel).
  5. Checking any reverse proxy or load balancer rules between Cloudflare and origin.

Test from at least three representative pages: homepage, a content pillar, a service area page or product page. Some misconfigurations apply globally; others apply only to specific paths.

What WAF and firewall rules block AI bots?

Web Application Firewalls and infrastructure firewalls operate independently of robots.txt and Cloudflare's AI Crawl Control panel. Common patterns that inadvertently block AI bots:

Generic "bot challenge" rules. Many WAF configurations challenge or block any user-agent containing "bot", which catches every AI bot. Explicitly exempt named AI bot user agents from challenge rules.

Rate-limiting on legitimate AI crawler request volume. AI bots can request many pages in short windows during training crawls. Aggressive rate-limiting returns 429s and effectively blocks the crawl. Configure rate limits with AI bot user agents specifically exempted or with higher thresholds.

Geographic blocking. Some WAF configurations block requests from countries where major AI provider infrastructure is hosted. This can silently block ClaudeBot (Anthropic infrastructure) or GPTBot (OpenAI infrastructure). Google-Extended cannot be blocked at a firewall because it is a control token set in robots.txt.

TLS / cipher restrictions. AI bots use specific TLS configurations. Older WAF rules that block non-modern TLS may block AI bots while allowing browser traffic.

Audit the full request pipeline (origin, Vercel firewall, Cloudflare, WAF, CDN) quarterly. Each layer can independently block AI bots without the others noticing.

How often do AI bots crawl?

Crawl frequency varies by bot and by site authority:

  • GPTBot: crawls high-authority sites frequently (daily or near-daily); medium-authority sites weekly; low-authority sites monthly or less.
  • PerplexityBot: crawls aggressively during real-time retrieval. As of September 2026, Perplexity says it is not used to crawl content for AI foundation models. Pattern depends heavily on query frequency for the brand's domain.
  • ClaudeBot: training-data crawls happen in batches tied to Anthropic's model update cycles. Less consistent than GPTBot.
  • Google-Extended: has no crawl schedule of its own because it is a robots.txt control token. Google crawls with its existing crawlers.
  • Bytespider: aggressive crawl frequency for sites in its target verticals.

The implication: changes to robots.txt or infrastructure permissions take effect on the timescale of the next crawl cycle, not instantly. A same-day Cloudflare toggle change may take 1-4 weeks to fully appear in ChatGPT search answers as OAI-SearchBot re-crawls.

Frequently asked questions

Should I block any AI bots?

For most brands, no. Blocking a search crawler such as OAI-SearchBot or PerplexityBot takes the site out of that platform's search answers, and blocking a training crawler such as GPTBot or ClaudeBot makes the brand less likely to be known to the model itself, with no compensating benefit. The arguments for blocking (training data privacy, intellectual property concerns) apply primarily to brands publishing sensitive proprietary content or to brands philosophically opposed to AI training. For commercial brands optimizing for AEO/GEO, the right default is allow.

Does blocking Google-Extended affect Google Search rankings?

No. Google-Extended is a robots.txt control token for Gemini training and grounding; Googlebot is for Search. Blocking Google-Extended does not affect Google Search ranking. It does affect Gemini, in model training and in grounding answers in the Gemini app. AI Overviews and AI Mode are features of Google Search and follow the rules for Googlebot.

What about CCBot?

CCBot is Common Crawl's bot. Common Crawl produces a publicly accessible web archive used by many AI training datasets. Allowing CCBot makes your content available for general AI training across many providers. Most brands should allow it; a minority block it as part of broader AI training opt-out.

Should robots.txt list disallowed paths for AI bots?

If specific paths shouldn't be crawled (admin areas, search result pages with infinite parameter combinations, gated content), yes. Standard disallow patterns work the same for AI bots as for traditional search bots. Just be specific: Disallow: /admin/ not Disallow: /.

What's the relationship between robots.txt and llms.txt?

robots.txt controls crawler access. llms.txt is a proposed standard for declaring site identity to AI crawlers, but as we documented in our methodology change, llms.txt has no measurable impact on AI citation per SE Ranking's November 2025 study of nearly 300,000 domains. robots.txt matters; llms.txt is hygiene-flag only.


Companion guides: Schema markup for AI search · The Cloudflare GPTBot trap · AI bot user-agent reference · The 10-Point AI Citation Framework.

Your turn

Get your free AI visibility scan.

See whose name ChatGPT, Perplexity, Claude, Gemini, Grok and Google AI Mode say when someone asks about your business. Ten minutes, report by email, no card.

Free AI Citation Audit, see how ChatGPT, Perplexity, Claude, Gemini, Grok & Google AI Mode rate your site, and who they recommend instead. By continuing you agree to the Terms and Privacy Policy.