← All articles·technical·Supporting

AI bot User-Agent reference: every crawler that matters in 2026

The complete reference for AI bot user agents in 2026. GPTBot, ChatGPT-User, OAI-SearchBot, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Google-Extended, Applebot-Extended, Bytespider, CCBot, cohere-ai, Meta-ExternalAgent. Owner, purpose, default treatment, recommended configuration.

Data for AI Search Editorial Team··10 min read

This is the reference document for AI bot user agents that matter for AEO and GEO optimization in 2026. Thirteen user agents account for the AI bot traffic that affects citation behavior across ChatGPT, Perplexity, Claude, Gemini, Grok, Microsoft Copilot, and emerging AI assistants. Each user agent has a distinct purpose (training-data ingestion, real-time retrieval, image indexing), owner, and default treatment under common infrastructure configurations. Brands optimizing for AI citation need to verify each user agent can reach the site, and need to know what each one does so configuration decisions are deliberate rather than accidental. The reference is structured for the use case: a brand auditing their robots.txt, Cloudflare AI Crawl Control panel, or WAF configuration needs to know what each user agent does and which AI assistant citation depends on it. Per the 10-Point AI Citation Audit, crawler accessibility (Check 1) carries the veto: a robots.txt block on a platform's search crawler forces that platform's score to zero.

Which AI bot user agents matter in 2026?

Thirteen AI bot user agents matter for citation as of mid-2026, and the table lists each one with its owner, its purpose, the assistant it affects and its default treatment.

User-Agent stringOwnerPurposeAffectsDefault treatment
GPTBotOpenAITraining data ingestionWhat OpenAI's models learn, not ChatGPT search answersDisallowed when Cloudflare's managed robots.txt setting is on
ChatGPT-UserOpenAIUser-action retrievalChatGPT real-timeUsually allowed
OAI-SearchBotOpenAIChatGPT Search retrievalWhether a site appears in ChatGPT search answersUsually allowed
ClaudeBotAnthropicTraining data ingestionWhat Claude models learn, not Claude search answersUsually allowed
Claude-SearchBotAnthropicSearch index crawlingClaude web searchUsually allowed
Claude-UserAnthropicUser-action retrievalClaude real-timeUsually allowed
PerplexityBotPerplexitySearch index crawlingPerplexity citationSometimes blocked
Google-ExtendedGoogleRobots.txt token for Gemini training and groundingGeminiUsually allowed
Applebot-ExtendedAppleRobots.txt token for Apple model trainingApple IntelligenceUsually allowed
BytespiderByteDanceDoubao / TikTok trainingDoubao + TikTok AISometimes blocked
CCBotCommon CrawlOpen web archiveMany AI providers via Common CrawlUsually allowed
cohere-aiCohereCohere model trainingCohere model citationUsually allowed
Meta-ExternalAgentMetaMeta AI trainingMeta AI citationUsually allowed

Source: compiled by Data for AI Search, July 2026, for the crawler accessibility check in the 10-Point AI Citation Audit. The default treatment column is our own characterization, not a vendor statement. Names and purposes were checked in September 2026 against the documentation published by OpenAI, Anthropic, Perplexity and Google.

Below is the per-user-agent reference.

What is GPTBot?

GPTBot is OpenAI's crawler for training data ingestion, and it affects what OpenAI's models learn, not whether a site appears in ChatGPT search answers.

Owner: OpenAI Purpose: Training data ingestion for GPT models. This is the primary OpenAI crawler for absorbing web content into future model training cycles. Affects: What OpenAI's models know without searching (training-corpus signal). A brand that blocks GPTBot is less likely to be known to the model itself, but blocking it does not remove the site from ChatGPT search answers, which depend on OAI-SearchBot. Default treatment: Depends on the site's configuration. When Cloudflare's managed robots.txt setting is switched on, it writes a rule that disallows GPTBot, as of September 2026. See The Cloudflare GPTBot trap. Recommended treatment: Allow. Documentation: https://developers.openai.com/api/docs/bots

Sample robots.txt:

User-Agent: GPTBot
Allow: /

What is ChatGPT-User?

ChatGPT-User is the OpenAI user agent for real-time retrieval triggered by ChatGPT user actions.

Owner: OpenAI Purpose: Real-time retrieval triggered by ChatGPT user actions (browsing, fetching specific URLs, performing web actions). Affects: ChatGPT real-time citation behavior when users explicitly request URL fetches or web actions. Default treatment: Usually allowed. Recommended treatment: Allow.

What is OAI-SearchBot?

OAI-SearchBot is the crawler that OpenAI names, as of September 2026, as deciding whether a site is shown in ChatGPT search answers.

Owner: OpenAI Purpose: ChatGPT Search retrieval, the crawler that surfaces websites in ChatGPT's search features. Affects: Whether a site appears in ChatGPT search answers. OpenAI says a site opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though it can still appear as a navigational link. Critical for time-sensitive queries. Default treatment: Usually allowed. Recommended treatment: Allow.

What is ClaudeBot?

ClaudeBot is Anthropic's primary crawler, used for training data ingestion for Claude models.

Owner: Anthropic Purpose: Training data ingestion for Claude models. The primary Anthropic crawler. Affects: What Claude models know without searching (training-corpus signal). A brand that blocks ClaudeBot is less likely to be known to the model itself, but Anthropic ties visibility in Claude's search answers to Claude-SearchBot and Claude-User, not to ClaudeBot. Default treatment: Usually allowed. Recommended treatment: Allow.

What is Claude-SearchBot?

Claude-SearchBot is the Anthropic crawler that indexes web content to improve search results for Claude users.

Owner: Anthropic Purpose: Crawls the web to improve the quality of search results for Claude users. It replaced the older anthropic-ai name, which Anthropic no longer documents as of September 2026. Affects: Claude web search. Anthropic's crawler documentation says that disabling it may reduce a site's visibility and accuracy in user search results. Default treatment: Usually allowed. Recommended treatment: Allow.

What is Claude-User?

Claude-User is the Anthropic user agent that fetches a page when a Claude user asks a question that needs it.

Owner: Anthropic Purpose: Retrieval started by a Claude user. It replaced the older Claude-Web name, which Anthropic no longer documents as of September 2026. Affects: Claude real-time citation behavior. Anthropic says that disabling it prevents Claude from retrieving a site's content in response to a user query. Default treatment: Usually allowed. Recommended treatment: Allow.

What is PerplexityBot?

PerplexityBot is Perplexity's crawler for surfacing and linking websites in its search results, and it affects Perplexity citation.

Owner: Perplexity Purpose: Search results. Perplexity's crawler documentation says, as of September 2026, that PerplexityBot surfaces and links websites in Perplexity search results and is not used to crawl content for AI foundation models. A second agent, Perplexity-User, fetches a page when a user asks a question. Affects: Perplexity citation. Critical because Perplexity is heavily real-time-retrieval-dominant per the Two-Track Law. Default treatment: Sometimes blocked by aggressive WAF rules that catch "bot" user agents broadly. Recommended treatment: Allow. Note: PerplexityBot can generate high-volume requests for popular domains. Our recommendation as of July 2026 is to configure rate-limiting generously (>1000 requests/hour from the user agent).

What is Google-Extended?

Google-Extended is a robots.txt control token for Gemini, not a crawler, and it is distinct from Googlebot.

Owner: Google Purpose: Controls whether content Google crawls may be used to train Gemini models and to ground answers in the Gemini app and Vertex AI. It has no user agent string of its own: Google crawls with its existing crawlers, so the token is set in robots.txt and cannot be tested with a request or managed at a firewall. Affects: Gemini citation. Google AI Overviews and AI Mode are Search features and follow Googlebot, not this token, as of September 2026 (Google's crawler documentation). Default treatment: Usually allowed because most sites that allow Googlebot also implicitly allow Google-Extended via User-agent: *. Recommended treatment: Allow. Important: Blocking Google-Extended does NOT affect Google Search ranking, AI Overviews or AI Mode. Google states that the token "does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search."

What is Applebot-Extended?

Applebot-Extended is Apple's robots.txt control for AI training, not a crawler, and it is distinct from Applebot.

Owner: Apple Purpose: Controls whether content crawled by Applebot may be used to train Apple's foundation models. Apple's documentation says, as of September 2026, that Applebot-Extended does not crawl webpages. Applebot itself serves Spotlight, Siri and Safari. Affects: Apple Intelligence citation. Still emerging surface as of mid-2026. Default treatment: Usually allowed. Recommended treatment: Allow.

What is Bytespider?

Bytespider is ByteDance's crawler for training its AI models, including Doubao and TikTok AI features.

Owner: ByteDance Purpose: Training data ingestion for ByteDance AI models including Doubao (the Chinese-market ChatGPT competitor) and TikTok AI features. Affects: Doubao citation + TikTok AI features. Relevant primarily for brands targeting Chinese-speaking markets or with significant TikTok presence. Default treatment: Sometimes blocked by WAF rules targeting Chinese-origin crawlers broadly. Recommended treatment: Allow for international brands; allow for brands with TikTok strategy; defer or block for brands with strict Chinese-market avoidance policies.

What is CCBot?

CCBot is the Common Crawl crawler, which builds a public web archive that many AI training datasets incorporate.

Owner: Common Crawl Purpose: Building the Common Crawl public web archive, which many AI training datasets incorporate. Affects: Many AI providers indirectly via Common Crawl ingestion. Default treatment: Usually allowed. Recommended treatment: Allow. Blocking CCBot is a defensible choice for brands philosophically opposed to AI training but reduces the open-web representation of the brand across many AI training datasets.

What is cohere-ai?

The cohere-ai user agent is Cohere's crawler for training its enterprise-focused language models.

Owner: Cohere Purpose: Training data ingestion for Cohere's enterprise-focused language models. Affects: Cohere model citation. Limited relevance for consumer-facing AEO; significant for enterprise B2B brands where Cohere adoption is meaningful. Default treatment: Usually allowed. Recommended treatment: Allow.

What is Meta-ExternalAgent?

Meta-ExternalAgent is Meta's crawler for training its AI models, including Llama and the Meta AI assistant.

Owner: Meta Purpose: Training data ingestion for Meta AI models including Llama and the consumer-facing Meta AI assistant. Affects: Meta AI citation. Default treatment: Usually allowed. Recommended treatment: Allow.

What does a complete robots.txt look like?

A complete robots.txt names each AI bot and gives it an explicit allow. This is the gold-standard explicit-allow pattern recommended in the AI bot robots.txt complete guide, applied as a complete reference file:

User-Agent: *
Allow: /

User-Agent: GPTBot
Allow: /

User-Agent: ChatGPT-User
Allow: /

User-Agent: OAI-SearchBot
Allow: /

User-Agent: ClaudeBot
Allow: /

User-Agent: Claude-SearchBot
Allow: /

User-Agent: Claude-User
Allow: /

User-Agent: PerplexityBot
Allow: /

User-Agent: Google-Extended
Allow: /

User-Agent: Applebot-Extended
Allow: /

User-Agent: CCBot
Allow: /

User-Agent: Bytespider
Allow: /

User-Agent: cohere-ai
Allow: /

User-Agent: Meta-ExternalAgent
Allow: /

Sitemap: https://yourdomain.com/sitemap.xml

Frequently asked questions

What about newer or less common AI bots?

The list above covers the AI bots that materially affect citation as of mid-2026. New bots may emerge: xAI's crawler for Grok training (currently undocumented user agent), Mistral's training crawler, others. The list will update as new bots cross meaningful citation impact thresholds.

Should I differentiate per-section access for AI bots?

Generally no for typical brands. Allowing AI bots to crawl the entire site (except genuinely private paths like admin areas) maximizes citation surface. The exceptions: paywalled content (block AI bots from the paywalled paths to protect the business model), highly sensitive proprietary content, and explicitly opt-out user-generated content.

How do I know if a user agent claiming to be an AI bot is legitimate?

User agent strings can be spoofed. The legitimate AI providers publish the IP addresses their crawlers use. As of September 2026, OpenAI, Anthropic and Perplexity each publish IP address lists, so a request that claims to be GPTBot or ClaudeBot should come from an address on that vendor's list. Google documents reverse DNS verification for Googlebot as well as published IP ranges. WAF configurations can include verification rules that require both the user-agent string AND verified origin IP, rejecting spoofed requests.

Does blocking AI bots affect my SEO?

Not directly. AI bots are separate from search engine bots. Blocking GPTBot doesn't affect Googlebot's crawl or Google Search ranking. Google-Extended is a separate case: it is a robots.txt token for Gemini training and grounding. Blocking it doesn't affect Google Search, AI Overviews or AI Mode, but does affect Gemini citation.

What if I want to allow training for some AI providers but not others?

The toggles are independent. A brand can allow Claude (ClaudeBot, Claude-SearchBot, Claude-User) while blocking GPTBot, or any other combination. The choice should reflect strategic AI citation priorities and any training-data preferences the brand holds. Most commercial brands allow all major AI bots; some philosophically-conscious brands block the training controls (GPTBot, ClaudeBot, Google-Extended) while allowing the search and retrieval agents (OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User).


Companion guides: The AI bot robots.txt complete guide · The Cloudflare GPTBot trap · Schema markup for AI search · The 10-Point AI Citation Framework.

Your turn

Get your free AI visibility scan.

See whose name ChatGPT, Perplexity, Claude, Gemini, Grok and Google AI Mode say when someone asks about your business. Ten minutes, report by email, no card.

Free AI Citation Audit, see how ChatGPT, Perplexity, Claude, Gemini, Grok & Google AI Mode rate your site, and who they recommend instead. By continuing you agree to the Terms and Privacy Policy.