AI crawlers
AI crawler bot guides
User-agent guides for GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and more—what they do, how to configure robots.txt, and how BatSignal measures access.
GPTBotOpenAI · training — Fetches public pages for OpenAI training-related crawling. Often managed separately from ChatGPT search bots in robots.txt.OAI-SearchBotOpenAI · search — Supports ChatGPT search retrieval by fetching pages that may ground answers with web sources.ChatGPT-UserOpenAI · agent — Used when ChatGPT browses the live web during an interactive user session.ClaudeBotAnthropic · training — Anthropic training-oriented crawler that fetches public web content subject to robots.txt.Claude-WebAnthropic · search — Associated with Claude web retrieval rather than bulk training crawl alone.anthropic-aiAnthropic · training — Additional Anthropic-related crawler token seen in robots policies and bot lists.PerplexityBotPerplexity · search — Fetches sources that may appear in Perplexity answer citations.Google-ExtendedGoogle · training — Used in robots.txt to signal preferences related to Google AI / Gemini training controls, separate from Googlebot search.CCBotCommon Crawl · training — Builds the Common Crawl public web archive used widely in research and some training corpora.BytespiderByteDance · training — ByteDance crawler that appears in AI-related robots policies.AmazonbotAmazon · search — Amazon crawler associated with Alexa and related discovery surfaces.meta-externalagentMeta · training — Meta AI crawler user-agent used for AI-related fetching.Applebot-ExtendedApple · training — Apple user-agent related to Apple Intelligence training controls, distinct from Applebot search crawling.DuckAssistBotDuckDuckGo · search — Associated with DuckDuckGo assistive AI answer features.cohere-aiCohere · training — Cohere-related crawler token seen in AI robots discussions and policies.GooglebotGoogle · search — Classic Google Search crawler—still foundational because AI overviews and discovery often depend on a healthy Google index.BingbotMicrosoft · search — Microsoft Bing crawler; Bing grounding can influence Copilot-style experiences.SlurpYahoo · search — Yahoo’s crawler user-agent, still seen in logs and robots examples.YandexYandex · search — Yandex search crawler important for certain geographic markets.FacebookBotMeta · search — Fetches URLs shared on Facebook for previews and related features.