An AI crawler (AI bot) is an automated program that AI companies use to read, index, and fetch web content so their models can learn from it or cite it in answers. These AI bots now decide whether a brand appears in answers generated by ChatGPT, Gemini, Copilot, Perplexity, and other AI engines, which makes them a core concern for generative engine optimization (GEO). A page that AI bots cannot reach cannot be cited, no matter how strong its traditional SEO is. TOS maps every major AI crawler by role and priority so a website earns AI visibility instead of losing it by accident.
Updated: 08/2026.
- AI bots split into three roles: training, retrieval/search, and user-fetch.
- Retrieval and user-fetch bots decide citations now; training bots only shape long-term model knowledge.
- The GEO priority order: keep Tier 1 (OpenAI, Google, Microsoft, Perplexity, Anthropic) fully crawlable.
- Blocking a training bot is a copyright choice; never block an answer bot by mistake.
What is an AI crawler (AI bot)?
An AI crawler is a bot operated by an AI provider to collect web pages for one of three purposes: training a model, building a retrieval index that feeds an AI answer engine, or fetching a specific page in real time when a user asks about it. Functionally it behaves like any web crawler, requesting URLs, reading HTML, and following links, but its output feeds a large language model rather than a classic ranked results page.
The difference from a traditional search-engine crawler is the destination of the data. Googlebot historically fed a list of ten blue links; an AI crawler feeds a system that composes a synthesized answer and, in the better cases, cites the sources it drew from. Because the answer is generated rather than listed, being crawled is no longer enough on its own, the content also has to be structured clearly enough for a model to quote. That shift is why the same names now carry two jobs: Googlebot and Bingbot still serve their search engines, but the same indexes also feed AI Overviews and Copilot.
AI bots multiplied across 2024 to 2026 because every major AI product needs fresh web data. Training-only crawling came first; retrieval and user-fetch bots followed as ChatGPT, Perplexity, Claude, and Gemini gained live browsing and web search. A single site can now be visited by a dozen or more distinct AI user-agents, each with a different job and a different impact on visibility, which is why AI crawlers belong in the broader landscape of types of website bots that reach a site every day.
The 3 roles of AI bots: training, retrieval, and user-fetch
Every AI bot fills one of three roles, training, retrieval/search, or user-fetch, and only the last two decide whether a brand is cited in an AI answer today. Separating these roles is the single most important idea for managing AI crawlers correctly.
Training bots collect content for a model’s foundational training set. They shape what a model knows in general but do not decide whether a specific page is cited in a live answer, so blocking them is a copyright and data-control decision that does not remove existing citations. Retrieval and search-index bots build the index an AI engine consults in real time when it answers; this group directly determines whether a brand shows up in AI search results, which makes it the most important group for GEO. User-fetch and agent bots pull a page on demand, typically when a user pastes a link or asks the assistant to read a specific URL, so blocking them means users cannot look a page up through the assistant.
The practical insight, drawn from GEO practice in 2026, is blunt: blocking training bots such as GPTBot, ClaudeBot, and CCBot does not cost citations, while blocking retrieval or agent bots such as OAI-SearchBot, PerplexityBot, and ChatGPT-User does.
| Role | Example bots | Impact | Recommendation |
|---|---|---|---|
| Training | GPTBot, ClaudeBot, CCBot, Google-Extended (token), Applebot-Extended (token) | Feeds long-term model knowledge; does not decide immediate citation | Allow or block as a copyright choice; blocking does not lose citations |
| Retrieval / search index | OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot, Bingbot | Directly decides whether a brand is cited in AI answers | Always allow; the most important group for GEO |
| User-fetch / agent | ChatGPT-User, Perplexity-User, Claude-User, meta-externalfetcher | Decides whether AI can read a page when a user asks about it directly | Keep open |
The complete list of common AI bots (AI crawlers) in 2026
The list below covers the AI crawlers most commonly seen on websites in 2026, grouped by provider, with the role and priority tier that matter for GEO. Reading it by the role and priority columns is faster than memorizing user-agents, because the role decides the impact and the tier decides the order of attention.
| Bot / User-Agent | Provider | AI engine served | Role | Respects robots.txt | Priority |
|---|---|---|---|---|---|
| GPTBot | OpenAI | GPT model training | Training | Yes | Tier 2 |
| OAI-SearchBot | OpenAI | ChatGPT Search (indexing) | Retrieval | Yes | Tier 1 |
| ChatGPT-User | OpenAI | ChatGPT browsing on user request | User-fetch | Yes | Tier 1 |
| ClaudeBot | Anthropic | Claude training | Training | Yes | Tier 2 |
| Claude-SearchBot | Anthropic | Indexing for Claude web search | Retrieval | Yes | Tier 2 |
| Claude-User | Anthropic | Claude page fetch on user request | User-fetch | Yes | Tier 2 |
| Googlebot | Google Search and AI Overviews (shared index) | Retrieval | Yes | Tier 1 | |
| Google-Extended | robots.txt token for Gemini/Vertex training and grounding (no separate user-agent) | Training (token) | Token | Tier 2 | |
| GoogleOther | Internal research and development crawling | Training/other | Yes | Tier 3 | |
| Google-CloudVertexBot | Fetches pages for customer-configured Vertex AI grounding | User-fetch | Yes | Tier 3 | |
| Bingbot | Microsoft | Bing Search and Copilot (shared Bing index) | Retrieval | Yes | Tier 1 |
| PerplexityBot | Perplexity | Indexing for the Perplexity answer engine | Retrieval | Yes | Tier 1 |
| Perplexity-User | Perplexity | Page fetch on user request | User-fetch | Yes | Tier 1 |
| Applebot | Apple | Siri/Spotlight and Apple Intelligence | Retrieval | Yes | Tier 2 |
| Applebot-Extended | Apple | robots.txt token controlling use of content for Apple AI training | Training (token) | Token | Tier 2 |
| meta-externalagent | Meta | Crawling for Meta AI and Llama training | Training | Yes | Tier 2 |
| meta-externalfetcher | Meta | Page fetch on user request for Meta AI | User-fetch | Yes | Tier 2 |
| Amazonbot | Amazon | Alexa and Amazon AI services | Retrieval/other | Yes | Tier 2 |
| Bytespider | ByteDance | AI training (Doubao) | Training | Often ignores; very high frequency | Tier 4 |
| CCBot | Common Crawl | Open dataset feeding training of many models | Training | Yes | Tier 3 |
| YouBot | You.com | You.com answer engine | Retrieval | Yes | Tier 3 |
| DuckAssistBot | DuckDuckGo | DuckAssist AI summaries | Retrieval | Yes | Tier 3 |
| cohere-ai | Cohere | Training and support of Cohere models | Training | Yes | Tier 3 |
| AI2Bot (Ai2Bot-Dolma) | Allen Institute for AI | Open dataset for research | Training | Yes | Tier 3 |
| Diffbot | Diffbot | Building an AI knowledge graph | Training/other | Yes | Tier 4 |
| Timpibot | Timpi | Decentralized search index | Retrieval/other | Yes | Tier 4 |
| Omgilibot | Webz.io | Sells web data for AI training | Training | Yes | Tier 4 |
| ImagesiftBot | Hive/ImageSift | Collects images for vision models | Training (images) | Yes | Tier 4 |
OpenAI (ChatGPT)
OpenAI runs three distinct bots. GPTBot handles model training and sits in Tier 2, while OAI-SearchBot builds the ChatGPT Search index and ChatGPT-User fetches pages when a user asks, both Tier 1 for GEO. Allowing the two retrieval and user-fetch bots is what keeps a brand citable in ChatGPT, regardless of the training choice made for GPTBot.
Anthropic (Claude)
Anthropic separates ClaudeBot for training, Claude-SearchBot for indexing Claude’s web search, and Claude-User for on-demand page fetches, all respecting robots.txt. The older anthropic-ai and Claude-Web user-agents are historical aliases being phased out, not separate bots to plan around.
Google (Gemini, AI Overviews)
Googlebot is the key crawler, feeding both Google Search and AI Overviews from one shared index, which makes it Tier 1. Google-Extended is a robots.txt token rather than a separate crawler, controlling whether content trains Gemini and Vertex and is used for grounding. GoogleOther and Google-CloudVertexBot handle internal research and customer-configured Vertex grounding, both in Tier 3.
Microsoft (Copilot / Bing)
Bingbot is the crawler that matters, because Copilot draws on the Bing index rather than a separate Copilot crawler. Keeping Bingbot open therefore serves both Bing Search and Copilot at once, which places it in Tier 1.
Perplexity
Perplexity runs PerplexityBot to index content for its answer engine and Perplexity-User to fetch pages on demand. As a pure answer engine with a high rate of source citation, Perplexity is Tier 1, and both bots respect robots.txt.
Apple (Siri, Apple Intelligence, Spotlight)
Applebot feeds Siri, Spotlight, and Apple Intelligence, making it the retrieval bot to allow. Applebot-Extended is a robots.txt token controlling only whether content trains Apple’s AI models, not a separate crawler; blocking it does not remove a site from Applebot’s search role.
Meta (Meta AI, Llama)
Meta uses meta-externalagent to crawl for Meta AI and Llama training, and meta-externalfetcher to pull pages on user request. Both respect robots.txt. The older FacebookBot mainly generates link previews and is not an AI crawler.
Amazon
Amazonbot collects content for Alexa and Amazon’s AI services. It respects robots.txt and sits in Tier 2, reflecting Amazon’s large but more contained ecosystem.
ByteDance
Bytespider crawls for ByteDance’s AI training, including Doubao. It frequently ignores robots.txt and crawls at very high frequency, so it usually has to be managed at the server level. For businesses outside the TikTok and Chinese market it is a Tier 4 candidate to block.
Common Crawl
CCBot builds the open Common Crawl dataset, which feeds the training of a very large number of models, including many open-source ones. It respects robots.txt and sits in Tier 3; allowing it gives a site broad presence across the training data many models share.
Others
Several smaller engines and data collectors round out the list: YouBot (You.com) and DuckAssistBot (DuckDuckGo) for answer features, cohere-ai and AI2Bot for model training and research datasets, and Diffbot, Timpibot, Omgilibot, and ImagesiftBot for knowledge graphs, decentralized indexes, data resale, and image collection. Some scrapers also spoof AI user-agents or send none at all; those ignore robots.txt and can only be stopped at the server or WAF level. The community catalog Known Agents (formerly Dark Visitors) at knownagents.com tracks new AI user-agents as they appear.
Which AI bots visit websites the most? (data observed on TOS)
On TOS’s own systems, ByteDance’s Bytespider generated by far the most AI bot traffic, while several high-value answer bots visited far less often. The figures below are data observed on TOS’s systems in the first fifteen days of August 2026, counting only clearly identified AI user-agents; they describe one agency’s servers, not the industry as a whole.
| Bot | Provider | Approx. hits (15 days, Aug 2026) |
|---|---|---|
| Bytespider | ByteDance | ~428,800 |
| Applebot | Apple | ~68,000 |
| GPTBot | OpenAI | ~48,000 |
| ClaudeBot | Anthropic | ~46,000 |
| YouBot | You.com | ~39,000 |
| Amazonbot | Amazon | ~14,000 |
| ChatGPT-User | OpenAI | ~11,700 |
| meta-externalagent | Meta | ~4,200 |
| CCBot | Common Crawl | ~3,400 |
| PerplexityBot | Perplexity | ~2,400 |
| OAI-SearchBot | OpenAI | ~1,100 |
| Google-Extended | ~250 (token / opt-in, low count is normal) |
- Bytespider created the largest load, roughly 428,800 hits, about 13% of total traffic, yet delivers little visibility or citation value for most businesses outside the TikTok and Chinese ecosystem, making it the first candidate to block when conserving server resources.
- Hit count is a poor measure of importance: high-value answer bots such as OAI-SearchBot and PerplexityBot appear far less often than training bots such as GPTBot and ClaudeBot, yet they are the ones that decide whether a brand is cited.
- TOS blocked Bytespider at the server level while keeping every GEO-relevant AI bot open, a practical example of blocking by intent rather than by volume.
Priority order: which AI bots to allow and optimize for first?
The priority order for allowing AI bots runs from the retrieval and user-fetch bots of the most widely used engines down to high-load training crawlers with little citation value. Three principles produce that ranking.
First, retrieval and user-fetch bots outrank training bots, because they decide citations immediately while training only shapes long-term model knowledge. Second, within the same role, bots are ranked by the user reach of the AI engine they serve, that is, how many people actually use that assistant. Third, relevance to a specific customer base can move a bot up or down; Bytespider, for example, is only worth prioritizing for a business targeting the TikTok or Chinese market.
| Priority / Tier | Provider and bots | Primary role | Why this tier |
|---|---|---|---|
| Tier 1 | OpenAI: OAI-SearchBot, ChatGPT-User (plus GPTBot for training) | Retrieval and user-fetch | The most widely used AI assistant; blocking removes a brand from ChatGPT answers immediately |
| Tier 1 | Google: Googlebot (plus enabling Google-Extended for Gemini) | Retrieval | Feeds AI Overviews across billions of searches from the shared Search index |
| Tier 1 | Microsoft Copilot / Bing: Bingbot | Retrieval | Copilot uses the Bing index, so one bot serves both |
| Tier 1 | Perplexity: PerplexityBot, Perplexity-User | Retrieval and user-fetch | A pure answer engine with a high source-citation rate |
| Tier 1 | Anthropic / Claude: Claude-SearchBot, Claude-User (plus ClaudeBot for training) | Retrieval and user-fetch | Feeds Claude web search and on-demand reading |
| Tier 2 | Apple Intelligence: Applebot (consider Applebot-Extended) | Retrieval | Growing fast on a very large device ecosystem |
| Tier 2 | Meta AI: meta-externalagent, meta-externalfetcher | Training and user-fetch | Large ecosystem coverage for Meta AI and Llama |
| Tier 2 | Amazon: Amazonbot | Retrieval/other | Feeds Alexa and Amazon AI services |
| Tier 3 | Common Crawl: CCBot | Training | Feeds many models; open for broad presence in training data |
| Tier 3 | You.com (YouBot), DuckDuckGo (DuckAssistBot), Cohere (cohere-ai), AI2 (AI2Bot) | Retrieval or training | Smaller answer engines and indirect data sources |
| Tier 4 | Bytespider (if not targeting TikTok/China), Omgilibot, ImagesiftBot, Diffbot, and spoofing scrapers | Training/other | High load or low citation value, or non-compliant with robots.txt |
For most businesses, the number-one GEO priority is making sure the Tier 1 group, OpenAI, Google, Microsoft, Perplexity, and Anthropic, can crawl all important content, and is never blocked by accident through robots.txt, a firewall, or a country block. Whether to allow training bots is a separate copyright decision; it should never be the reason an answer bot gets blocked by mistake.
How to manage AI bots the right way
The right way to manage AI bots is to open access by default for answer and retrieval bots, block only high-load or non-compliant crawlers on purpose, and verify bot identity before blocking. The controls differ by bot type, so the method matters as much as the intent.
- Control compliant bots per user-agent in robots.txt, for example a User-agent: GPTBot block with its own Disallow rules.
- Treat Google-Extended and Applebot-Extended as opt-out tokens for training only; they do not block search indexing, so using them does not remove a site from AI Overviews or from Applebot’s search role.
- Block non-compliant or spoofing bots at the server or WAF level, because they ignore robots.txt and cannot be stopped by it.
- Verify a bot is genuine by matching its requests against the provider’s published IP ranges and a reverse DNS check, to avoid blocking a real bot or trusting a fake one.
- Keep answer and retrieval bots open by default, and only block high-load, low-value crawlers deliberately.
- Do not use hit count to decide what to keep or block; a low-volume bot such as OAI-SearchBot can still be the one that decides a citation.
- Review server logs quarterly, since new AI user-agents appear continuously, and use community catalogs of AI user-agents to keep the list current.
How AI bots affect GEO/AIO
Being crawled by the right AI bots is the precondition for being cited in AI answers, so AI bot management is the foundation of GEO and AIO. Generative engine optimization and AI optimization both assume the content can actually be reached; if a retrieval bot is blocked, even the best-optimized page never enters the pool an engine can quote.
The failure mode is silent: a page keeps ranking in classic search while quietly vanishing from AI answers, because a robots.txt rule, a firewall, or a country block is turning away OAI-SearchBot or PerplexityBot. Nothing in a standard SEO report flags this, which is why AI crawl access has to be audited as its own layer. TOS builds this check into its GEO service, verifying that Tier 1 and Tier 2 bots reach every important URL and that structured content is easy for a model to quote. A fast way to gauge current standing is the free AI visibility check tool, which shows whether AI engines can see and cite a brand.
TOS can help with your GEO
TOS designs and runs GEO and AIO programs that keep a brand visible and citable across every major AI engine. The work starts by confirming that no Tier 1 or Tier 2 AI crawler is blocked, then structures content so models quote it, and monitors AI citations over time. Teams that want a fast baseline can run the free AI visibility check tool or reach out to TOS to build a full plan for earning AI citations.
Conclusion
Managing AI crawlers is now a core part of digital visibility, because these bots decide whether a brand appears in the answers people read on ChatGPT, Gemini, Copilot, and Perplexity. The complete list of AI bots is long and still growing, but the framework that makes it usable is simple: sort every AI crawler into training, retrieval, or user-fetch, and treat retrieval and user-fetch as the groups that decide citations today. From that framework the priority order follows naturally. Tier 1, meaning OpenAI, Google, Microsoft, Perplexity, and Anthropic, must be able to crawl every important page, because blocking any of them removes a brand from a major AI answer engine at once. Training bots are a separate, optional copyright decision that does not affect current citations. High-volume, low-value crawlers such as Bytespider can be blocked at the server level without harming GEO. The single most costly error is blocking an answer bot by accident, and a periodic log review is the cheapest way to prevent it. Getting AI crawler access right is what turns strong content into an AI citation.
Frequently asked questions
How is an AI crawler different from Googlebot?
An AI crawler feeds a generative model, while Googlebot historically fed a ranked list of links. The distinction is blurring: Googlebot now supplies both classic search results and AI Overviews from the same index, so it is both a search crawler and an AI retrieval crawler. Dedicated AI bots such as OAI-SearchBot or PerplexityBot exist only to feed answer engines. The practical difference is that an AI crawler’s goal is to let a model quote or summarize a page, not merely rank it.
Should I block GPTBot?
Blocking GPTBot is optional and purely a training decision. GPTBot collects content to train OpenAI’s models; it does not decide whether a page is cited in ChatGPT answers. That job belongs to OAI-SearchBot and ChatGPT-User, which must stay allowed for GEO. Blocking GPTBot only opts content out of future training, a copyright choice some publishers make. A common mistake is blocking GPTBot and assuming ChatGPT visibility is unaffected; it is unaffected, but only if the retrieval and user-fetch bots remain open.
Which AI bot matters most for GEO?
Retrieval bots matter most for GEO, because they build the index an AI engine cites in real time. In practice that means Googlebot (which feeds AI Overviews), Bingbot (which feeds Copilot), OAI-SearchBot, PerplexityBot, and Claude-SearchBot. These sit in Tier 1 and Tier 2 of the TOS priority order. Hit count is misleading here: OAI-SearchBot may visit far less often than a training bot, yet it decides whether ChatGPT Search cites a page. Keeping every retrieval bot fully crawlable is the first GEO priority.
How do I know if AI bots crawl my site?
Check server access logs and filter by the known AI user-agents, such as GPTBot, ClaudeBot, PerplexityBot, and OAI-SearchBot. Server logs are more reliable than analytics tools, because most bots do not run JavaScript and never appear in tag-based analytics. To confirm a bot is genuine and not a spoofed user-agent, match its IP against the provider’s published ranges and run a reverse DNS check. Community catalogs of new AI user-agents help keep the list current, and reviewing logs quarterly catches new bots as they appear.
Does blocking AI bots cost me traffic?
Blocking retrieval and user-fetch bots costs AI visibility, which increasingly translates into lost traffic and lost brand mentions as users move questions to AI assistants. Blocking training bots does not cost citations today, so it does not directly reduce AI-driven traffic; it only opts content out of future model training. The real risk is accidental over-blocking, when a broad robots.txt rule, a firewall filter, or a country block catches answer bots along with training bots. Precision matters: block by intent, never by a blanket rule.
Should I block Bytespider?
Blocking Bytespider is reasonable for most businesses that do not target the TikTok or Chinese market. In data observed on TOS’s systems in August 2026, Bytespider generated by far the highest bot load, roughly 428,800 hits in fifteen days, about 13% of total traffic, while offering little citation value outside ByteDance’s ecosystem. It also frequently ignores robots.txt, so it usually has to be blocked at the server or WAF level rather than through robots.txt. TOS blocked Bytespider while keeping every GEO-relevant bot open.

