The AI Crawler Reference for Publishers: Who to Block, Who to Allow (2026)

Updated September 2026. Last verified against each vendor's own documentation: September 2026.

When we first published this guide in 2025, the advice for most publishers was simple: block the AI bots. That advice has aged badly. Today the same companies run separate crawlers for training their models and for citing sources in AI search, and blocking the wrong one can quietly cut you out of ChatGPT, Perplexity, and Claude answers that send readers back to your site.

This reference sorts every major crawler by what it actually does, gives you our call on each, and flags the ones robots.txt can't stop.

Sign up for our newsletter to get the printable PDF

This is a starting point, not a mandate. Where you land depends on the posture you set: full protection, the managed middle, or full extraction.

robots.txt is the sign on the door. Your CDN is the lock. Some crawlers ignore robots.txt, so enforce anything that matters at the CDN or bot-management layer, not just in robots.txt.

Re-audit every 90 days. These definitions change fast. Confirm current tokens against each vendor's own documentation before you deploy.

Start with your posture

There is no single right answer. Where you land depends on the posture you set:

  • Full protection: Block training and AI search crawlers alike. Best for paywalled archives or publishers pursuing licensing deals, at the cost of AI search visibility.
  • The managed middle: Block training crawlers, allow search and citation crawlers. You stay visible in AI answers without contributing your archive to model training. For most publishers, this is the right starting point.
  • Full extraction: Allow everything. Maximum exposure in AI products, and a reasonable choice for publishers who prioritize reach over archive control.

The tables below give our default call for each crawler.

robots.txt is the sign on the door. Your CDN is the lock.

robots.txt is an honor-system request. Most major AI companies follow it, but some crawlers ignore it, and others show up looking like ordinary browser traffic. Anything that truly matters to you should be enforced at the CDN or bot-management layer, not just in robots.txt. It is the same approach we recommend for suspicious bot traffic like the recent Gansu, China spike: confirm what is hitting your servers, then block or challenge it at the edge.

Search engine crawlers

Your existing search traffic. Allow all unless you have a specific reason not to.

User agent Operator and purpose TFD call Notes
Googlebot Google Search Allow Your core search traffic, and the index behind AI Overviews and AI Mode.
Bingbot Microsoft Bing Allow Feeds Bing and Microsoft Copilot's answers.
DuckDuckBot DuckDuckGo Allow Traditional results come largely from Bing; DuckDuckGo also runs its own indexes.
Slurp Yahoo Search Allow Also collects partner content for Yahoo News, Finance, and Sports.
Baiduspider Baidu (China) Allow Relevant only for a China audience.
YandexBot Yandex (Russia) Allow Low relevance for US audiences.
PetalBot Huawei Petal Search Allow Also feeds Huawei Assistant. Follows robots.txt; no documented AI-training use. Low US relevance.

AI search and citation crawlers

This is how AI cites you and sends readers back. Allow these if you want to appear in AI answers.

User agent Operator and purpose TFD call Notes
OAI-SearchBot OpenAI, ChatGPT Search index Allow Opting out drops you from ChatGPT search answers, though you can still appear as a plain link. Separate from GPTBot.
ChatGPT-User OpenAI, user-initiated fetches Allow Fires when a ChatGPT user, Custom GPT, or action asks. May not follow robots.txt.
PerplexityBot Perplexity, search index Allow Cites and links. Cloudflare (Aug 2025) reported additional undeclared Perplexity crawling.
Perplexity-User Perplexity, user-initiated fetches Allow Perplexity's docs say it generally ignores robots.txt.
Claude-SearchBot Anthropic, Claude search index Allow Separate from ClaudeBot (training).
Claude-User Anthropic, user-initiated fetches Allow Fetches a page when a Claude user asks. Follows robots.txt.
Meta-WebIndexer Meta, AI search index Allow Allowing it helps Meta AI cite you. Distinct from Meta's training crawler.
Applebot Apple, Siri, Spotlight, Safari Allow Training is controlled separately (Applebot-Extended, below).
Amzn-User Amazon, user-initiated fetches Allow Fires when an Amazon user or assistant asks. May not follow all robots.txt rules. Separate from Amazonbot.
MistralAI-User / -Index Mistral, Vibe fetches and index Allow Behind Mistral's Vibe (formerly Le Chat).
DuckAssistBot DuckDuckGo, AI answers Allow Cites sources. Not used to train AI models.
YouBot You.com search Allow Powers You.com search.

AI training, scraping, and extraction crawlers

Your call: block to protect your archive, allow if you want the exposure, or license where it pays. Blocking training does not remove you from AI search indexes, with one exception: Google-Extended also controls whether Gemini grounds answers in your content.

User agent Operator and purpose TFD call Notes
GPTBot OpenAI, model training Your call Blocking does not affect ChatGPT Search visibility (that is OAI-SearchBot).
ClaudeBot Anthropic, model training Your call Anthropic's current training crawler. It also honors robots.txt rules written for the retired "anthropic-ai" and "Claude-Web" tokens (Anthropic, 2024).
CCBot Common Crawl Your call An open crawl reused by many AI trainers. Blocking stops future crawls only.
Google-Extended Google, Gemini training and grounding Your call A robots.txt token, not a bot in your logs. Controls Gemini training and grounding in Gemini Apps. Google says it does not affect Search ranking or inclusion.
GoogleOther Google, internal R&D Your call Not documented as a training crawler; Google says blocking it affects no product.
Applebot-Extended Apple, Apple Intelligence training Your call A token, not a bot. Controls training only; you stay in Siri, Spotlight, and Safari.
Meta-ExternalAgent Meta, AI training and content indexing Your call Meta's crawler for AI training and direct indexing. Not the same as facebookexternalhit, see Watch-outs.
Meta-ExternalFetcher Meta, user-initiated fetch Your call Fetches pages when users ask. May bypass robots.txt. Meta also uses these fetches to improve its AI agents.
Amazonbot Amazon, products/services and AI training Your call Respects robots.txt. A page-level noarchive meta tag opts that page out of training.
Bytespider ByteDance / TikTok, reportedly AI training Enforce at CDN Widely reported to ignore robots.txt. Enforce at the CDN if it matters.
Diffbot Diffbot, knowledge-graph and search extraction Your call Extraction, not foundation-model training. Respects robots.txt by default; its Crawlbot or Extract customers can override that.
Webzio Webz.io, data collection and resale Your call Formerly "Omgilibot." Its "Webzio-extended" companion tags crawled data as usable or not for AI/ML training.
img2dataset Open-source scraping tool Enforce at CDN A tool anyone can run, not a company bot. Ignores robots.txt and uses a generic browser user agent, so a name-based block misses default runs. By default it does skip images served with an X-Robots-Tag: noai or noimageai header, the one control you have here.

Watch-outs: do not blanket-block these

The mistakes that cost publishers traffic they meant to keep.

User agent What it really is TFD call Notes
facebookexternalhit Meta, link preview and unfurling Allow Builds the preview cards when your links are shared on Facebook, Instagram, and Messenger. That is its documented job, not AI training, so blocking it breaks those previews. May bypass robots.txt for security checks.
Undeclared / stealth Crawlers that pose as ordinary browsers Enforce at CDN Some crawlers appear as normal browser traffic or ignore robots.txt. Cloudflare reported this of Perplexity in Aug 2025. robots.txt cannot reliably stop them; use CDN or bot-management rules.

A starting robots.txt for the managed middle

If you land in the managed middle, this blocks the major training crawlers while leaving search and citation crawlers untouched. Anything not listed is allowed by default.

# Managed middle: block AI training, allow search and citation
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

Two deliberate omissions:

  • Google-Extended is left out because it also controls whether Gemini grounds answers in your content. Blocking it is not a clean training-only block, so make that choice on purpose.
  • Bytespider and img2dataset are left out because robots.txt will not stop them. Handle those at the CDN.

Keep it current

AI crawlers and their robots.txt tokens change frequently. This reference is point-in-time as of September 2026. Confirm current details against each vendor's documentation before deploying, and re-audit roughly every 90 days.

Deciding who can crawl your site is only half of the job. The other half is making sure the AI systems you allow in get your story right. That is what a dedicated AI page is for.

Get the printable reference

Newsletter subscribers get the printable PDF. It's the full AI Crawler Reference in a format you can share with your editorial, product, and dev teams.

Need help?

Not sure what's crawling your site today? We audit server logs and robots.txt for publishers and set up CDN rules for the crawlers that don't follow them. Talk to us.