Updated September 2026. Last verified against each vendor's own documentation: September 2026.
When we first published this guide in 2025, the advice for most publishers was simple: block the AI bots. That advice has aged badly. Today the same companies run separate crawlers for training their models and for citing sources in AI search, and blocking the wrong one can quietly cut you out of ChatGPT, Perplexity, and Claude answers that send readers back to your site.
This reference sorts every major crawler by what it actually does, gives you our call on each, and flags the ones robots.txt can't stop.
Sign up for our newsletter to get the printable PDF
This is a starting point, not a mandate. Where you land depends on the posture you set: full protection, the managed middle, or full extraction.
robots.txt is the sign on the door. Your CDN is the lock. Some crawlers ignore robots.txt, so enforce anything that matters at the CDN or bot-management layer, not just in robots.txt.
Re-audit every 90 days. These definitions change fast. Confirm current tokens against each vendor's own documentation before you deploy.
Start with your posture
There is no single right answer. Where you land depends on the posture you set:
- Full protection: Block training and AI search crawlers alike. Best for paywalled archives or publishers pursuing licensing deals, at the cost of AI search visibility.
- The managed middle: Block training crawlers, allow search and citation crawlers. You stay visible in AI answers without contributing your archive to model training. For most publishers, this is the right starting point.
- Full extraction: Allow everything. Maximum exposure in AI products, and a reasonable choice for publishers who prioritize reach over archive control.
The tables below give our default call for each crawler.
robots.txt is the sign on the door. Your CDN is the lock.
robots.txt is an honor-system request. Most major AI companies follow it, but some crawlers ignore it, and others show up looking like ordinary browser traffic. Anything that truly matters to you should be enforced at the CDN or bot-management layer, not just in robots.txt. It is the same approach we recommend for suspicious bot traffic like the recent Gansu, China spike: confirm what is hitting your servers, then block or challenge it at the edge.
Search engine crawlers
Your existing search traffic. Allow all unless you have a specific reason not to.
| User agent | Operator and purpose | TFD call | Notes |
|---|---|---|---|
Googlebot |
Google Search | Allow | Your core search traffic, and the index behind AI Overviews and AI Mode. |
Bingbot |
Microsoft Bing | Allow | Feeds Bing and Microsoft Copilot's answers. |
DuckDuckBot |
DuckDuckGo | Allow | Traditional results come largely from Bing; DuckDuckGo also runs its own indexes. |
Slurp |
Yahoo Search | Allow | Also collects partner content for Yahoo News, Finance, and Sports. |
Baiduspider |
Baidu (China) | Allow | Relevant only for a China audience. |
YandexBot |
Yandex (Russia) | Allow | Low relevance for US audiences. |
PetalBot |
Huawei Petal Search | Allow | Also feeds Huawei Assistant. Follows robots.txt; no documented AI-training use. Low US relevance. |
AI search and citation crawlers
This is how AI cites you and sends readers back. Allow these if you want to appear in AI answers.
| User agent | Operator and purpose | TFD call | Notes |
|---|---|---|---|
OAI-SearchBot |
OpenAI, ChatGPT Search index | Allow | Opting out drops you from ChatGPT search answers, though you can still appear as a plain link. Separate from GPTBot. |
ChatGPT-User |
OpenAI, user-initiated fetches | Allow | Fires when a ChatGPT user, Custom GPT, or action asks. May not follow robots.txt. |
PerplexityBot |
Perplexity, search index | Allow | Cites and links. Cloudflare (Aug 2025) reported additional undeclared Perplexity crawling. |
Perplexity-User |
Perplexity, user-initiated fetches | Allow | Perplexity's docs say it generally ignores robots.txt. |
Claude-SearchBot |
Anthropic, Claude search index | Allow | Separate from ClaudeBot (training). |
Claude-User |
Anthropic, user-initiated fetches | Allow | Fetches a page when a Claude user asks. Follows robots.txt. |
Meta-WebIndexer |
Meta, AI search index | Allow | Allowing it helps Meta AI cite you. Distinct from Meta's training crawler. |
Applebot |
Apple, Siri, Spotlight, Safari | Allow | Training is controlled separately (Applebot-Extended, below). |
Amzn-User |
Amazon, user-initiated fetches | Allow | Fires when an Amazon user or assistant asks. May not follow all robots.txt rules. Separate from Amazonbot. |
MistralAI-User / -Index |
Mistral, Vibe fetches and index | Allow | Behind Mistral's Vibe (formerly Le Chat). |
DuckAssistBot |
DuckDuckGo, AI answers | Allow | Cites sources. Not used to train AI models. |
YouBot |
You.com search | Allow | Powers You.com search. |
AI training, scraping, and extraction crawlers
Your call: block to protect your archive, allow if you want the exposure, or license where it pays. Blocking training does not remove you from AI search indexes, with one exception: Google-Extended also controls whether Gemini grounds answers in your content.
| User agent | Operator and purpose | TFD call | Notes |
|---|---|---|---|
GPTBot |
OpenAI, model training | Your call | Blocking does not affect ChatGPT Search visibility (that is OAI-SearchBot). |
ClaudeBot |
Anthropic, model training | Your call | Anthropic's current training crawler. It also honors robots.txt rules written for the retired "anthropic-ai" and "Claude-Web" tokens (Anthropic, 2024). |
CCBot |
Common Crawl | Your call | An open crawl reused by many AI trainers. Blocking stops future crawls only. |
Google-Extended |
Google, Gemini training and grounding | Your call | A robots.txt token, not a bot in your logs. Controls Gemini training and grounding in Gemini Apps. Google says it does not affect Search ranking or inclusion. |
GoogleOther |
Google, internal R&D | Your call | Not documented as a training crawler; Google says blocking it affects no product. |
Applebot-Extended |
Apple, Apple Intelligence training | Your call | A token, not a bot. Controls training only; you stay in Siri, Spotlight, and Safari. |
Meta-ExternalAgent |
Meta, AI training and content indexing | Your call | Meta's crawler for AI training and direct indexing. Not the same as facebookexternalhit, see Watch-outs. |
Meta-ExternalFetcher |
Meta, user-initiated fetch | Your call | Fetches pages when users ask. May bypass robots.txt. Meta also uses these fetches to improve its AI agents. |
Amazonbot |
Amazon, products/services and AI training | Your call | Respects robots.txt. A page-level noarchive meta tag opts that page out of training. |
Bytespider |
ByteDance / TikTok, reportedly AI training | Enforce at CDN | Widely reported to ignore robots.txt. Enforce at the CDN if it matters. |
Diffbot |
Diffbot, knowledge-graph and search extraction | Your call | Extraction, not foundation-model training. Respects robots.txt by default; its Crawlbot or Extract customers can override that. |
Webzio |
Webz.io, data collection and resale | Your call | Formerly "Omgilibot." Its "Webzio-extended" companion tags crawled data as usable or not for AI/ML training. |
img2dataset |
Open-source scraping tool | Enforce at CDN | A tool anyone can run, not a company bot. Ignores robots.txt and uses a generic browser user agent, so a name-based block misses default runs. By default it does skip images served with an X-Robots-Tag: noai or noimageai header, the one control you have here. |
Watch-outs: do not blanket-block these
The mistakes that cost publishers traffic they meant to keep.
| User agent | What it really is | TFD call | Notes |
|---|---|---|---|
facebookexternalhit |
Meta, link preview and unfurling | Allow | Builds the preview cards when your links are shared on Facebook, Instagram, and Messenger. That is its documented job, not AI training, so blocking it breaks those previews. May bypass robots.txt for security checks. |
Undeclared / stealth |
Crawlers that pose as ordinary browsers | Enforce at CDN | Some crawlers appear as normal browser traffic or ignore robots.txt. Cloudflare reported this of Perplexity in Aug 2025. robots.txt cannot reliably stop them; use CDN or bot-management rules. |
A starting robots.txt for the managed middle
If you land in the managed middle, this blocks the major training crawlers while leaving search and citation crawlers untouched. Anything not listed is allowed by default.
# Managed middle: block AI training, allow search and citation
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Applebot-Extended
Disallow: /
User-agent: Meta-ExternalAgent
Disallow: /
Two deliberate omissions:
- Google-Extended is left out because it also controls whether Gemini grounds answers in your content. Blocking it is not a clean training-only block, so make that choice on purpose.
- Bytespider and img2dataset are left out because robots.txt will not stop them. Handle those at the CDN.
Keep it current
AI crawlers and their robots.txt tokens change frequently. This reference is point-in-time as of September 2026. Confirm current details against each vendor's documentation before deploying, and re-audit roughly every 90 days.
Deciding who can crawl your site is only half of the job. The other half is making sure the AI systems you allow in get your story right. That is what a dedicated AI page is for.
Get the printable reference
Newsletter subscribers get the printable PDF. It's the full AI Crawler Reference in a format you can share with your editorial, product, and dev teams.
Need help?
Not sure what's crawling your site today? We audit server logs and robots.txt for publishers and set up CDN rules for the crawlers that don't follow them. Talk to us.
