Every AI assistant that talks about your brand learned about it somewhere — from training data, from a live crawl, or from a fetch triggered by a user's question. Each path has its own crawler, and whether those crawlers can read your site is the most basic technical input to your AI visibility. Blocking them doesn't make AI stop talking about you; it makes AI talk about you from older, third-party information instead.
The crawlers worth knowing
The user agents you'll most often see in logs and robots.txt policies:
- GPTBot (OpenAI) — collects content that can be used for model training.
- OAI-SearchBot and ChatGPT-User (OpenAI) — power ChatGPT's search index and live page fetches when a user's question needs current information. These affect whether ChatGPT can cite you today.
- ClaudeBot and Claude-SearchBot (Anthropic) — Anthropic's crawl and search-index agents for Claude.
- PerplexityBot and Perplexity-User (Perplexity) — index and on-demand fetch for Perplexity's answer engine, which cites sources prominently.
- Google-Extended (Google) — a control token: blocking it opts your content out of Gemini training, without affecting Google Search rankings or AI Overviews.
- CCBot (Common Crawl) — a nonprofit crawl whose archives feed training sets for many models.
Training access vs. search access
The key distinction is between crawlers that gather training data (GPTBot, CCBot, the Google-Extended token) and crawlers that serve live answers (OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot). Blocking training bots is a legitimate content-policy choice with slow, diffuse effects. Blocking search-and-fetch bots has an immediate one: assistants lose the ability to read and cite your current pages, and your pricing, features, and positioning get represented by whatever third parties say about you.
For most brands that want to be found, the sensible default is to allow both groups — and at minimum the search group — then verify the policy is actually in effect. Check robots.txt, but also CDN bot-management rules and WAF settings, which often block these agents silently. Your server logs are the ground truth: if PerplexityBot gets a 403 on your docs, no robots.txt line will fix it.
A checklist to run this week
- Read your live robots.txt and list every AI user agent that's disallowed — then confirm each block is intentional.
- Grep recent server or CDN logs for the user agents above and look for 403/429 responses.
- Confirm your pricing, product, and docs pages return clean HTML to these agents (not a JavaScript shell or a challenge page).
- Keep your sitemap current so new pages get discovered quickly.
- Re-check after any CDN, firewall, or bot-protection change — this is where working policies quietly break.
Monitoring the result
Crawler access is an input; the output is whether answers improve. Cited tracks both sides — it flags AI crawlers blocked from your pages and shows whether your mention rate and citations move once access is fixed. Run your first report free and see where you stand.
See what AI says about your brand
Your first visibility report — mentions, sentiment, competitors, and citations across ChatGPT, Perplexity, Gemini, and AI Overviews — is free and ready in 24 hours.
Get your free report