Block AI Crawlers Without Vanishing From AI Answers
Plenty of businesses have chosen to block AI crawlers, and plenty more have done it without choosing at all. A CDN tick box, an over-keen security plugin or a host’s server-wide bot list can all stop ChatGPT, Claude, Perplexity or Gemini reading your website, and then your business isn’t in the answer when a customer asks an AI assistant for a recommendation. On 15 September 2026 Cloudflare changed its defaults for AI crawlers, which makes this a good week to check what your own site is really doing.
What changed at Cloudflare on 15 September 2026?
From 15 September 2026, new domains added to Cloudflare block AI “Training” and “Agent” bots by default on pages that display ads, while AI “Search” bots stay allowed. The change was announced on 1 July 2026 in Cloudflare’s Content Independence Day update, which replaced the single “Block AI bots” switch with three separate categories that every customer, including those on the Free plan, can now set for themselves.
Cloudflare defines the three categories like this:
- Training. A crawler taking your content to train or fine-tune an AI model.
- Search. A crawler that collects or indexes your content so it can answer questions about it later.
- Agent. Software acting in real time on a person’s behalf, for example fetching your page because someone just asked an assistant about you.
Cloudflare’s post says the new defaults apply to new domains; PYMNTS reported that they also reach new sites added by existing customers and free-tier accounts. There is a second detail that matters to far more sites. Crawlers that do more than one job are now judged by the strictest of their behaviours, so Cloudflare says multi-purpose crawlers such as Googlebot, Applebot and Bingbot will be blocked for customers who have chosen to block Training, whether through the new options or the old “Block AI bots” setting. If anyone switched that on and forgot, look again.
Training crawlers, search crawlers and user agents: what’s the difference?
Each major AI company runs separate bots for training, for search and for fetching pages when a user asks, and blocking one does not block the others. Here is how the main vendors describe theirs:
- OpenAI. GPTBot collects content that may be used to train its models. OAI-SearchBot surfaces websites in ChatGPT’s search features. ChatGPT-User fetches a page when a person asks ChatGPT to do something.
- Anthropic. ClaudeBot collects content that could contribute to training. Claude-SearchBot improves Claude’s search results. Claude-User fetches pages when someone asks Claude a question. Anthropic warns that disabling either of the last two “may reduce your site’s visibility”.
- Perplexity. PerplexityBot surfaces and links websites in Perplexity’s results, and Perplexity-User fetches pages for a user’s question. Perplexity says neither is used to train foundation models.
- Google. Google-Extended is a robots.txt token that controls whether your content can be used for Gemini training and grounding. Google states that it does not affect inclusion or ranking in Google Search.
Is it ever right to block AI crawlers?
Yes: blocking training crawlers is a legitimate business decision, but blocking search crawlers and user agents takes you out of AI answers. If your content is your product (original research, paid courses, journalism), you may reasonably decide it shouldn’t feed someone else’s model. That costs little visibility, because training isn’t how assistants find and cite you today.
Search crawlers and user agents are different. They are how an assistant looks you up when a customer asks “who’s a good accountant in Stockport?” or “does this company offer same-day repairs?”. Block them and the assistant answers from other people’s pages, or not at all. For most service businesses the sensible default is to allow search and user agents, then make a deliberate call on training. Our AI search visibility service starts with exactly this audit.
What we found on sites we look after
The most common cause of AI invisibility we see isn’t a policy decision at all: it’s a layer of the stack that nobody realised was saying no. Three examples from the last few months:
- A security plugin refusing the real ClaudeBot. On one site, Cloudflare was set to allow every AI bot and robots.txt allowed everyone, yet around 96% of roughly 200 visits from ClaudeBot over a day and a half were refused with a 403 error. Every one of those visits came from Anthropic’s published IP ranges, so they were genuine. The refusal was coming from the WordPress security plugin’s firewall on the server, which nothing in the CDN dashboard would ever show you. Googlebot got through fine, so every ordinary check looked healthy.
- A hosting platform’s server-wide bot list. On other sites, the host’s own server configuration was turning away anything that called itself ClaudeBot, before WordPress was even involved. No plugin or theme change would fix that.
- A CDN setting rewriting robots.txt. On another, a single “managed robots.txt” option at the CDN had quietly added rules telling the training bots to stay away. The robots.txt file inside WordPress looked clean; the one the world actually saw did not.
Why your raw logs can’t be trusted on their own
Anyone can put “ClaudeBot” or “GPTBot” in a request, so a log line with an AI bot’s name proves nothing until you check where it came from. On the same sites we saw scanners on ordinary cloud servers cycling through every well-known AI bot name while probing for configuration files and passwords. Blocking those is correct. But count them as AI visits and it’s easy to wave every AI-bot 403 away as “only scanners”. In the case above, scanners and the genuine ClaudeBot were both being refused, and only one of those was a problem.
The fix is to verify by address. OpenAI, Anthropic and Perplexity each publish the IP ranges their bots use, linked from the documentation pages above. A request is only really from ClaudeBot or OAI-SearchBot if it comes from one of those ranges. Good web application firewall rules work the same way: they let verified bots through and stop impostors using their names.
A practical checklist to stay visible in AI answers
Work from the outside in, because each layer can overrule the one behind it.
- Read your live robots.txt. Open yourdomain/robots.txt in a browser, not the file in your CMS, and look for Disallow rules against OAI-SearchBot, ChatGPT-User, Claude-SearchBot, Claude-User or PerplexityBot.
- Check your CDN’s bot settings. If you use Cloudflare, look at the AI crawler options in your Security settings, including whether Training is blocked and what that now means for Googlebot and Bingbot.
- Check your security plugin and server firewall. Look for bad-bot lists, rate limits or “block AI” options, and ask your host whether they block bots at server level.
- Test against the real IP ranges. Compare refused requests in your logs with each vendor’s published list. A 403 from a verified address is a real block; a 403 from anywhere else is usually an impostor being stopped.
- Keep watching. Plugin updates, host changes and CDN defaults move without warning. Waggle watches every crawler visit at the edge, so a new block shows up in days rather than months.
And once AI assistants start sending people your way, Scout shows which companies arrived and what they looked at.
Frequently asked questions
Will blocking GPTBot or ClaudeBot remove me from ChatGPT or Claude answers?
Not by itself. Both are training crawlers. ChatGPT and Claude use separate bots (OAI-SearchBot and ChatGPT-User; Claude-SearchBot and Claude-User) to find and fetch pages for answers, and blocking one bot does not block the others.
Does the Cloudflare change affect my existing website?
Cloudflare’s post says the new defaults apply to new domains. However, if you ever chose to block Training, including with the older “Block AI bots” switch, Cloudflare says multi-purpose crawlers such as Googlebot and Bingbot are now blocked too, so existing settings are worth reviewing.
How can I tell if a visit really came from an AI company?
Check the visitor’s IP address against the ranges the company publishes. The name in the request (the user agent) is easy to fake; the address is not.
Not sure what your own site is telling AI assistants? Book a call and we’ll check your robots.txt, CDN, firewall and hosting layer with you, and show you which AI crawlers are actually getting through.
Has something in this article peaked your interest? We’re never more than a few clicks or a quick call away so please don’t hesitate to get in touch!



