# Cloudflare's AI Crawler Block: When It Stops Googlebot

> Since September 15, Cloudflare's Block setting for AI training also stops Googlebot. Who is affected, a five-minute check, and settings by site type.

- Source: https://www.citlyze.com/blog/cloudflare-ai-crawler-block-googlebot
- Author: Mara Ellison
- Published: 2026-10-02
- Tags: Crawlers, GEO

A Cloudflare AI crawler block stops Googlebot when your Training setting reads "Block" or "Block on pages with ads". Since September 15, 2026, Cloudflare applies both options to mixed-use crawlers (Googlebot, Bingbot and Applebot), so those pages stop getting crawled for search. The newer "Disallow AI Training" option keeps search crawling open.

Cloudflare says [more than 20% of web domains](https://blog.cloudflare.com/content-independence-day-ai-options/) sit behind its network, so a change to its defaults reaches a large share of the pages AI engines read. Check crawler access before anything else in an AI visibility plan. A rewrite cannot help a page that no crawler can fetch, and the setting lives in Cloudflare's security menu, far from the tools an SEO or content team opens each week.

## What changed on September 15, and who is affected by default

Two things changed: new domains now get preset AI bot policies, and Cloudflare's Block settings for training now apply to crawlers that also index for search.

The story has two dates, and trade coverage from July describes a plan that shifted before launch.

On July 1, Cloudflare split AI traffic into three categories (Search, Agent and Training) and made the controls available on every plan, including Free. Its [press release](https://www.cloudflare.com/press/press-releases/2026/cloudflare-allows-the-agentic-internet-to-flourish-with-a-simple-philosophy-your-content-your-rules/) said that on September 15, new customers, new sites and existing free customers who had not changed their settings would get defaults that allow search but block training and agent use on pages with ads. [Search Engine Journal](https://www.searchenginejournal.com/cloudflares-ai-crawler-rules-can-block-googlebot/581385/) flagged the consequence: a crawler that does both search and training would fall under the strictest rule, and Googlebot does both.

On September 15, Cloudflare shipped a revised version. Its [launch post](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/) added a fourth training option, "Disallow AI Training", which writes a no-training preference to robots.txt instead of blocking at the network edge. Google, Apple and Microsoft received an "Accountable" label for offering, or committing to, a training opt-out that leaves search untouched. The same post states the part that matters for your traffic: "Block and 'Block on pages with ads' now apply to mixed-use crawlers, including Applebot, Bingbot, and Googlebot, so either setting impacts search as well as training."

| Your situation on or after Sep 15 | What Cloudflare applies | Googlebot on those pages |
| --- | --- | --- |
| New domain, ad-supported preset | Search: Allow. Training: Disallow AI Training. Agent: Block on pages with ads | Allowed |
| New domain, no-ads preset | Search, Training and Agent: Allow | Allowed |
| Existing domain, legacy "Block AI bots" off | Migrated to Allow for all three | Allowed |
| Existing domain, legacy "Block AI bots" on (all pages or ad pages) | Migrated to Training: Disallow AI Training, Agent: Block on pages with ads | Allowed |
| Anyone who set Training to "Block" | Blocks training crawlers and mixed-use crawlers on all pages | Blocked |
| Anyone who set Training to "Block on pages with ads" | Same, limited to pages that show ads | Blocked on ad pages |

So Cloudflare does not block Googlebot by default. The risk sits with sites where someone chose one of the two Block options for training, for example between July 1 and September 15, when those were the only ways to stop training in the new controls.

A second default costs AI visibility on ad-supported sites: the preset blocks the Agent category on pages with ads. ChatGPT-User and Claude-User sit in that category, so when a reader asks ChatGPT or Claude to open one of your ad-carrying articles, the fetch fails.

Cloudflare's own documents leave one point open. The July announcement said existing free customers would receive the new defaults, while the September post describes a migration based on each domain's legacy setting and says nothing about free-plan zones. Check your own dashboard rather than trusting either summary.

## Which crawlers the setting touches, and what each one does

Cloudflare sorts AI bots into Search, Agent and Training, and Googlebot and Bingbot sit in two categories at once.

The table combines each operator's documentation with the category Cloudflare assigns on [Cloudflare Radar](https://radar.cloudflare.com/bots/directory), checked on September 26, 2026.

| Bot | Operator | What the operator says it does | Cloudflare category | What blocking it costs you |
| --- | --- | --- | --- | --- |
| GPTBot | OpenAI | Crawls content that "may be used in training" OpenAI's models ([OpenAI](https://developers.openai.com/api/docs/bots)) | Training | Your content leaves future training sets. OpenAI documents no effect on ChatGPT search |
| OAI-SearchBot | OpenAI | Surfaces websites in ChatGPT search | Search | OpenAI says opted-out sites "will not be shown in ChatGPT search answers" |
| ChatGPT-User | OpenAI | Visits pages when a user asks ChatGPT to; OpenAI says robots.txt "may not apply" | Agent | ChatGPT cannot open your page for a user who asks about it |
| ClaudeBot | Anthropic | Collects web content for model training ([Anthropic](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler)) | Training | Training exclusion only |
| Claude-SearchBot and Claude-User | Anthropic | Search indexing, and fetches for a user's question | Search and Agent | Anthropic says blocking either one reduces your visibility in Claude's search results |
| PerplexityBot | Perplexity | Surfaces and links sites in Perplexity search; "not used to crawl content for AI foundation models" ([Perplexity](https://docs.perplexity.ai/guides/bots)) | Not listed as a verified bot | You drop out of Perplexity's results |
| Google-Extended | Google | A robots.txt token with no crawler of its own. It controls Gemini training and grounding in Gemini Apps and Vertex AI ([Google](https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers)) | None (no user agent) | Google says it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal" |
| Googlebot | Google | Builds Google's search index | Search and Training | Pages that return 4xx are "removed from the index" ([Google](https://developers.google.com/crawling/docs/troubleshooting/http-status-codes)) |

Two rows answer questions that come up in most crawler discussions. Blocking GPTBot does not remove you from ChatGPT: GPTBot collects training data, OAI-SearchBot decides whether your pages appear in ChatGPT search answers, and ChatGPT-User fetches pages when a user asks. And blocking Google-Extended does not touch your Google rankings: it opts you out of grounding in the Gemini app and Vertex AI, while AI Overviews and AI Mode keep drawing on the regular search index. If you track [how Gemini mentions your brand](https://www.citlyze.com/blog/gemini-rank-tracking), treat that token as a lever on the result.

PerplexityBot needs a separate note. It did not appear in Radar's verified bot directory when we checked, and Cloudflare [de-listed Perplexity as a verified bot](https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/) in August 2025 after documenting undeclared crawling. Whether it reaches your pages depends on your wider bot rules as much as on the AI category toggles, so confirm it in your security events rather than assuming.

## One decision per bot type

Make one decision per category, write it down, and match the settings to it.

- Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot, Bingbot): allow them if you want to appear in AI answers or search at all. Cloudflare reports that [fewer than 1% of its sites block Search bots](https://blog.cloudflare.com/accountable-mixed-use-ai-crawlers/).
- User-triggered agents (ChatGPT-User, Claude-User, Perplexity-User): these fetch a page because a person asked about it. Blocking them stops the engine from reading the page a buyer pointed it to. Robots.txt alone will not hold them back either, since OpenAI and Perplexity say their user-triggered fetchers may not follow it; enforcement for those needs an edge rule such as Cloudflare's Agent setting.
- Training crawlers (GPTBot, ClaudeBot, the Google-Extended token): a policy choice about future models. Cloudflare says 17% of its sites enable some mechanism to block training. If you choose that, use Disallow AI Training, which leaves Googlebot's search crawl alone.

The [AI crawlers guide](https://www.citlyze.com/blog/ai-crawlers-guide) covers the wider bot list, including crawlers Cloudflare does not classify.

## The five-minute check

Read the setting, read the robots.txt it produces, then confirm with logs from both sides.

1. In the Cloudflare dashboard, open your domain, then Security, then Settings, and find the AI bot policies. Write down the Search, Training and Agent values. If Training reads "Block" or "Block on pages with ads", Googlebot, Bingbot and Applebot are blocked on those pages.
2. Check whether Preference Sync is on, then open `yourdomain.com/robots.txt`. With sync on, Cloudflare writes your AI preferences into that file. Look for the lines it added and for any older `Disallow` rules that contradict them.
3. In Security, open Events and filter to the last seven days for Googlebot and for verified bots with a block or challenge action. A blocked Googlebot request on a page you want indexed means the Training setting, or another firewall rule, is stopping search crawling.
4. In Google Search Console, open Settings, then Crawl stats, and read the breakdown by response. A jump in 403 responses after September 15 points at the edge. Then run URL Inspection with a live test on one page that carries ads.
5. Ask ChatGPT to open and summarize one of your ad-carrying pages. If it reports that it cannot access the page, the Agent block is in effect for user-triggered fetches.

A robots.txt reader cannot see a Cloudflare network block, because CDNs verify real crawlers by IP address, not by the user-agent string a checker could send. Steps 3 and 4 cover that gap.

For the robots.txt side, the free [AI crawler access checker](https://www.citlyze.com/free-tools/ai-crawler-access-checker) reads your file the way crawlers do and reports a verdict and the deciding rule for 44 documented AI crawlers, grouped by what each block costs. No account needed.

## Settings for a brand site, a publisher and an online store

None of the three should pick "Block" for training if Google traffic matters to them. The table is a starting point; the right answer depends on how you make money from the page.

| Site type | Search | Training | Agent | Reasoning |
| --- | --- | --- | --- | --- |
| Brand or B2B site | Allow | Allow, or Disallow AI Training if your policy requires it | Allow | Your pages exist to be found, quoted and opened by buyers' assistants |
| Ad-supported publisher | Allow | Disallow AI Training | Decide per section. Cloudflare's preset blocks agents on pages with ads | Blocking agents protects ad impressions but removes those pages from answers built on a user's request |
| E-commerce store | Allow | Allow, or Disallow AI Training | Allow | Shopping assistants fetch product pages on a shopper's behalf; a blocked fetch loses that shopper |

Publishers carry the hardest trade-off. An agent visit shows no ads, and an agent block also means ChatGPT or Claude cannot read the article a reader asked about, which removes the citation and the link back. Decide it section by section, and leave a note in your team's settings log explaining why.

If you maintain robots.txt by hand next to Cloudflare, keep the two consistent. A preset such as "block training, allow search" in the file should match the dashboard, or the stricter of the two wins and nobody notices.

## Confirming recovery after you change the setting

Watch Googlebot's responses return to 200, then watch indexing and AI citations catch up over the following weeks.

- At the edge: blocked Googlebot events in Cloudflare should stop once the new setting is live.
- In Search Console: Crawl stats should show 403s falling and 200s rising. Google removes indexed URLs that return 4xx, and it publishes no timeline for bringing them back, so each page has to be recrawled first. Request indexing for your most important URLs to put them at the front of the queue.
- In your logs or crawler analytics: GPTBot, OAI-SearchBot and ClaudeBot visits should reappear on the pages you unblocked.
- In AI answers: track the prompts those pages used to earn citations for. A crawl is an input, the citation is the outcome, and the two can lag each other, as the [crawl vs citation](https://www.citlyze.com/blog/crawl-vs-citation) piece explains.

Add the check to your regular [GEO audit](https://www.citlyze.com/blog/geo-audit-guide), because firewall rules and CDN defaults drift.

## Start with the two checks no outside tool can run

Spend five minutes in the Cloudflare dashboard on steps 1 and 3 above: read the three AI policy values, then look for blocked Googlebot events. Then run the [AI crawler access checker](https://www.citlyze.com/free-tools/ai-crawler-access-checker) against your homepage and one article with ads, and rewrite any robots.txt rule you did not mean to publish with the [AI robots.txt generator](https://www.citlyze.com/free-tools/ai-robots-txt-generator). Neither tool needs an account.

If you want the check to run on its own, Citlyze's [bot access monitoring](https://www.citlyze.com/features/bot-access-monitoring) re-tests each crawler against your key pages and names the layer that blocked it, next to the citations those pages earn. Plans start at $29 a month on [the pricing page](https://www.citlyze.com/pricing), with a 14-day trial and no card.

---

Canonical: https://www.citlyze.com/blog/cloudflare-ai-crawler-block-googlebot
