web analytics
Back to blog

On 15 September, Cloudflare Picks Your AI Crawler Policy for You

Extractable summary

From 15 September 2026, Cloudflare will block AI crawlers used for model training and agent access by default on any page that carries advertising. It applies to new domains, new sites from existing customers, and every free-tier account that hasn’t changed its settings, but search crawling stays allowed. If your B2B site runs no ads, the new default does nothing to you at all, and your exposure sits in two other places: a legacy toggle that will start taking Googlebot down with it, and the fact that the third-party pages currently citing your category may be the ones going quiet.

What changes on 15 September

Cloudflare announced the change on 1 July 2026. From 15 September, the default configuration for AI traffic splits into three categories, and two of them get switched off on monetised pages. In Cloudflare’s own wording, “Training and Agent will be blocked by default on the pages that display ads”, while Search stays allowed (Cloudflare).

It hits all new domains onboarding to Cloudflare, any new site added by an existing customer, and every free-tier customer who hasn’t touched their settings by that date. Paying customers with existing zones keep what they’ve got, though the category controls are live now and worth opening regardless.

Opting out takes about a minute in zone Security settings, any time before the deadline. Cloudflare has said it’ll keep notifying customers as the date approaches, which is a reasonable promise and not one I’d build a plan around, because those emails go to whoever registered the domain four years ago.

Search, agent and training are now three separate questions

The old control was a single “Block AI Bots” toggle, which asked site owners a question that had stopped making sense, so the new taxonomy splits automated traffic by what a bot does with what it takes.

CategoryWhat it doesWhat you get back
SearchIndexes your pages so an engine can answer questions about them later.Referral traffic, or the AI-era version of it: citations.
AgentFetches a page in real time on a person’s behalf, with a human waiting.A visit, sometimes a conversion, no lasting index entry.
TrainingAbsorbs your content into a model’s weights.Nothing directly, ever.

Publishers forced this split because the three deals are wildly different and the old toggle priced them identically. The numbers got impossible to ignore: 52% of crawler requests were for AI training as of June 2026, up from 22% in spring 2025, and mixed-use crawlers now account for over a third of crawler activity (Cloudflare).

Cloudflare is also testing a fourth robots.txt signal through the Content Signals standard, expressing how a crawler may reuse what it takes: immediate, reference, or full.

The word doing all the work is “ads”

Most of the coverage frames this as the web putting up a turnstile, and for publishers that’s fair. For a B2B SaaS marketing site it’s close to a non-event, because the default only fires on pages that display advertising, and your product pages, blog and case studies almost certainly display none.

Ad-supported publisherB2B marketing site
How the page earnsHuman attention on the page.A named buyer entering a pipeline weeks later.
What a training crawl costsDirect revenue substitution.Bandwidth, and some pride.
What a citation is worthA fraction of a lost click.A qualified visitor who arrived pre-sold.
Correct default postureRestrict, then negotiate.Stay open, then measure.

Nobody reading the blog generates revenue by being there, which is exactly why blocking the crawlers that put you inside an answer is a self-inflicted wound. I’ve argued the same point from the other direction in Your Next Buyer Might Be a Bot: if agents are doing early-stage evaluation on a buyer’s behalf, an agent block is a lead block.

Should a B2B marketing site block AI crawlers?

For training crawlers on a marketing site, my answer is a soft yes if you’ve got proprietary research behind a form, and a firm no for everything else. For search and agent crawlers, blocking is close to indefensible unless you’re going dark on purpose.

The one place I’d take training blocking seriously is original research: a survey you ran, a benchmark you published, or a dataset nobody else has; that’s the only material on your site with real scarcity value, and it’s worth treating differently from the rest of the library. I made the broader case in Original Research Is the One Content Asset AI Can’t Copy.

The reflex to block AI crawlers came from publishers who lose money when a bot reads a page instead of a person. If you sell software, the loss runs the other way, arriving when nobody reads the page at all.

The Googlebot trap nobody’s putting in the headline

From 15 September, multi-purpose crawlers get evaluated against all their declared behaviours, and the most restrictive rule wins. Googlebot, Applebot and Bingbot all crawl for search and for AI features through a single user agent, so any site with Training blocked, including through the legacy “Block AI Bots” toggle, will find Googlebot blocked on ad-carrying pages too (Cloudflare docs).

Somebody at your company probably flipped that toggle in July 2025 during the first wave of coverage, told nobody, and left. If your site has advertising anywhere, including a sponsored placement on a resources page, that decision now has a search consequence attached to it.

Google is why this is happening at all: it accounts for roughly 88% of referral traffic and runs a mixed-use crawler, giving it about twice the information access of AI companies that separated their bots (Cloudflare). Matthew Prince framed the goal as wanting to “encourage mixed-use crawlers to separate out search from agent use and training” (Cloudflare press release), but that message is aimed at one company, and every site with a legacy toggle is standing in the blast radius.

Pay-per-crawl became pay-per-use, and it still isn’t a revenue line for you

The other half of the announcement replaces Pay-Per-Crawl with Pay-Per-Use, paying content owners when their work influences an AI answer instead of a bot fetching the page. Cloudflare’s reasoning is sound: “crawling is a crude measure of value”, since a page can be crawled once and cited a thousand times, or crawled endlessly and never used (Cloudflare).

For a publisher with a million-page archive, this is the start of a real revenue stream, but a B2B marketing site with 200 pages would get an annual cheque that wouldn’t cover a Semrush seat, and chasing it means restricting the access that gets you cited in the first place, so skip it.

The useful part is the reporting attached: participating sites get the queries that led to their content appearing, the page and snippet used, and an average result position. That’s the first native answer-engine equivalent of Search Console I’ve seen offered. Until it’s generally available, the manual method still works, and I wrote up mine in How I Track Whether AI Is Citing My Content.

What to check before September

Whether or not you intend to block AI crawlers, six things, in the order I’d do them:

  1. Find out whether you’re behind Cloudflare at all. A surprising number of marketing teams don’t know, because IT set it up. Check your DNS or ask.
  2. Open zone Security settings and read what’s currently set. Note whether the legacy AI-bots toggle is on and who turned it on.
  3. Confirm whether any page on your domain serves ads. Sponsored logos, a partner banner, a monetised resource hub, etc. If the answer is no, the new default can’t reach you and you can stop worrying about this.
  4. Decide Search, Agent and Training separately, in writing. Then put the decision somewhere a future colleague will find it.
  5. Pull your crawl volume. Cloudflare’s Attribution Business Insights dashboard shows which bots take your work and how little they send back. Most teams have never looked at this, and the number is usually bigger than the human one: bots crossed 50% of all internet traffic this year, and over half of good-bot crawl requests re-fetch pages that haven’t changed since the last visit (Cloudflare).
  6. Set a citation baseline before the date. If your visibility moves in October, you’ll want a September number to compare it against, and a page-level read of what’s actually extractable. That’s the job of an AI readiness audit.

That last one gets skipped and later regretted. When I built AEO/LLMO from zero at a B2B EdTech SaaS, what made the work defensible in a leadership meeting was having tracked visibility from the start rather than reconstructing it afterwards, which is the argument I make in The Metrics That Actually Tell You If Your Content Strategy Is Working.

Where this is irrelevant to you

If you’re not behind Cloudflare, nothing here forces you to block AI crawlers or stops you from doing it, and paying customers with an existing zone and no ads see no change on 15 September either.

The second-order effect is where it does reach you, and almost nobody is writing about that. The pages that currently supply AI answers in your category include forums, trade press and ad-supported publishers, and those are precisely the properties the new defaults are built for. As they restrict access, the pool of citable sources in your category thins out. That’s an argument for publishing more original material, and it’s the same logic underneath the AEO/LLMO framework I use, and the reason I keep arguing that structure without substance gets you nowhere.

The limits

I’m a content marketer, not an infrastructure engineer. I can read a bot policy and tell you what it does to a content strategy, and I’d want a sysadmin’s eyes on anything before it goes live.

Cloudflare has said it’s running tests and taking feedback until the deadline, so classification details may move, and Pay-Per-Use is explicitly an experiment with two partners. Treat the mechanics above as accurate as of writing and check the settings page yourself before acting on any of it.

I don’t know whether staying fully open will look smart in two years, and the economics of being crawled are being rewritten in public by a company that also sells the toll booth. I’m confident about the near-term call, because a B2B marketing site that blocks the crawlers feeding AI answers pays a real cost today against a benefit that’s still hypothetical.

Frequently asked questions

Only if your site is behind Cloudflare, carries advertising on at least some pages, and falls into one of the affected groups: new domains, new sites added by existing customers, or free-tier accounts that haven’t changed their settings. Training and agent crawlers get blocked on ad-carrying pages, while search crawling stays allowed. Sites with no advertising are unaffected, and any customer can opt out through zone Security settings before the date.

It can, from 15 September. Googlebot crawls for both search indexing and AI features through one user agent, and Cloudflare applies the most restrictive matching rule to multi-purpose crawlers, so a site with training blocked through the older “Block AI Bots” toggle will have Googlebot blocked on ad-carrying pages as a result. Applebot and Bingbot behave the same way. If you rely on organic search, check that setting before September.

For most B2B SaaS marketing sites, no. Your content exists to be found by buyers, and AI answers are a discovery channel rather than a threat to a paywall you don’t have. The exception worth weighing is proprietary material such as original survey data or a benchmark you produced, where scarcity has real value. Everything else earns more by being citable than by being protected.

How did this land?No reactions yet
Solange Rainha
Solange Rainha
Content Marketing Manager | 10+ Years B2B SaaS & AEO/LLMO