Cloudflare’s Crawler Classification Puts Publishers Between Discovery and Training

Cloudflare’s new purpose-based crawler defaults expose a structural impossibility: you can permit a crawler, but you cannot permit its intentions.

Key Context — September 15, 2026

July 2025 — Pay-Per-Crawl Marketplace

Cloudflare launches a marketplace allowing publishers to charge AI crawlers for access to their content; no current pricing details are public.

July 1, 2026 — Cloudflare Traffic Report

Cloudflare reports that mixed-use crawlers — blending search, agent use, and training — account for more than 36% of crawler activity across its network.

September 15, 2026 — New Default Goes Live

On pages displaying advertising, Cloudflare’s new default blocks Training and Agent crawlers while allowing Search crawlers — applied to all new domains onboarding to Cloudflare, with an opt-out available through Security settings.

What Happened

Effective today, Cloudflare has introduced a purpose-based crawler classification system and attached a new default to it. The company defines three categories in its own words: Search is “any behavior that collects or indexes your content, so it can answer questions about it later”; Agent is “automated behavior that is acting, usually in real time, on a person’s behalf, to get something done right now”; and Training is “a crawler taking your content to train or fine-tune a model.” On pages that display advertising, the default as of September 15, 2026, blocks Training and Agent crawlers while leaving Search crawlers permitted.

Cloudflare’s own post scopes this change to all new domains onboarding to Cloudflare. Some trade coverage has reported a wider application — to new customers and existing free-tier customers — though that broader reach has not been confirmed by Cloudflare directly and should be treated as reported rather than established. Site owners who wanted no change to the treatment of Training crawlers that also crawl for Search purposes could opt out through their Security settings ahead of the deadline.

The complication Cloudflare names directly is the one that matters most analytically. Googlebot, Applebot, and BingBot are identified as multi-purpose crawlers that combine Search with Training. Under the new default, if a publisher blocks Training, such a crawler is, in Cloudflare’s framing, “blocked according to all of their behaviors” — which means the search function is blocked alongside the training function. Cloudflare’s own July 2026 data showed that mixed-use crawlers blending search, agent use, and training represented more than 36% of crawler activity; that figure is a floor, not a precise count of a static population.

The Numbers That Frame the Default

>36%

Mixed-use crawler share (search + agent + training), per Cloudflare July 2026

3

Crawler purposes Cloudflare now classifies: Search, Agent, Training

1

Real choice for publishers on mixed-use crawlers: found or trained on — not both

~1 yr

Age of Cloudflare’s Pay-Per-Crawl marketplace, running alongside the new default

The key insight: A permission system organised by purpose cannot bind an infrastructure that has no notion of purpose. At the moment of the fetch, what a crawler intends to do with the bytes it collects is unobservable. That is not a flaw in Cloudflare’s implementation — it is a property of the layer the implementation has to sit on.

Purpose-based permissions assume you can allow a crawler's search function and refuse its training function. F
Purpose-based permissions assume you can allow a crawler’s search function and refuse its training function. For more than a third of crawler activity, that choice does not exist — which is why the publisher’s real decision is between being found and being trained on.

The Structural Read

The taxonomic work Cloudflare has done here is genuinely useful. Naming crawlers by their downstream function — not by the company that operates them — is a more honest framing than the bot-allowlist model that preceded it. But the taxonomy runs into a protocol reality the moment it meets a mixed-use crawler: purpose lives in what the operator subsequently does with the bytes, somewhere else entirely. A publisher can permit a crawler. A publisher cannot permit a crawler’s intentions. For the more than 36% of traffic that is already mixed-use, “allow search, refuse training” is simply not an available option, however sensible it sounds when stated as a preference. The real choice is between being found and being trained on — and it is one choice, not two.

Everything that flows from this is what Cloudflare’s system, operating honestly at the network layer, has to treat as an honour system dressed as a control. The operator on the other side of the fetch decides how the bytes are used. The publisher set a preference. The two things happen at different layers, at different times, with no enforcement mechanism connecting them for mixed-use traffic.

Permission Layer — Business Engineer Framework

The Permission Layer framework holds that whoever controls access conditions — defaults, settings, compliance requirements — exercises effective authority over which capabilities propagate and which do not. Cloudflare’s new default is a textbook instance: a settings change applied to newly onboarding domains shapes training-data supply for a slice of the web through inertia rather than through any publisher’s active decision. No legislature was involved at any stage.

The second structural observation concerns bundling, and it needs to be stated carefully. Operators that run a general search index and a frontier model behind a single crawler identity have made refusal expensive in a way no independent AI company can match. An AI-only crawler presents a publisher with a clean decision carrying a bounded cost: refuse it, and nothing visible is lost. A crawler that also carries discovery makes refusal cost traffic — which for an advertising-supported publisher means revenue. The asymmetry is entirely about who else is in the bundle.

This is a structural consequence of operating both functions from a single crawler identity, not an allegation about anybody’s intentions. Nothing here suggests that Google, Apple, or Microsoft combined these behaviours in order to deter refusal. Nothing here suggests conduct that is improper, anticompetitive, or unlawful. The observation is about the position those firms occupy, not the motive behind it.

The third observation is about where Cloudflare chose to draw the line, and it is the most revealing design decision of the three. The default does not apply to all content. It applies to pages carrying ads. That locates the boundary where the economic injury is — not where the copyright argument is. It converts a slow, abstract, and genuinely unresolved intellectual-property dispute into a concrete revenue-protection feature a publisher can switch on this afternoon. The questions of whether training on published content is lawful, what fair use covers, and how copyright applies to AI training data are contested and remain unsettled; nothing here states any legal conclusion about them. But the design implicitly concedes something worth noticing: content that does not monetise through human attention is not what this particular default defends. This is an advertising problem being solved, not an authorship one.

Search Crawlers (Googlebot, Applebot, BingBot)

BUNDLED

Permitted under the new default — but on mixed-use crawlers, blocking training blocks search too. The bundle is indivisible at the fetch layer.

Training Crawlers

BLOCKED (DEFAULT)

Blocked by default on ad-bearing pages for new domains onboarding to Cloudflare. Opt-out available through Security settings.

Agent Crawlers

BLOCKED (DEFAULT)

Also blocked by default on ad-bearing pages. Defined by Cloudflare as real-time, task-directed automated behaviour acting on a person’s behalf.

Three Implications

IMPLICATION 1 — DEFAULTS ARE POLICY, NOT SETTINGS

Most site owners never change a default. By applying this one to newly onboarding domains, Cloudflare shapes training-data supply for a portion of the web through inertia rather than through any publisher’s deliberate choice. A private infrastructure company is exercising effective authority over what AI models can learn from — without a legislative process, a regulatory comment period, or a vote. Set that beside an antitrust accommodation refused within a day of being sought, and an audit-and-insurance standard for AI agents now spreading through enterprise procurement checklists: whoever controls a default, a procurement checklist, or a disclosure obligation moves faster than whoever requires a statute. This piece makes no characterisation of that dynamic as desirable or otherwise, and nothing here suggests Cloudflare intends to set policy for anyone. It is an observation about where effective authority over this technology currently sits.

IMPLICATION 2 — BUNDLING DISCOVERY WITH TRAINING IS A STRUCTURAL MOAT

Independent AI companies operating training-only crawlers face a clean refusal — publishers can block them with no visible cost to search ranking or referral traffic. Operators running a search index and a frontier model behind one crawler identity face a different dynamic entirely: the cost of refusal is traffic, which for ad-supported publishers is revenue. That asymmetry requires no allegation of intent to be structurally significant. It is a property of operating both functions, and it is durable for as long as those functions remain bundled behind a single crawler identity.

IMPLICATION 3 — THE PAY-PER-CRAWL MARKETPLACE REFRAMES THE NEGOTIATION

Cloudflare’s Pay-Per-Crawl marketplace, operating for roughly a year alongside the new default, offers a third path beyond binary allow-or-block. The new default gives publishers a stronger baseline position from which to approach that marketplace — the economic logic of a blocked Training crawler provides leverage that a world of unrestricted crawling did not. No pricing details are public and no claim is made here about how publishers or AI operators will respond to the combination. The design logic, however, is coherent: establish a default that creates sc

91,000+ executives read Business Engineer for the AI strategy frameworks cited by ChatGPT, Claude, and Perplexity.

Cloudflare’s own description scopes the new default to domains newly onboarding to Cloudflare. Some trade coverage reports a wider application, including to new customers and existing free-tier customers; that wider scope is reported rather than confirmed here, and nothing in this article should be read as saying the default applies to all Cloudflare customers, to all existing sites, or to the web generally. Site owners were able to opt out through their Security settings ahead of the deadline. The figure for mixed-use crawlers is Cloudflare’s own, published on 1 July 2026, and is stated as more than 36% of activity, covering crawlers that blend search, agent use and training; 36% should therefore be read as a floor rather than a precise estimate. The observation that operating a search index and a model behind a single crawler identity makes refusal costly is a structural one about the position those firms occupy. Nothing here alleges that Google, Apple, Microsoft or anyone else combined these behaviours in order to deter refusal, and nothing here suggests conduct that is improper, anticompetitive or unlawful. This article states no legal conclusion about copyright, fair use, or whether training a model on published content is lawful. Those questions are contested and remain unsettled, and nothing here is legal advice. No pricing or compensation details for Cloudflare’s Pay-Per-Crawl marketplace are reported or asserted. No prediction is made about how many publishers will block, about traffic or revenue effects for anyone, or about how AI companies will respond. Cloudflare, Alphabet, Apple and Microsoft are publicly listed companies. No claim is made about any share price or market effect. This is business analysis, not investment advice, no view is expressed on any security, and no recommendation is made.

Sources: blog.cloudflare.com · blog.cloudflare.com · techcrunch.com

Scroll to Top

Discover more from FourWeekMBA

Subscribe now to keep reading and get access to the full archive.

Continue reading

FourWeekMBA