Back to blog
Web Scraping·October 7, 2026·13 min read

Cloudflare Now Blocks AI Agents by Default: What Builders Need to Know

Cloudflare now blocks AI agent traffic by default on ad-supported pages for new domains and free-tier sites. See the traffic data driving the change, the Perplexity stealth-crawling case, the IETF standard forming around it, and how to tell a real block from a misclassification before building against a domain.

M

Mahendra Sreekumar

Anakin Team

Cloudflare's new default blocking AI Agent and Training bots behind a wall on ad-supported pages, while Search crawlers pass through, effective September 15, 2026

Introduction

A RAG pipeline or research agent that fetched a page cleanly last month can now come back with nothing, on the same domain, with the same code. The reason is not a bug. Cloudflare changed what happens by default when an automated agent, not a search crawler, requests a page that carries ads, and that change landed in the middle of the fastest growth in AI bot traffic the web has measured.

This covers exactly what changed and when, why Cloudflare felt forced to make it the default, how the wider industry (IETF, Anthropic, OpenAI, and Cloudflare's own Pay Per Crawl program) is converging on the same idea, what a block actually looks like from an agent's side, and how to tell a legitimate access problem from a site's deliberate opt-out before writing a line of fetch logic against it.

Anyone pulling pages through Anakin's URL Scraper, or any comparable fetch layer, hits this exact wall first. The diagnosis below is written from that seat.

What changed, and when

Cloudflare's own documentation states the new default plainly: bots classified as Training or as Agent are blocked on pages that display ads, and Search remains allowed. That took effect September 15, 2026.

  • New Cloudflare customers get these defaults from day one.
  • New domains added by existing customers get these defaults, even if the account is years old.
  • Existing customers on the Free tier who never touched their bot settings were switched over automatically.
  • Existing paid customers who already had a domain live before September 15 keep whatever they had configured. Nothing changed under them without action.

That last point matters more than the headlines suggest. This is not Cloudflare blocking agents across the entire web overnight. It is a default that applies going forward, to new and newly onboarded properties, and to the free tier specifically.

Why Cloudflare moved: the traffic numbers behind the policy

This default did not appear in a vacuum. It is a response to a measurable shift in what is actually hitting origin servers.

  • Imperva's 2026 Bad Bot Report found automated traffic now accounts for more than 53% of all web traffic, up from 51% the year before, with human activity continuing to decline. The report frames AI agents as a distinct new category: systems that retrieve data, execute workflows, and act on a user's behalf rather than just scanning pages.
  • Fastly's network data shows AI requests grew roughly 30% between January and May 2026, about 6.5x faster than human traffic in the same window, with more than 51% of AI requests requiring a trip back to origin infrastructure versus under 9% for human requests. AI traffic hits the expensive part of the stack far more often per request.

An agent that cannot be cached the way a browsing human can, multiplied across a 6.5x growth curve, is an infrastructure cost problem before it is a content-licensing problem. The default targets exactly the traffic pattern driving that cost. Every call Anakin's URL Scraper makes on an agent's behalf fits that same profile: a fresh, uncached trip to origin each time, exactly the pattern this default was built to catch.

Bar stat comparison: 53 percent of all web traffic is automated per Imperva's 2026 Bad Bot Report, and AI traffic growing 6.5x faster than human traffic per Fastly's 2026 network data

The three categories Cloudflare actually checks

Cloudflare's bot classification splits automated traffic into named behaviors, and the new default only touches two of them.

  • Search - crawlers that index content to answer questions later, the kind that sends referral traffic back. Allowed by default.
  • Agent - automated activity acting in real time on a person's behalf: chat-fetch bots, browser-use agents, anything retrieving a page to answer one specific request right now. Blocked by default on ad pages.
  • Training - crawlers harvesting content to train or fine-tune a model. Blocked by default on ad pages.

A RAG pipeline that fetches a page at query time to ground an answer, whether through Anakin's URL Scraper or a hand-rolled fetch, is Agent traffic by Cloudflare's own definition, not Search. That is the category this change targets.

Why ad-supported pages specifically

Cloudflare's stated logic: an ad on a page is evidence that page was built for a person to look at it. An agent that fetches the page, reads the price or the review, and answers a user who never loads the page themselves never sees that ad. Search crawlers get a pass because they send a visitor back eventually. Agent and Training traffic, by Cloudflare's framing, consume the content without returning anything the ad model depends on.

Pages with no ads on the same domain are not covered by this default at all. A news site's article pages might block agents while its documentation pages do not.

Diagram of Search, Agent, and Training bot request lanes hitting a web page with an ad unit, Search passing through while Agent and Training traffic are blocked by Cloudflare's default

The catch for mixed-purpose crawlers

Googlebot, Bingbot, and Applebot each crawl for Search and for their own AI training under a single bot identity. Blocking Training by default blocks that same bot on ad pages too, unless the site owner explicitly carves out an exception. A domain that adopted the new defaults without reviewing them can end up suppressing standard search indexing it never meant to touch, on exactly the pages that monetize the site.

The real compliance problem robots.txt alone can't fix

A 2026 study out of Duke University ran a controlled, large-scale test of 130 self-declared bots over 40 days and found bots grow less compliant as robots.txt directives get stricter. The researchers' conclusion is blunt: certain categories of bots, including AI search crawlers, rarely check robots.txt at all. Relying on a voluntary text file as the only control was already shaky before this policy existed.

The clearest real-world example of that risk played out between Cloudflare and Perplexity in 2025.

Cloudflare's own investigation found that Perplexity is repeatedly modifying their user agent and changing their source ASNs to hide their crawling activity, as well as ignoring, or sometimes failing to even fetch, robots.txt files. On test domains Cloudflare built specifically to be unreachable by any declared crawler, Perplexity's assistant still returned answers describing the page content. Cloudflare responded by de-listing Perplexity as a verified bot and adding detection heuristics for the stealth pattern. Perplexity disputed the findings publicly.

Whichever side of that specific dispute turns out to be right, the incident is why Cloudflare, Anthropic, and OpenAI have all converged on one position: a declared, honest user agent that respects robots.txt is the baseline expectation now, not a courtesy.

The wider shift: standardizing how sites say yes or no

Cloudflare's Pay Per Crawl program repurposes the long-unused HTTP 402 Payment Required status code: a site owner sets a per-crawl price, and a crawler that doesn't pay gets a 402 instead of the page. Site owners choose, per crawler, to allow for free, charge, or block outright.

Cloudflare's Content Signals Policy works one layer up, as a declared-preference extension to robots.txt itself:

curl https://example.com/robots.txt

# User-Agent: *
# Content-Signal: search=yes, ai-train=no
# Allow: /
  • search - permission to index and return excerpts in search results.
  • ai-input - permission to use the content as live input for a generated answer, the exact thing a RAG fetch or agent fetch does.
  • ai-train - permission to use the content in model training or fine-tuning.

Cloudflare's managed robots.txt already sets search=yes, ai-train=no across more than 3.8 million domains by default. None of this is Cloudflare acting alone, either.

The IETF's AI Preferences working group (AIPREF) is drafting a standard vocabulary for exactly this kind of signal, with a Content-Usage HTTP header and a matching robots.txt directive, meant to work the same way across every site and every crawler rather than per-vendor. That standardization effort is the direction this is all heading: fewer one-off vendor rules, one shared way for a site to state what it allows.

Know the crawler before building against the domain

Each major AI company runs separate, named bots for separate jobs, documented directly by the companies that operate them.

OperatorBot nameJobRespects robots.txt
OpenAIGPTBotCrawls content for model trainingYes
OpenAIOAI-SearchBotSurfaces sites in ChatGPT searchYes
OpenAIChatGPT-UserFetches a page a user asked about, liveNot necessarily - user-initiated
AnthropicClaudeBotCrawls content for model trainingYes
AnthropicClaude-UserFetches a specific page a user asked Claude aboutYes
AnthropicClaude-SearchBotImproves Claude search result qualityYes
PerplexityPerplexityBot / Perplexity-UserDeclared crawler and user-fetch agentDisputed - see Cloudflare's findings above
GoogleGooglebot / Google-ExtendedSearch indexing / Gemini trainingYes

A site that disallows ClaudeBot but not Claude-User, for instance, is drawing exactly the Agent-versus-Training line Cloudflare's own default enforces. Checking which named bot a request will actually present as, before assuming a domain-wide block, is worth doing before writing retry logic that treats every one of these the same.

Anakin's URL Scraper identifies itself with a declared, checkable user agent for the same reason every bot in that table does: an honest identity is what keeps a request eligible for Search-style treatment instead of defaulting into Agent traffic.

Table of named AI crawlers including OpenAI's GPTBot, Anthropic's ClaudeBot, Perplexity's PerplexityBot, Google's Googlebot, and Anakin's URL Scraper, with their role and robots.txt compliance status

What this looks like from the agent's side

Not a clean 403 on every request to a domain. It shows up as inconsistent: a product page with ad units returns a block or a challenge, while a docs page with no monetization on the same domain returns content normally. A scraper tuned to treat any non-200 as a retry-and-move-on case will often get this wrong, since it looks like a transient rate limit rather than a deliberate policy.

The fastest way to confirm it is the Agent-block default, not a rate limit or a one-off fingerprint flag, is to check the same URL from a plain Search-style request pattern versus a request that behaves like a live fetch-and-answer agent. If only the second consistently fails on ad-carrying pages, this is the default at work.

On Anakin's URL Scraper, that pattern shows up as a blocked or challenge response on ad-carrying pages specifically, with the same domain's non-ad pages returning clean markdown, no code change required to tell the two apart.

Check the site's declared preference before building against it

Content Signals in robots.txt are voluntary on the crawler's side; nothing stops a request at the network layer, and the Duke study above is a direct warning about trusting that file alone. But checking it first is still a five-second request that tells a builder whether a block seen later is policy or something else entirely, like a plain rate limit or fingerprint mismatch, before any pipeline gets built against a domain that has already said no.

Where Anakin's URL Scraper fits

None of this changes what Anakin's URL Scraper is built for: turning a URL into clean markdown, HTML, or structured JSON, for data access a team is authorized to pull. The far more common problem most pipelines actually hit is not an explicit, deliberate block. It is a plain HTTP client getting misclassified as automated traffic on a domain that has not restricted agent access at all, which looks identical to a real block until it is actually tested.

URL Scraper's useBrowser: true parameter switches a request from a raw HTTP fetch to a full headless browser render, with fingerprint, locale, and timezone signals aligned to a real session, at no extra credit cost over the base scrape. That is the difference between a request that looks like a script and one that looks like a browsing session with a real execution context, which is exactly the distinction most generic bot detection is actually keying on.

curl -X POST https://api.anakin.io/v1/url-scraper \
  -H "X-API-Key: your_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/product-page",
    "formats": ["markdown"]
  }' 

If that comes back blocked specifically on ad-carrying pages while docs pages on the same domain succeed, switch to a full browser render:

curl -X POST https://api.anakin.io/v1/url-scraper \
  -H "X-API-Key: your_api_key" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/product-page",
    "useBrowser": true,
    "formats": ["markdown"]
  }' 
Toggle switch illustration showing Anakin URL Scraper's useBrowser parameter flipping a request outcome from flagged as bot to real session

This is not a workaround for a site that has explicitly declared ai-input=no in its Content Signals, or disallowed the relevant bot by name, and it should not be treated as one. It is the right tool for the mismatch case: a domain that has not restricted agent access but flags a plain HTTP client as a bot anyway, which is the case most RAG pipelines and research agents hit day to day, long before they ever encounter a site that has deliberately opted out.

For a pipeline that needs the rank-and-describe step before it knows which URLs are even worth fetching,

Anakin's Search API sits in front of URL Scraper for exactly that purpose, and Anakin's Browser API is available when a pipeline needs direct Playwright or Puppeteer control over the same stealth browser infrastructure instead of the managed URL Scraper endpoint. URL Scraper remains the starting point for the single most common job: turning one URL into clean content a pipeline can actually use.

What doesn't change

  • Search crawling is unaffected by this default. A domain's SEO visibility does not drop because of it, on its own.
  • Existing paid Cloudflare customers with domains already live before September 15 are not auto-switched. The new defaults apply to new domains and new customers, plus the free tier.
  • A site owner can always override the default in either direction, in Security settings, whenever they choose.
  • Content Signals in robots.txt are a declared preference, not a technical barrier, and the Duke study and the Perplexity incident both show declared preferences get ignored in practice more than site owners assume.

This sits next to the wider Cloudflare anti-bot picture. A request failing for TLS fingerprint, IP reputation, or browser-rendering reasons unrelated to this specific default needs a different fix: see How to Fix Cloudflare Blocking Your Automation System and 6 Ways to Reduce Browser Fingerprint Detection. For the search layer above this one, Google Custom Search API Shutdown: What to Use Instead covers the migration path.

Conclusion

Cloudflare drew a real line between Search traffic and Agent traffic, and that line now defaults to closed on ad-supported pages for new and free-tier domains, against a backdrop of AI traffic growing 6.5x faster than human traffic and bot traffic crossing 53% of the web. For anyone building agents or RAG pipelines, the practical move is to check a target domain's robots.txt Content Signals before writing a single line of fetch logic against it, know which named crawler the pipeline will actually present as, and reach for a real browser-rendering path when the problem turns out to be a mismatch rather than a deliberate opt-out.

Start with Anakin's URL Scraper and test against the domains that actually matter to the pipeline before assuming either way.

Frequently asked questions

When exactly did Cloudflare start blocking AI agents by default?

September 15, 2026. The default applies to new Cloudflare customers, new domains added by existing customers, and existing Free tier customers who had not changed their bot settings.

Does this affect every website on Cloudflare?

No. It only changes the default for new domains, new customers, and the free tier. Existing paid customers with domains already live before September 15 keep their prior configuration unless they change it. It also only applies to pages that display ads.

Will my agent get blocked on every page of an affected domain?

No. The block applies specifically to pages that carry ads. A documentation page or a login page on the same domain with no ad units is not covered by this default.

How can I tell if a specific block is this policy and not a rate limit or fingerprint issue?

Check the domain's robots.txt for a Content-Signal line. If ai-input is set to no, or the request fails consistently on ad-carrying pages specifically while succeeding elsewhere on the same domain, this default is the likely cause rather than a generic anti-bot trigger.

Does Googlebot get blocked too?

On ad pages where Training is blocked, yes, unless the site owner explicitly exempts it. Googlebot, Bingbot, and Applebot crawl for both Search and their own AI training under one bot identity, so blocking Training by default catches them on those pages too.

Is robots.txt actually enough to keep an unwanted AI crawler out?

Not reliably. A 2026 Duke University study of 130 self-declared bots over 40 days found AI search crawlers in particular rarely check robots.txt at all, and Cloudflare's own investigation into Perplexity documented a named crawler rotating user agents and IPs specifically to get around declared no-crawl rules on test domains. Treat robots.txt as a stated preference, not an enforcement mechanism.

What is Cloudflare's Pay Per Crawl, and is it related to this default?

Pay Per Crawl is a separate but related Cloudflare feature that lets a site charge AI crawlers per request using the HTTP 402 status code, with allow, charge, or block set per crawler. It runs alongside the Agent/Training default covered here, as part of the same broader shift toward giving site owners explicit, enforceable control over AI access instead of relying on robots.txt alone.

Is there a legitimate way to keep accessing a site that has blocked Agent traffic?

Respect the site's declared Content Signals and named-bot rules first. If ai-input is set to no, that is the site owner's stated position, not a technical obstacle to route around. For domains that have not restricted agent access but misclassify a plain HTTP client as a bot, a real browser-rendering request through a tool like Anakin's URL Scraper resolves that specific mismatch.

Sources