← All writing
MadTech Aug 25, 2026 6 min read

Cloudflare’s new crawler defaults don’t touch most brand-owned content

Christopher Dorsey

Christopher Dorsey

AI & MadTech Advisor · Enterprise Sales Leader

TL;DR

On September 15, Cloudflare will start blocking Training and Agent crawlers by default on any page that carries ads, while Search crawlers stay allowed by default; the change applies as new domains onboard to a network Cloudflare says carries more than half of all web traffic. The same month, the Wall Street Journal reported that USA Today, Politico, Reuters and People Inc. are each weighing whether to block Google’s crawlers outright, after some publishers lost more than 40% of search traffic between June 2025 and June 2026 on Semrush data; Reddit is reevaluating a $60 million-a-year Google licensing deal, and Penske Media already sued Google over AI Overviews last year. USA Today CEO Mike Reed described the logic plainly: publishers with a licensing deal get crawled, everyone else gets blocked. Every one of those fights is about pages that carry ads, because that is the population Cloudflare’s new default targets and the population publishers monetize. Most brand-owned content, meaning blog posts, comparison pages, documentation and buying guides, doesn’t run third-party ads, so it sits outside the fight and stays open to the same crawlers by default. As the ad-funded web gets harder for AI answer engines to reach, whatever a brand already publishes on its own domain becomes a larger share of what is left to crawl and cite, without anyone touching a robots.txt file. The catch: any brand property that does carry ad script inherits the same default block, opting itself out of AI citations without anyone deciding to.

Cloudflare is changing its default settings on September 15. For any page that carries ads, the network will block Training and Agent crawlers automatically, while Search crawlers stay allowed by default. Cloudflare describes an ad as a signal that a site owner wants a person to land on the page and see something monetizable, so on those pages it now keeps out the bots that compete with human attention for advertisers. The default doesn’t touch pages without ads on them at all.

That single word, ads, ends up doing most of the work in this story, because it splits the entire fight over AI crawlers into two populations that are getting treated completely differently, and almost nobody writing about the fight has said so.

The publisher fight is entirely about pages that carry ads

The Wall Street Journal reported this summer that USA Today, Politico, Reuters and People Inc. are each weighing whether to block Google’s crawlers outright, on top of whatever Cloudflare changes by default. Some publishers lost more than 40% of their Google search traffic between June 2025 and June 2026, per Semrush data cited by the Journal, while The Guardian and the BBC gained. Reddit is reevaluating a $60 million-a-year licensing deal with Google. Penske Media sued Google over AI Overviews last year, arguing the feature repackages Rolling Stone and Variety journalism without permission.

USA Today CEO Mike Reed put the logic in the plainest terms I’ve seen from any publisher executive: “For those with licensing agreements, they get our content. For those without, we block them.” That is a company treating crawl access as inventory to be sold, which is exactly what it is once your page has ads on it and your traffic is the product.

Every name on that list runs an ad-supported page. Cloudflare’s new default and the publishers’ own blocking decisions are reactions to the same fact: an ad-monetized pageview is worth less once an AI answer replaces the click. Nobody is fighting over whether to let crawlers onto a page with no ad on it, because until this month there wasn’t much reason to.

Most brand content was never part of this fight

Company blog posts, product comparison pages, documentation sites and buying guides mostly carry no third-party ad script. They exist to get found and read. Under Cloudflare’s new default, none of that falls into the category getting blocked on September 15, and none of it is part of the licensing fight USA Today just picked with Google.

As more of the ad-monetized web either gets blocked by Cloudflare’s default or blocks itself on purpose, the pool of pages an AI answer engine can reach for a given question shrinks on the publisher side and stays exactly the same size on the brand side. Nobody has to win a negotiating position with Google for that to happen. It happens automatically, as a side effect of who is fighting and who isn’t.

I spent a stretch of my career on the AI-visibility side of this problem, and the part that never quite made it into the trade coverage is how much of citation share is decided by what’s reachable at all, before you get anywhere near what’s well-written or well-linked. I’ve written before that a page’s AI citation rate and its Google ranking are already answering two different questions. This is the same mechanic one level up: the answer engines can only cite what they can still crawl, and the publisher fight is shrinking one side of that pool while leaving the other untouched.

The exception is the brand property that runs ads

Some brand-owned content isn’t exempt at all. A media arm, a content hub monetized with programmatic display, a blog running a retargeting pixel through an ad network: all of that carries the same ad signal Cloudflare is now using to trigger the block, whether or not anyone on the marketing team thinks of it as a publisher. If your content sits on Cloudflare and carries even one ad script, you inherit the September 15 default along with every news site making headlines this month.

That makes this a bigger decision than it looks. Running a single display unit on a blog to cover hosting costs could quietly opt that content out of AI citations the day the default flips, with no announcement and no vote.

Check your own pages before September 15

If you run brand content on Cloudflare, go find out which of your pages carry any ad script, not just the ones your team thinks of as monetized. For anything that does, decide on purpose whether you want the September 15 default or want to override it in your Content Signals settings. For everything that doesn’t carry ads, you don’t need to do anything to benefit from this, which is a strange position to be in during a month when an entire industry is renegotiating its relationship with Google over exactly this kind of access.

Share this post

About the author

Christopher Dorsey

Christopher Dorsey

Enterprise Sales Leader · AI Go-To-Market · Startup Advisor · Denver, CO

Fifteen years selling technology to Fortune 500 brands across AI, advertising, and data infrastructure — most recently at Zeta Global, Oracle, and Fastly. Currently advising founders and sales leaders on AI go-to-market and Generative Engine Optimization.

Questions, pushback, or just want to compare notes?