Cronkite

Bulletin of August 21, 2026

6 minBusiness

Publishers lose the bot war, so the fight shifts to AI answers

A new report shows unauthorized AI crawlers are increasingly hard to block, even for large media companies. The real battle for publishers is no longer about preventing scraping but about controlling how their content appears in AI-generated answers and building sustainable business models around retrieval.

The battle to keep unauthorized artificial intelligence bots out of publisher websites is effectively over, according to a new industry report that suggests even the best-resourced media companies cannot fully win that fight. The finding is pushing the debate over AI and journalism away from blocking crawlers and toward a more practical question: how publisher content is used in AI answers and how that use can be turned into revenue.

The report, State of the Bots from TollBit, a company that builds payment systems between publishers and AI crawlers, includes an account from Jonathan Roberts, chief innovation officer at People Inc., about his company's approach to bot traffic. Roberts said the company aggressively blocks unauthorized bots while allowing access to legitimate crawlers, usually through licensing agreements. But he conceded that unauthorized crawlers have become much harder to identify. The report documents cases of bad bots disguising themselves as Google's crawler, rotating IP addresses after a blocked first attempt, and industrial-scale scraping operations using networks of devices in people's homes to make traffic appear human.

If a large media company with dedicated expertise cannot keep all bad bots out, the report argues, smaller publishers have little chance. The conclusion is that the future of publishing in the AI era cannot be defined by the fight over bot access. That fight, the report says, has already been lost. The focus instead needs to shift from what content is scraped to how that content is used, and what publishers can do to preserve value and encourage good behavior from AI companies.

The report draws a clear line between two distinct uses of publisher content: training AI models and retrieving information in real time. Training large language models is expensive, with runs costing hundreds of millions or even billions of dollars, and it is done by very few companies. The report notes that training is not the main threat to media business models anymore. The real threat is retrieval, when people use AI as a discovery surface and expect accurate, up-to-date answers. That requires real-time access to fresh content, which is a different kind of bot interaction with different stakes.

Licensing deals have already shifted in that direction. Rob Kelly of the Media and the Machine Substack tracked 94 publicly announced deals between publishers and AI companies and found that only about four in ten now include training rights. The market, Kelly observed, is moving from buying content to build better models toward licensing content to deliver better answers.

Being cited in an AI answer, however, does not automatically translate into revenue. A recent Digiday story highlighted the struggles brands face in connecting AI presence to business outcomes, echoing what publishers have long felt: appearing as the authoritative source in an answer has some value, but it is not in itself monetizable. Studies consistently show that most AI users never click through to sources. Pew Research found that clicks on links inside a Google AI summary were about 1 percent of visits, compared with 15 percent on a results page with no AI answer at all.

Despite the lack of direct financial upside, there is clear value to users in getting accurate information from AI answers. The challenge is quantifying that value. Current approaches include pay-per-crawl, pay-per-use, licensing, and serving ads to bots. None has become an industry standard, largely because enforcement is difficult. There are too many paths for content to enter the ecosystem: stealthy bots crawling information they should not, AI companies purchasing data from gray-market scraper firms, or honest retrieval of republished content.

The report suggests a different approach: police the answer, not the crawl. In an ideal system, an answer engine would find content, check sources, and then verify that it has legitimate access to those sources through licensing or another agreed business model. Attribution already works well in retrieval systems, the report notes. All major AI engines give citations, and companies like ProRata have built business models around accurately showing which sources contributed to an answer and how much. If content does not pass the access check, it would not appear in the answer.

This approach would not solve every problem. Unauthorized scraping would still happen, and gray-market data would still circulate. But the report argues that focusing on the answer itself gives publishers a more enforceable point of control than trying to police every crawl. It also aligns incentives: AI companies that want to deliver high-quality answers would have a reason to ensure their sources are legitimate, and publishers would have a clearer path to getting paid for the value they provide.

Erin Baxter

Author

News Editor

Erin Baxter covers public affairs, politics, business, culture and daily news for Cronkite. The role focuses on verification, context, and clear explanations for readers.

Tail slate

Reporter
Erin Baxter
Filed
Runs
6 min
Source
Fast Company
Block
Business

Next in the Business block

  1. ——:—— Aug 21 Walmart opens 100th company-owned EV fast-charging site 3 min
  2. ——:—— Aug 21 Small employers are winning over the next generation 4 min
  3. ——:—— Aug 13 Warner Bros. Puts Hogwarts Legacy Sequel in Its Profitability Pipeline 4 min
  4. ——:—— Aug 13 Pokémon Card Index Rises 27.9% as Onchain Marketplaces Expand 4 min