The Newsroom Paywall Defense: How Media Publishers Monetize Archives and Structure B2B AI Licensing

Enterprise AI for News Media Paywall

Hey AI, I want to write an article for some news media. Could you give me an idea? The AI and I then began to explore topics and subjects for the article that likely suit from some news media. While exploring, I suddenly realized: wait, how come you (the AI model that I used) can access a news outlet? Is it the AI part or the search engine part? These two come in pairs, like Romeo and Juliet need each other for a great love story. AI without a search engine is not up-to-date and loses its main function. This makes me think that this topic is definitely entertaining and economically important to discuss.

In the past, AI companies have taken the benefit of using news outlet data to train their LLM models. The AI companies know exactly that the most powerful data lies not in an open internet forum, but behind newsroom archives. Prior to 2023, AI companies used to crawl news outlet sites to get their data from a decade ago, just like search engines crawl news outlet data. Of course, news outlets let search engines crawl their data for SEO and marketing purposes. The search engine drives traffic and enriches the news outlet's commercial interest to sell ad impressions and gain subscribers. It is not the same with AI companies. AI companies do not preserve the commercial bottom line of news media because the data is used to train their models for their advantage, bringing no traffic or commercial gain to the news outlet.

When news media protested, AI companies argued that they thought news was open public data for everyone to grab, just as search engines have done. Well, I believe it is not open public data. Readers tend to subscribe to read those news pieces and articles. We call this paywalled news. It is a commercial asset built with substantial capital investment, legal risk, and human editorial labor. It is only fair for news outlets to ask for a subscription when readers want to read their news.

Based on the current legal landscape, news outlets can prohibit AI company bots from crawling their data while still allowing search engines to use it for marketing purposes. Is it too late? I don’t think so. One fact remains: AI needed news outlet data yesterday, needs it now, and will still need it in the future. AI needs to be trained constantly or it becomes obsolete.

So here comes the punchline: If AI companies really need news outlet data, why don't news outlets simply let them have it, as long as they pay for it! In the last few years, OpenAI has reached deals with News Corp, owner of The Wall Street Journal and The Times, for $250 million across 5 years. Meta has signed deals with Reuters and Newsmax, Anthropic with publishers, and Google with Reddit. I think Indonesian media should follow suit.

Many Indonesian news outlets are currently struggling with declining advertising revenues and other challenges. A research study published by the Indonesian Cyber Media Association (AMSI) in collaboration with Monash University Indonesia and PR2Media clearly shows that Indonesian news sites lose up to 40% of search traffic to AI search summaries. The main problem here is scale, particularly for smaller regional digital publications that lack individual leverage to negotiate with trillion-dollar tech giants.

For this, the music industry set an example a decade ago: when individual musicians could not negotiate with giant platforms like Spotify, Collective Management Organizations (CMO) or Lembaga Manajemen Kolektif (LMK) emerged. The product is fundamentally the same: Intellectual Property Rights covered under copyright law. It makes sense now that regulators explicitly seek to clarify journalistic works as part of it. Alternatively, news outlets can pool together and create a consortium for their paywalled archives to build unified API feeds that command much higher licensing valuations.

Another strategy is structuring B2B API Licensing Agreements. Rather than giving a long-term or perpetual blanket for training rights, news outlets could structure a time-bound, pay-per-use, or annual subscription for API Licensing Agreements. This ensures that AI companies pay recurring royalties for real-time news access while protecting news outlet copyright ownership.

Artificial Intelligence does not have to be the end of digital journalism. The true value still lies behind newsroom archives to be monetized with subscription-grade reporting, embracing collective licensing, and paid B2B API contracts. This way, Indonesian news outlets can turn a technological threat into sustainable, recurring revenue.

Previous
Previous

The Due Process of Data: Architecting Legally Sound Pipelines to Eliminate Tainted AI Liabilities

Next
Next

Network Scale & Brand Integrity: How Dual-Grounding AI Protects Franchise Royalties and Trademark Licensing