Two American newspapers with roots far removed from Silicon Valley have brought OpenAI and Microsoft before a New York federal court. The Seattle Times and Newsday allege that the two companies copied their journalism without authorization, including content accessible behind paywalls, to train or operate artificial intelligence products such as ChatGPT, Copilot, and Bing's AI features. It is an allegation, not a judicial finding, but the proceedings are significant because they further expand the legal front attempting to define how much copyrighted material can be used to build generative models.

The case comes after years of disputes between publishers, authors, photographers, musicians, and AI companies. The fundamental question is now well known: can training a model on protected works qualify as fair use under US law, or does it require licensing? For newspapers, however, the controversy holds a second, perhaps even more tangible dimension. If a system can answer a question by synthesizing information gathered from their work, users may have less incentive to visit the site that bore the cost of collecting, verifying, and publishing that reporting.

The lawsuit covers both training and the final product

According to Reuters' reporting, the complaint filed in the Southern District of New York claims that OpenAI and Microsoft scraped the two newspapers' websites and incorporated the articles into datasets used to train and run their systems. The outlets also challenge the ability of generative products to reproduce or paraphrase parts of their work without directing audiences to the original source.

The distinction between training and output is pivotal. A model does not operate like a traditional archive where every article remains stored as a retrievable page, but under certain conditions it can reproduce phrasing, facts, or structures learned from the data. The courts are still establishing case law capable of distinguishing between transformation, memorization, reproduction, and substitutive use.

Paywalls make the dispute even more sensitive

Seattle Times and Newsday emphasize that part of the disputed content was behind a paywall. This factor reinforces the economic argument: access to those articles is precisely what a reader is expected to pay for. If the material was collected through methods that circumvent terms of access or technical restrictions, the case could involve issues beyond the simple use of publicly available pages.

AI companies generally argue that training on vast amounts of data is transformative and that the models create new tools rather than substitute copies of the original works. OpenAI has also struck licensing agreements with numerous publishers, a sign that the market and the law are moving in parallel: while the courts determine which uses are permitted without authorization, companies are forging commercial agreements to reduce uncertainty and secure structured access to content.

Local journalism has a more fragile economy

The involvement of Seattle Times and Newsday makes the case different, at least symbolically, from a clash of titans. American regional journalism has endured two decades of eroding advertising revenue, closures, and newsroom cutbacks. The outlets that have maintained a significant structure increasingly rely on digital subscriptions and direct relationships with readers.

In this context, the prospect of a tech intermediary ingesting content and directly answering queries can be seen as a second disintermediation, following that of social media and search engines. In the 2000s, platforms captured the lion's share of advertising; in the AI era, the fear is that they might also capture part of the informational relationship.

Microsoft is involved because the ecosystem is intertwined

Microsoft's presence among the defendants shows how difficult it is to separate model, infrastructure, and distribution. Microsoft has invested heavily in OpenAI, integrates AI technologies into its products, and provides computing capacity and cloud services. For publishers, therefore, the issue is not just about who trains the model, but the entire chain that delivers a generative response to the public.

From a legal perspective, of course, liability and conduct will have to be proven separately. The fact that two companies collaborate does not automatically mean they share every liability. It will be up to the legal proceedings to determine which disputed activities are attributable to each party.

Citations and links do not solve everything

Many AI products are introducing links to sources, more prominent citations, and traffic-sharing programs. These are important tools because they acknowledge that a reliable answer depends on an ecosystem of verifiable sources. However, they do not completely resolve the economic problem. If users already find the answer in the summary, the presence of a link does not guarantee they will open it.

The publishing industry is therefore seeking a new balance: maintaining visibility in AI systems without merely becoming a free provider of raw materials. Some publishers are opting for licensing deals, others are blocking crawlers, and still others are filing lawsuits. The future market will likely combine all three approaches.

The question of value does not coincide with the question of law

Even if a court were to rule that certain forms of training fall under fair use, the industrial question of how to fund the production of original reporting would remain open. Conversely, a victory for publishers would not automatically guarantee a sustainable business model: it could create a licensing market dominated by the largest outlets, leaving smaller ones with less bargaining power.

This is one of the reasons why the copyright dispute is so crucial yet insufficient. The real issue is the architecture of information: who bears the cost of reporting, who captures user attention, who gets to monetize the answer, and how value is attributed to sources.

Courts are writing the rules while technology races ahead

Lawsuits against AI companies are moving forward as models are integrated ever more deeply into search, browsers, and operating systems. This creates a disconnect: a judicial ruling may arrive years after an industry practice has become standard. For newsrooms, waiting for the conclusion of all proceedings is not an option.

The Seattle Times and Newsday are thus using the law as one of the available tools to negotiate their place in the new ecosystem. OpenAI and Microsoft will have the opportunity to challenge the facts and legal interpretations contained in the complaint. Whatever the outcome, the case makes it harder to continue treating the issue as an abstract conflict between innovation and copyright. Behind every dataset, there are industries with costs, workers, and revenue models. In journalism, those costs are precisely what makes it possible to produce the information that a model can then synthesize.

The key question courts will have to untangle is the extent to which transformative use allows new tools to be built on copyrighted content. The issue the market will have to resolve is even broader: how to ensure AI improves access to information without consuming the economic system that continues to produce it.

Sources