Microsoft, OpenAI Workers Fear ‘Largest Theft’
Unsealed court documents released this week have shed light on a growing concern within Microsoft and OpenAI: the use of millions of news articles to train their next‑generation AI models. The filings, part of a broader litigation over data licensing, detail how proprietary and public domain content is harvested, processed, and fed into machine‑learning pipelines without explicit author compensation.
According to the documents, the companies have amassed a dataset that includes millions of news stories spanning several decades. The data is then tokenized, anonymized, and used to fine‑tune language models that power everything from chatbots to code‑generation tools. While the companies argue that the data is transformed beyond recognition, critics point out that the sheer volume of content represents a form of intellectual labor that has not been monetized for the original creators.
Workers at both firms have voiced alarm, citing union representatives who describe the practice as a “largest theft of labor” in the history of the tech industry. The union’s statement highlights that the training process consumes significant computational resources, yet the original journalists and editors receive no royalties or acknowledgment. The workers fear that the continued use of such data could set a precedent that erodes the economic value of creative labor.
Industry observers note that the debate over training data is not new. OpenAI’s GPT‑4, for example, was trained on a mixture of licensed, publicly available, and user‑generated content. However, the scale of data now being used raises questions about the adequacy of existing copyright frameworks and the need for clearer licensing models that protect content creators while fostering innovation.
Regulators in the United States and the European Union are already discussing potential reforms. Proposed legislation would require AI developers to obtain explicit licenses for large corpora of copyrighted text or to pay a statutory fee. Such measures could reshape the economics of AI training and create new revenue streams for journalists and publishers.
As the legal battle unfolds, both Microsoft and OpenAI have pledged to review their data acquisition practices. The outcome of this case may set a critical precedent for how AI companies balance the need for vast datasets with respect for the labor that produces them.

