A Microsoft executive called the use of news to train AI ‘the biggest job theft in history’

Brent Hechtdirector of Applied science of Microsoftwarned in internal documents that using news to train artificial intelligence models could be the ‘biggest job theft in human history’. His words have transcended the judicial process he faces in USA to various media with Microsoft and OpenAI for the use of its contents without authorization.

The plaintiffs, led by The New York Times, filed a brief that was made public on Thursday, September 17, Ars Technica reports. They request that the judge resolve several issues in the case in their favor without waiting for the trial, including whether companies infringed your copyright and whether they can rely on fair use. They argue that the evidence already allows us to decide on these points and leave the issues pending for the trial. The document collects its arguments and quotes internal company communications.

Hecht’s warnings described news mining as ‘a staggering theft of unprecedented proportions’. He also wrote that the plan to collect this content massively ‘completely mocks the idea of ​​”fair use”‘. In another document he acknowledged that ‘almost no one created their content with the intention of it being used in this way, nor does he receive compensation for its use‘.

Microsoft and OpenAI defend that training their models with news is covered by ‘fair use’ or fair use. This figure of American law allows protected works to be used without permission in certain circumstances. The judges assess the purpose of the use, the nature of the work, how much it is copied and the effect on its market, as explained by the United States Copyright Office.

The media argues that chatbots reproduce your articles and satisfy the demand for information that they previously served. To demonstrate this, they conducted tests in which they obtained long verbatim fragments by asking for summaries, key points or evaluations of news bias. Their request focuses on articles with broad overlaps and they consider that that reproduction and substitution of its services contradict the defense of fair use.

The plaintiffs also point to Microsoft data with drops of between 83% and 93% in click-through rates for some media outlets. In internal messages, Nick Turleyresponsible for ChatGPTacknowledged that ‘there is no good reason to click’ when the chatbot provides the information.

The CEO of Microsoft, Satya Nadellaadmitted under oath that Chatbots have served as substitutes for journalistic platforms by offering information directly and preventing the user from going to the source.. He also stated that AI companies should not violate website conditions by bypassing their paywalls.

Internal documents also reflect concern for a ‘vicious circle’ which would end up hurting the AI ​​companies themselves. If the media lose income and stop producing information, models will lose one of their content sources.

Another warning from Hecht concerned a filter that, according to the plaintiffs, made it difficult to verify what content the chatbots were reproducing. The manager pointed out that it could be perceived as a ‘accidental cover-up’ because it would mean that ‘people who have rights to the content have less information about what was used for training’.

Microsoft’s public response distances itself from those assessments collected in internal documents. A spokesperson told Ars Technica that Hecht’s comments ‘they reflect the individual perspective of an employee, do not constitute legal analysis and do not represent the position of the company’. He defended that his products do not replace the media and that Nadella’s words were general observationsnot legal conclusions.