Showing posts with label AI training data. Show all posts
Showing posts with label AI training data. Show all posts

Saturday, September 19, 2026

Scoop: DOJ's copyright filing took key agencies by surprise; Axios, September 19, 2026

 Sara Fischer, Kerry Flynn, Axios; Scoop: DOJ's copyright filing took key agencies by surprise

"The Department of Justice's statement of interest supporting OpenAI and Microsoft in the New York Times' copyright infringement lawsuit surprised critical agencies like the U.S. Patent and Trademark Office and the Copyright Office, sources told Axios.

Why it matters: Statements of interest allow the government to declare an official position on a legal matter in private lawsuits. While not binding, they can hold significant weight and help persuade cases.

  • The DOJ's SOI argues copyrighted works to train models should be considered fair use because that practice is new and transformative, but also says outputs aren't necessarily covered by that same legal argument.

  • Unlike many SOIs, no career antitrust attorneys signed the filing alongside senior DOJ officials.

Between the lines: Publishers have criticized the claims in the SOI, including the idea that enforcing copyright laws is too cumbersome and would threaten America's AI dominance over foreign rivals."

Friday, September 18, 2026

The AI Copyright Cases Are Starting to Tell Us Something; Legalytics, September 14, 2026

 ADAM FELDMAN, Legalytics; The AI Copyright Cases Are Starting to Tell Us Something

"The emerging question is therefore no longer simply whether AI companies can train on copyrighted material. A potentially more useful set of questions is coming into view: What exactly did the defendant copy? How did it obtain the material? What copies did it retain? What can the resulting system reproduce or retrieve? Does the AI product compete with the copyright owner’s market? And what evidence exists that the challenged conduct has actually produced market harm?

The answers do not yet produce a reliable formula for predicting who will win though. Most of the litigation remains unresolved, and motions to dismiss, preliminary injunctions, summary judgments, and settlements answer different legal questions. But the accumulated cases now provide enough variation to identify where the pressure points are developing. And those key points suggest that the next phase of AI copyright litigation may look considerably less like one giant test of AI training than the first wave of cases made it appear.

Two Courts, but an Increasingly Diverse Set of Cases

The litigation is remarkably concentrated geographically. Of the 105 captured U.S. infringement actions, 49 were filed in the Northern District of California and 32 in the Southern District of New York. Together those districts account for 77% of the filings in the dataset.

Consolidation does not explain away the pattern. When related proceedings are collapsed into litigation families, roughly the same share—76%—remains centered in those two courts."

Thursday, September 17, 2026

Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History; The New York Times, September 17, 2026

 Karen Weise and  , The New York Times; Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in History

Newly unsealed court documents showed concern within Microsoft and OpenAI over the use of millions of news articles to develop A.I. systems.

"Newly unsealed court documents showed considerable concern within Microsoft and its close partner OpenAI over the use of millions of news articles to develop artificial intelligence systems.

As OpenAI was forging ahead with its work, Microsoft employees debated whether what OpenAI was doing represented the “largest theft of labor in human history” and could create a “doom loop” that could ultimately threaten the quality of the large language models they were building...

Snippets of those discussions were made public on Thursday as part of a closely watched lawsuit The New York Times filed against OpenAI and Microsoft in late 2023. Eleven other publishers have joined the suit. Judge Sidney H. Stein of U.S. District Court for the Southern District of New York is considering motions for a summary judgment. Documents related to the case are slowly being unsealed as the judge considers those motions.

The publishers argue that the tech companies violated copyright law by scraping millions of their stories off the internet and other databases, and using the text, without approval or pay, to train advanced A.I. systems.

Microsoft and OpenAI contend their work was covered under legal protections for “fair use” of copyrighted material. They say the articles were sufficiently transformed into entirely new work by A.I., and were not substitutes that harm the value of the original work."

Wednesday, September 16, 2026

Why the DOJ’s OpenAI copyright stance is the real threat to national security; ZDNET, September 15, 2026

 David Gewirtz, ZDNET; Why the DOJ’s OpenAI copyright stance is the real threat to national security

The DOJ argues that AI training is transformative fair use, but publishers say unlicensed scraping threatens their survival. This copyright fight could shape the future of online knowledge.

"Ever since generative AI arrived in early 2023, we’ve seen that its almost unlimited base of knowledge is due to how the big AI companies trained their models. To a large degree, AI models like those from OpenAI and Anthropic have been trained on anything they could ingest, including nearly all the copyrighted material on the web.

The AI companies are even reported to be buying up physical books by the millions, cutting them apart, scanning them in, and then disposing of the remains. For example, based on a search of the Anthropic settlement database, I know the company scanned my book, The Flexible Enterprise, and included it in the Claude corpus. I was never asked for permission. I don’t even get a free Claude account.

Due to the bulk ingestion of intellectual property, many companies filed suit against the AI companies. One such company is Ziff Davis, the owner of ZDNET. Disclosure: Ziff is also the company that pays me each week for my writing here.

Two weeks ago, on Sept. 1, 2026, the US Department of Justice filed a Statement of Interest with the US District Court for the Southern District of New York, where the case is being litigated. What makes this statement particularly interesting is that the DOJ is not one of the parties in the case. The government is putting its thumb on the scale, weighing in on a lawsuit between private parties.

On Monday, Fortune published a commentary by Vivek Shah, CEO of Ziff Davis, regarding the Justice Department’s unusual intervention in the case. In this article, I’ll briefly summarize the DOJ’s statement, then discuss Shah’s premise, and then pick up and expand upon it with some of my own thoughts."

Monday, September 7, 2026

Court Filings in A.I. Suit Invoke Copyright Law, Culture and Sports; The New York Times, September 4, 2026

 Mike Isaac and  , The New York Times; Court Filings in A.I. Suit Invoke Copyright Law, Culture and Sports

Filings made Friday in The New York Times’s closely watched lawsuit against OpenAI and Microsoft included a range of copyright law and cultural references.

"Court filings made Friday in a closely watched copyright trial pitting The New York Times against OpenAI and Microsoft invoked a wide range of material, including relevant copyright law, arts and sports.

The suit, filed in 2023 by The Times and joined by a group of other news outlets, claims that OpenAI, a leading artificial intelligence start-up, and its partner Microsoft infringed on the publishers’ copyrighted material by using millions of their articles to train A.I. technologies. A.I. companies now compete with The Times as a source of information, the news outlet argued in its suit.

The briefs, filed in the U.S. District Court for the Southern District of New York, largely boiled down to two questions: whether the publishers’ news articles were sufficiently “transformed” into an entirely new work by A.I., and whether A.I. produced content that “substituted” for news articles and harmed their value.

Friday was the last day the companies could file motions for a summary judgment that would head off a trial. Judge Sidney H. Stein is expected to make a ruling in the coming weeks."

Thursday, September 3, 2026

Justice Dept. Sides With OpenAI in New York Times Copyright Suit; The New York Times, September 2, 2026

Karen Weise and , The New York Times; Justice Dept. Sides With OpenAI in New York Times Copyright Suit

"The Justice Department told a Manhattan federal court that it was in the national interest for the judge to find that OpenAI did not violate copyright law when it used articles by The New York Times and other publishers to develop artificial intelligence systems.

The filing late Tuesday was the first time the Justice Department weighed in on the use of copyrighted material by A.I. companies, which has led to several lawsuits, including one brought by The Times.

The Justice Department argued that developing A.I. was critical to national security, and that training A.I. systems sufficiently transformed the written works to new material allowed under copyright law. It said the benefits of A.I. “far outweigh any competitive harm.”

The government’s intervention is an escalation in the landmark litigation that could determine whether OpenAI violated the law when it was developing its A.I. systems and had harmed the news industry and other content creators...

The Times’s lawsuit is one of many amid a wave of legal action against A.I. companies over copyright claims."

Saturday, August 29, 2026

Sony, Warner sue Anthropic, alleging "blatant theft" of intellectual property; Axios, August 29, 2026

Ben Berkowitz , Axios; Sony, Warner sue Anthropic, alleging "blatant theft" of intellectual property

"Some of the world's largest music publishers filed a blockbuster lawsuit against Anthropic late Friday night, alleging "one of the largest and most blatant ongoing thefts of intellectual property in history."

Why it matters: The suit is the opening salvo in what is now likely to be a yearslong fight over music, AI, and how intellectual property is protected in a new era of technology.

Driving the news: Units of Sony Music and Warner Music filed the suit in federal court in northern California late Friday night, naming Anthropic, CEO Dario Amodei and co-founder Benjamin Mann as defendants. 

The big picture: The Sony/Warner lawsuit is notable because it's broad.

It alleges Anthropic unlawfully trained its models off "tens of thousands" of music publishers' copyrighted compositions, whereas other lawsuits have focused on a narrower set of works.

BMG's lawsuit against Anthropic, for example, claims infringement against 493 compositions."

Thursday, August 27, 2026

From Napster to Sampling to AI: Copyright Law’s Role as the Sheriff to Emerging Technology; JD Supra, August 25, 2026

 Richard Busch , JD Supra; From Napster to Sampling to AI: Copyright Law’s Role as the Sheriff to Emerging Technology

"Copyright law rarely meets new technology at the start. More often, it arrives after technology, and those behind it seeking to push (or ignore) the law for monetary gain have already changed markets, habits, and expectations. In music, that pattern has repeated across many major technological shifts. This article focuses on several: digital sampling, which made fragments of existing recordings newly usable in the studio; digital downloads, which changed the economics of distribution; interactive streaming, which made all music available for license on demand; platform music libraries, which embedded music into social media tools; and artificial intelligence, where those building literally trillion dollar businesses have trained on copyrighted works, and used others’ voice and likeness. In each of these instances, creators and rights holders have had to ask courts, Congress, or regulators to apply existing legal frameworks to commercial realities.

My work in music copyright litigation has often involved a recurring collision between music, technology, and copyright law."

Wednesday, August 26, 2026

The Original Sin of Anthropic’s Claude; The New York Times, August 24, 2026

 Charles Graeber, The New York Times; The Original Sin of Anthropic’s Claude

"One day, hopefully copyright protections will be extended to prevent the unlicensed training of A.I.s.

Until then, we live in a legal gray area, in which giant cash-rich corporations gorge freely on the rights of individual creators. The big ones will continue to eat the little ones unless the little ones can band together.

It’s up to us to demand legislation extending copyright protections in the 21st century. Many members of Congress appear willing; property protections have appeal for both sides of the aisle.

Affirming those copyright protections to include A.I. might even force A.I. companies to hire human writers to create original works for their L.L.M.s to train on. Or perhaps they’ll start spending their treasure like modern Medicis, funding the quality work A.I.s need to stay sharp. At the very least, they might invest in publishing houses and writing programs, fending off model collapse while ushering in a Silicon Era of artistic enlightenment.

I worry for this generation of artists coming of age in yet another technological adolescence, on the brink of so many cultural and economic disruptions.

But I do not worry for the future of art itself. We will always need human mediators to translate the human experience and make the flesh word."

Monday, August 24, 2026

Is it legal to train AI models on copyrighted books? It’s complicated; TechCrunch, August 23, 2026

 Amanda Silberling, TechCrunch ; Is it legal to train AI models on copyrighted books? It’s complicated

"You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right? 

The reality isn’t that simple. 

“I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on,” Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch. “It’s very complex and there are a lot of raw feelings about what is happening, both for and against.”"

Monday, August 17, 2026

Hidden Airtag reveals Amazon is trashing rare books to train AI; Ars Technica, August 17, 2026

 ASHLEY BELANGER , Ars Technica; Hidden Airtag reveals Amazon is trashing rare books to train AI

"For the past year or so, booksellers have suspected that AI firms are buying up huge lots of rare books, then destroying them after scanning them to train AI. But this was hard to prove until now, as 404 Media reports that an Airtag hidden in a rare book shows that at least one tech giant, in the race to advance its frontier models, is behind some of the bulk orders: Amazon.

On Monday, 404 Media revealed that it had connected with a bookseller who agreed to plant an Airtag in a rare book that was part of a bulk order. That Airtag was then tracked to an Amazon AI training facility in Las Vegas that housed a team focused on tearing books from their spines and scanning pages, 404 Media reported. Apparently tone-deaf to the escalating backlash over destructive book scanning, a logo on the door of that team’s warehouse, VGT3, showed a Tyrannosaurus rex preparing to devour a book, 404 Media documented."

Apple Shareholders Sue Company Heads Over AI Copyright Claims; Bloomberg Law, August 17, 2026

 

 , Bloomberg Law; Apple Shareholders Sue Company Heads Over AI Copyright Claims

"Apple’s top directors and officers knew of copyright violations while training artificial intelligence models, leading to potentially liabilities and reputational damages, a new investor lawsuit said. 

Apple’s executives and board members violated their fiduciary duties and acted in bad faith by “knowingly or recklessly” causing the company to engage in copyright infringement, according to shareholder Phil Rosen’s derivative lawsuit filed Friday in the US District Court for the Northern District of California. 

Multiple tech giants have been sued by shareholders over AI-related claims, including copyright concerns and allegations that they over-hyped capabilities. Apple is also facing separate copyright infringement lawsuits. 

Apple trained its models on data sets that included pirated books as well as videos extracted without authorization from YouTube or the copyright of those works, the lawsuit said."

Wednesday, August 5, 2026

Why is Anthropic destroying books?; The Guardian, August 5, 2026

Kathryn James, The Guardian; Why is Anthropic destroying books?

"Should we be surprised that destroying printed texts seemed easier to Anthropic than working with their human authors?"

Legal Battle Over U.S. Copyright Chief Heats Up Again; Publishing Perspectives, August 4, 2026

 Andrew Albanese, Publishing Perspectives; Legal Battle Over U.S. Copyright Chief Heats Up Again

"In the United States, a legal battle over the future of the nation’s top copyright officer, Register of Copyrights Shira Perlmutter, is heating up again. In a legal filing last week, Administration lawyers told a Washington D.C. court that a recent Supreme Court decision bolstered their case that President Trump has the power to fire Perlmutter. But in a filing of their own, lawyers for Perlmutter reiterate that the president lacks such authority, arguing that, as an appeals court ruled last fall, the law gives that authority to the U.S. Librarian of Congress.

The legal drama began last May, when Trump purportedly fired Perlmutter, just two days after the shock firing of Librarian of Congress Carla Hayden. The firing surprised and outraged stakeholders in the copyright and Intellectual Property communities, who have given Perlmutter high marks for her work at the office, including significant progress on a much-needed modernization effort.

More concerning, however, Perlmutter’s attempted firing came immediately following the prepublication release of the Copyright Office’s third and final part of a wide-ranging AI review, which argued for the rights of copyright owners—an opinion that appears to clash with the president’s AI goals."

Tuesday, August 4, 2026

Authors weigh in on $1.5 billion Anthropic AI copyright settlement; WGCU, August 3, 2026

Mike Kiniry, WGCU ; Authors weigh in on $1.5 billion Anthropic AI copyright settlement

"As Generative AI language models have entered the scene in recent years, a wave of copyright lawsuits has arisen in response, brought by authors and publishers. These lawsuits hinge on whether downloading and ingesting millions of copyrighted books without explicit permission to train Large Language Models constitutes copyright infringement or is protected as fair use.

Authors and publishers argue that it is infringement — particularly when AI developers illegally pirate or copy their books to help train their language models. AI companies argue that reading and learning from text is transformative and therefore falls under fair use.

In one class action lawsuit that was recently settled, the AI Company Anthropic agreed to pay $1.5 billion dollars in a landmark copyright infringement settlement. It's one of the biggest in U.S. history.

There are other similar high-profile cases, including one by publishing houses including Hachette, Macmillan, and McGraw Hill, along with bestselling novelist and former President of the Author's Guild Scott Turow against Meta and its CEO, Mark Zuckerberg and another against Google. Those cases are ongoing.

The Anthropic settlement means payments of roughly $3,100 to the authors and publishers of nearly half a million books, including our guests. We have a conversation about that settlement, and other pending cases, and what this all means for the publishing world.

Guests:

Marty Ambrose-McLaughlin is an award-winning author and English instructor at Florida Southwestern State College
Scott Turow is a writer and former attorney. He is the author of fourteen works of fiction, including Presumed Innocent and his most recent, Presumed Guilty which was published in 2025."

Tuesday, July 28, 2026

Authors have mixed feelings about the $1.5B Anthropic copyright infringement ruling; NPR, July 27, 2026

  , NPR; Authors have mixed feelings about the $1.5B Anthropic copyright infringement ruling

"Graeber is among the more than 300,000 writers involved in the suit who may soon be getting a modest windfall. A federal judge in San Francisco rubber stamped a $1.5 billion settlement in July resulting from a landmark class action lawsuit the authors brought against the AI company Anthropic two years ago...

AI companies often invoke the fair use doctrine – which enables the use of copyrighted works without the copyright holder's consent in some situations – as they try to make the case in court for training their models on these materials...

Chinese AI companies often use a technique to build their models called "AI distillation." This involves feeding their models the outputs generated by other AI models, often high-quality U.S.-based ones like OpenAI's GPT-4 or Anthropic's Claude, instead of directly training them on pirated copies of books by American authors...

One possible way for authors to get a fairer shake in the age of AI could be through the licensing of their work to AI companies...

There are already some such deals between publishers and AI companies in place, such as Perplexity AI's agreement with media entities like the Los Angeles Times and Le Monde to license content for the training of its models. There are also online licensing marketplaces, such as Created by Humans."

Friday, July 24, 2026

UK's Bloomsbury among beneficiaries of $1.5 billion Anthropic copyright lawsuit settlement; Reuters, July 22, 2026

 Reuters ; UK's Bloomsbury among beneficiaries of $1.5 billion Anthropic copyright lawsuit settlement

"Britain's Bloomsbury ​Publishing confirmed on Wednesday it was among ‌the beneficiaries of a landmark $1.5 billion settlement that resolves claims artificial intelligence ​company Anthropic used copyrighted books ​to train its AI models without ⁠purchasing the content.

Here are some ​more details:

  • Bloomsbury said a U.S. court ​identified 14,087 of its titles covered by the settlement, with proposed compensation of ​about $3,000 per title, split equally ​between the author and publisher...
  • The settlement ​is the largest known copyright payout ​in ⁠U.S. history."