Showing posts with label AI training data. Show all posts
Showing posts with label AI training data. Show all posts

Tuesday, July 28, 2026

Authors have mixed feelings about the $1.5B Anthropic copyright infringement ruling; NPR, July 27, 2026

  , NPR; Authors have mixed feelings about the $1.5B Anthropic copyright infringement ruling

"Graeber is among the more than 300,000 writers involved in the suit who may soon be getting a modest windfall. A federal judge in San Francisco rubber stamped a $1.5 billion settlement in July resulting from a landmark class action lawsuit the authors brought against the AI company Anthropic two years ago...

AI companies often invoke the fair use doctrine – which enables the use of copyrighted works without the copyright holder's consent in some situations – as they try to make the case in court for training their models on these materials...

Chinese AI companies often use a technique to build their models called "AI distillation." This involves feeding their models the outputs generated by other AI models, often high-quality U.S.-based ones like OpenAI's GPT-4 or Anthropic's Claude, instead of directly training them on pirated copies of books by American authors...

One possible way for authors to get a fairer shake in the age of AI could be through the licensing of their work to AI companies...

There are already some such deals between publishers and AI companies in place, such as Perplexity AI's agreement with media entities like the Los Angeles Times and Le Monde to license content for the training of its models. There are also online licensing marketplaces, such as Created by Humans."

Friday, July 24, 2026

UK's Bloomsbury among beneficiaries of $1.5 billion Anthropic copyright lawsuit settlement; Reuters, July 22, 2026

 Reuters ; UK's Bloomsbury among beneficiaries of $1.5 billion Anthropic copyright lawsuit settlement

"Britain's Bloomsbury ​Publishing confirmed on Wednesday it was among ‌the beneficiaries of a landmark $1.5 billion settlement that resolves claims artificial intelligence ​company Anthropic used copyrighted books ​to train its AI models without ⁠purchasing the content.

Here are some ​more details:

  • Bloomsbury said a U.S. court ​identified 14,087 of its titles covered by the settlement, with proposed compensation of ​about $3,000 per title, split equally ​between the author and publisher...
  • The settlement ​is the largest known copyright payout ​in ⁠U.S. history."

Saturday, July 18, 2026

Hackers Expose How AI Music App Suno Stole Decades Worth of Copyrighted Music; Futurism, July 17, 2026

, Futurism; Hackers Expose How AI Music App Suno Stole Decades Worth of Copyrighted Music

The evidence is damning.

"A hack revealed in detail how AI music generating app Suno scraped millions of songs, likely including copyrighted ones, from across the web to feed into its AI model, 404 Media reports.

Suno, which is currently embroiled in multiple ongoing copyright lawsuits, has already admitted in response to legal action that it used “essentially all music files of reasonable quality that are accessible on the open internet” to train its music-generating AI."

Friday, July 10, 2026

The Work of Helping A.I. Destroy Work; The New York Times, July 10, 2026

  , The New York Times; The Work of Helping A.I. Destroy Work

"Every day, Mercor, a start-up that sells training data to artificial intelligence companies, pays 30,000 contractors more than $4 million to help make their jobs, and those of their colleagues, obsolete.

It’s gig work, but for professionals with rarefied skills. One recent Mercor posting offered $225 an hour for a voice actor able to maintain a customer service persona in fluent Hebrew. Another sought a Ph.D. physicist with a specialization in general relativity, astrophysics or cosmology. A third listing wanted a physician with more than three years of experience in the Rwandan primary care medical system.

Mercor and a handful of similar start-ups are the primary middlemen in a supply chain of “human data” that may power the next generation of A.I. As OpenAI, Anthropic and other major ventures compete to become the industry’s dominant platform, the market for premium data that has been vetted by experts is exploding.

No longer do the A.I. companies need armies of low-paid workers, often overseas, to do rote tasks like tag images of cars or transcribe audio. They need mathematicians to annotate proofs, lawyers to mark up briefs and professors to grade essays. That’s what Mercor and its rivals supply. To use the parlance of the industry, data labeling has moved up the “value chain,” and the start-ups that offer this service have become some of the fastest growing in Silicon Valley."

OpenAI may have made a fatal misstep in copyright fight with news orgs; Ars Technica, July 9, 2026

 ASHLEY BELANGER  , Ars Technica; OpenAI may have made a fatal misstep in copyright fight with news orgs

"OpenAI is facing calls for “serious sanctions” after fighting to keep news organizations from snooping through millions of logs to find evidence of users skirting their paywalls by prompting ChatGPT to regurgitate their articles.

This evidence is considered among the most important to both sides, potentially either dooming OpenAI as an infringer or exonerating its chatbot technology as a transformative fair use of news sites’ content."

Thursday, July 2, 2026

Microsoft Shareholder Sues Top Brass for AI Copyright Claims; Bloomberg Law, July 1, 2026

 

, Bloomberg Law; Microsoft Shareholder Sues Top Brass for AI Copyright Claims

"Microsoft Corp.'s executives and board directors misled shareholders in statements concealing its artificial intelligence tools were trained on copyrighted material, a new investor lawsuit said. 

The tech giant’s false statements about its AI strategy and violations of intellectual property law caused substantial damage to Microsoft and its shareholders, according to Eric Anderson’s stockholder derivative lawsuit filed Tuesday in the US District Court for the Western District of Washington."

Tuesday, June 30, 2026

Ford rehires human engineers after AI fails to match quality checks; BBC, June 29, 2026

Liv McMahon , BBC; Ford rehires human engineers after AI fails to match quality checks

"Ford says it has hired back some human engineers after AI failed to match their skills and experience.

In a bid to reap the benefits of the tech, which developers claim can cut costs and boost productivity, the US carmaker adopted it across some parts of its operations including for quality checks.

But, according to Bloomberg, its executives said the firm has rehired more than 300 "veteran" quality inspectors in recent years to make up for the pitfalls of automated systems.

"Artificial intelligence is a fantastic tool, but it's only as good as the information you use to train it," Charles Poon, vice president of vehicle hardware engineering, told reporters.

"Over prior years, we didn't pay as much attention as we should have to the experience of our most knowledgeable engineers that have been with us through many product cycles," he said.

The US automaker is among many to have seized on the buzz around AI, particularly amid Wall Street fervour about the tech's potential to increase margins."

Monday, June 29, 2026

NYT slams Microsoft for building copyright-infringing supercomputer for OpenAI; Ars Technica, June 26, 2026

 ASHLEY BELANGER , Ars Technica; NYT slams Microsoft for building copyright-infringing supercomputer for OpenAI

"In a heavily redacted court filing Thursday, The New York Times proposed to amend its copyright complaint against OpenAI and Microsoft to clarify a claim and allege that Microsoft actively encouraged OpenAI to steal NYT works by building a bespoke supercomputing system ranked among the most powerful in the world."

Saturday, June 27, 2026

How Teaching A.I. to Speak Cajun Can Help Save a Language; The New York Times, June 27, 2026

 , The New York Times; How Teaching A.I. to Speak Cajun Can Help Save a Language

By feeding centuries-old nursery rhymes and folklore recordings into their own model, linguists in Louisiana hope to help a community control its digital destiny.

"Louisiana French, the oral dialect of which Balfa was a cultural guardian, is part of the Bayou’s societal DNA, a link to its history, music and identity. Today, Caffery described the language as struggling and endangered, a notion reinforced by Alexa’s overlooking Balfa.

In response, Caffery assembled a small team at the center to train its own language learning model in automatic speech recognition for Louisiana French, drawing from a trove of historical artifacts and interviews.

Over the months, as the learning language model is trained on bits of the language — such as an old-age French nursery rhyme — it brings centuries-old dialect closer into the digital age."

Thursday, June 25, 2026

The New York Times Amends Lawsuit Against OpenAI and Microsoft; The New York Times, June 25, 2026

  , The New York Times; The New York Times Amends Lawsuit Against OpenAI and Microsoft

"The New York Times amended its lawsuit against OpenAI and Microsoft on Thursday, modifying one claim against Microsoft and dropping another against OpenAI, according to a legal filing in federal court...

In a filing in the U.S. District Court for the Southern District of New York on Thursday, The Times accused Microsoft of encouraging OpenAI to train its A.I. systems using copyrighted articles from The Times and of providing services designed to help with this training.

The Times also dropped a claim from its original lawsuit, filed in 2023, accusing OpenAI of “secondarily” infringing on its copyrights because it did not prevent consumers and businesses from generating copyrighted material using A.I."

Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft; New Jersey Globe, June 24, 2025

 David Wildstein, New Jersey Globe ; Nearly 400 local newspapers sue OpenAI, Microsoft over alleged copyright theft

"The massive coalition of local newspaper publishers filed a federal lawsuit today against OpenAI and Microsoft, alleging the technology companies systematically copied copyrighted reporting from nearly 400 local newspapers to train and develop commercial artificial intelligence products, including ChatGPT and Microsoft Copilot, without permission or compensation.

The publishers, represented by Platkin LLP, a law firm founded earlier this year by former New Jersey Attorney General Matthew J. Platkin, contend that OpenAI and Microsoft unlawfully appropriated original news content to build their AI systems, violating the Copyright Act and threatening the future of local journalism.

The lawsuit also alleges that OpenAI knowingly stripped copyright management information from publishers’ work — including author bylines, copyright notices, and terms of use information — in violation of the Digital Millennium Copyright Act.

The complaint cites remarks by OpenAI founder Sam Altman, who acknowledged during testimony before the British House of Lords that it would be “impossible to train today’s leading AI models without using copyrighted materials.”"

Tuesday, June 23, 2026

Archiving with AI; Library Journal, June 8, 2026

Matt Enis, Library Journal; Archiving with AI

"AI companies are offering some libraries funding for digitization projects, but archives and special collections are working through how to manage projects responsibly

“Imagine a world where you know things but cannot say where you learned them,” begins “Memory Without Origin,” a paper published in April by University of Virginia (UVA) Dean of Libraries and University Librarian Leo S. Lo. This isn’t a hypothetical question, Lo notes, it’s a predictable consequence if libraries allow generative artificial intelligence (AI) to ingest archival materials as training data without requiring provenance conditions. And libraries, which could always use funding for projects involving digitization, special collections, and archives, are being approached by AI companies with deep pockets.

“They’ve been approaching a lot of larger research libraries, including Oxford and many more,” Lo tells LJ. (Oxford’s Bodleian Libraries began a digitization pilot project funded by ChatGPT maker OpenAI last year.) “Usually the offer is: they will pay you to digitize materials—which we want, because we want to make them more accessible—and in return, depending on the deal…they would like to have the data to train their AI models.”

These partnerships can benefit both parties, but for libraries, the consequences of getting these arrangements wrong “are more permanent than anything the profession has previously encountered,” Lo writes. “Once archival materials are absorbed into foundation model weights, no subsequent institutional action can remove them from the model.” If proper care isn’t taken, that information becomes unmoored from its former context within an archive."

Friday, June 19, 2026

Millions of Copyrighted Songs Were Fed to AI Music Generators – Now There’s Proof; Gadget Review, June 16, 2026

 Al Landes , Gadget Review; Millions of Copyrighted Songs Were Fed to AI Music Generators – Now There’s Proof

Atlantic databases name 21 million tracks fed to Suno and rivals as Sony, UMG, and Warner seek $150,000 per song in damages

"Searchable databases verify roughly 21 million copyrighted songs trained AI music generators.

Sony, UMG, and Warner lawsuits seek up to $150,000 per song from Suno and Udio.

HarmonyCloak tool lets artists protect songs by adding inaudible AI-blocking audio perturbations.

Millions of copyrighted songs — including chart-topping hits — verifiably trained AI music generators, and now there are searchable databases to prove it. The Atlantic, through an investigation by staff writer Alex Reisner, published four catalogs documenting exactly which music fed these models:"

Tuesday, June 16, 2026

Publishers Sue WeLib for Copyright Infringement; Publishers Weekly, June 16, 2026

  Jim Milliot , Publishers Weekly; Publishers Sue WeLib for Copyright Infringement

"Fresh off of last month’s victory against pirate web site Anna’s Archive, 13 publishers across all segments of the industry have allied to sue yet another pirate site, WeLib, for copyright infringement.

The suit, filed in the U.S. District Court for the Southern District of New York, charges that the operators of WeLib “ copied the source code and most of the contents of” Anna's Archive."

The plaintiffs include the Big Five, Cengage, Elsevier, McGraw Hill, Pearson, Taylor & Francis, and Wiley.

“Defendants boast that they have reproduced ‘an endless collection of literature, research papers, and education materials,’ none of which they own or have licensed,” the complaint alleges. 

According to its website and repeated in the lawsuit, WeLib hosts over 43 million books and 98 million papers, and its stolen collection of literary works has purportedly attracted over 80,000 active monthly users. According to the website, WeLib’s users have illegally accessed over 51 million books in the last month alone, or an average of over 1.7 million books per day."

The Millions of Songs Mashed Into AI-Generated Music; The Atlantic, June 14, 2026

 Alex Reisner, The Atlantic; The Millions of Songs Mashed Into AI-Generated Music

"The actual recordings that go into any model are a closely guarded secret—AI companies have claimed they are proprietary—but the number of songs is almost certainly huge, spanning genres and time periods.

As part of my series of investigations into AI training data, I recently discovered four giant datasets of songs that are being shared within the AI-development community. One has 12 million tracks. Another has 9 million. The two smaller datasets each have more than 100,000. They include hits from major pop artists such as Bad Bunny, Nirvana, Taylor Swift, Billie Eilish, Pearl Jam, Elvis Costello, Sheryl Crow, and the Beatles. (The New Radicals’ “You Get What You Give” is in two of the datasets.) Jazz artists such as Miles Davis, John Zorn, and Vijay Iyer are featured, as are classical composers and tens of thousands of minor artists across genres. The 12-million-track dataset, on its own, would take 91 years to listen to...

In an attempt to prevent their products from generating songs that duplicate existing music, AI companies implement detection software. But neither Suno nor Udio prevents users from generating songs in the style of real artists. Earlier this year, Sony found 135,000 AI-generated tracks attributed to its artists on various streaming platforms. Although it’s not clear exactly which AI tools were used to generate those tracks, the technology is already harming artists’ ability to make a living from their music...

usicians and labels have filed at least 12 lawsuits against AI companies for training models on copyrighted music. The music industry’s three major labels have sued both Suno and Udio, and others have sued Google, OpenAI, and smaller AI vendors. No rulings have been issued in these cases, but some of the labels have reached settlements with Suno and Udio...

On the Free Music Archive, the guitarist and singer Derek Clegg has been sharing his original, home-recorded songs for more than 15 years. Clegg told me he’s happy for people to put his music in the background of their personal videos, as long as they credit him. When people expect to make money from the use of his music, then they pay him for a license. More than 250 of Clegg’s songs are in the FMA dataset I found. I asked whether he would opt out of AI training if a mechanism for doing so existed. “Yeah, definitely,” he said.

What bothers Clegg most is that AI companies take people’s music without consent, and without acknowledging that their tech products are entirely dependent on musicians. “It just seems dishonest. It seems like theft,” he said. “There’s going to have to be a reckoning.” That’s his hope, anyway."

Thursday, June 11, 2026

AI company argues its use of scraped Westlaw legal data was transformative; Courthouse News Service, June 11, 2026

  , Courthouse News Service; AI company argues its use of scraped Westlaw legal data was transformative

"“Fair use ruling here brings into question the core technology of the AI revolution,” Mark S. Davies of White & Case in Washington, attorney for ROSS, argued...

“This is a copyright case,” he said. “It’s an interesting case, it raises lots of issues, but it’s a copyright case and the point of copyright is progress.”

“Copyright is not a privilege reserved for the well-behaved,” Davies added."

Thursday, June 4, 2026

How to share the AI windfall. Are taxes enough?; The Economist, May 14, 2026

The Economist; How to share the AI windfall. Are taxes enough?

"Should artificial intelligence cause mass unemployment, workers will not be thrilled. But neither will the taxman, even if he hasn’t been automated. For most of the past century, rich countries have had simple rules for sharing prosperity: raise money mostly by taxing work and consumption, sprinkle in some borrowing and hand out the proceeds. That model may collapse if ai advances as quickly as its boosters suggest. Hence, many say, a new approach is needed, in which government makes its money primarily from the new technology."

Thursday, May 28, 2026

CNN Sues AI Firm Perplexity For Copyright Infringement; Deadline, May 28, 2026

  Jill Goldsmith, Deadline; CNN Sues AI Firm Perplexity For Copyright Infringement

"CNN is the latest to sue Perplexity for copyright infringement, alleging the AI firm “has unlawfully copied over 10,000 CNN stories, videos, images, and other works to power its products and tools.”

The suit said the two sides tried but failed to reach an agreement in 2025 and Perplexity continued ripping off CNN content and claiming a relationship with the news network that does not exist despite repeated warnings that the moves are illegal."

Tuesday, May 19, 2026

A 16th-Century Sketch Claims to Depict Anne Boleyn. A.I. Says It’s Her Mom.; The New York Times, May 19, 2026

, The New York Times; A 16th-Century Sketch Claims to Depict Anne Boleyn. A.I. Says It’s Her Mom.

Using facial-recognition technology, scholars have concluded that a 500-year-old drawing labeled “Anna Bollein Queen” more likely showed her mother, Elizabeth Howard.

"To dig into this mystery, Ms. Davies and her colleagues, including David G. Stork, a computer scientist and electrical engineer at Stanford University, turned to computational facial recognition. “This has one foot in art history and one foot in computer science,” Dr. Stork said...

Amit Roy-Chowdhury, a computer vision scientist at the University of California, Riverside, who was not involved in the research, said that facial recognition can play an important role in art history. But going forward, he added, it will be important to assemble larger training data sets of faces in artwork, as algorithms trained strictly on photographs can introduce uncertainties. And that could be a challenge, since there are millions of faces in photographs but far fewer in art. “For artwork, you don’t have that many examples,” Dr. Roy-Chowdhury said."