Showing posts with label scrapers. Show all posts
Showing posts with label scrapers. Show all posts

Sunday, August 23, 2026

Artists Built A Site To Escape AI. Scrapers Are Coming For It Anyway; Forbes, August 23, 2026

Rob Salkowitz, Forbes; Artists Built A Site To Escape AI. Scrapers Are Coming For It Anyway

"Scrapers defend their actions

Artists’ claims to own and control their own work online are disputed by individuals, groups and commercial entities who believe that advancing the progress of AI entitles them to any and all data they can obtain, regardless of consent. This was apparently the motivation of the original scraper, who posted an archive of nearly 12 million images from Cara to the sub-Reddit r/DefendingAIArt under the handle “MandarinDrawnPoppy994.”

“Scraping is necessary to develop good models. It’s like building a highway – some houses must be demolished, but in the end everyone benefits,” the poster wrote in a thread titled “[AMA] I scraped all of Cara.”

Zhang says she and others reached out to the original scraper and eventually prevailed on him to take down the post. She adds, “not only did the first scraper delete the dataset, but he has turned around now to offer help, and we’re now co-creating an open source tool separate from Cara that will help people check if they've been scraped in new datasets in the future.”

Unfortunately, that was not the end of the problem. Several days later, the site was scraped again by a different actor. This time the data was posted on Hugging Face, a hub of resources for AI developers rooted in the open source community, by a poster under the name “Ioannis/Captive Dreamer.”

After some people reported the post to Hugging Face Trust and Safety, the team responded that “we have reviewed these [copyright reports] carefully. Because no copies of the artworks are hosted here, and because the URLs [in the dataset] point to the copies the artists published on Cara, there is nothing hosted on Hugging Face that we can disable through our notice and takedown process. This is not a judgement about who owns the works (the artists do); it is about what is stored on our servers. Further copyright reports on the same basis will not change this outcome.”

Hugging Face did not respond to a request for further comment for this story.

Third attack in 10 days

Now on Saturday, August 22, Zhang says the site was scraped for a third time, with the perpetrator taking just 123K images, but also a second data set that includes users’ text posts, bio and information they share on the site. That archive has been posted on Academic Torrents, a site that “was established to meet the demands of science in the age of big data” by providing data for researchers, according to its “About” page."

Saturday, July 6, 2024

THE GREAT SCRAPE: THE CLASH BETWEEN SCRAPING AND PRIVACY; SSRN, July 3, 2024

Daniel J. SoloveGeorge Washington University Law School; Woodrow HartzogBoston University School of Law; Stanford Law School Center for Internet and SocietyTHE GREAT SCRAPETHE CLASH BETWEEN SCRAPING AND PRIVACY

"ABSTRACT

Artificial intelligence (AI) systems depend on massive quantities of data, often gathered by “scraping” – the automated extraction of large amounts of data from the internet. A great deal of scraped data is about people. This personal data provides the grist for AI tools such as facial recognition, deep fakes, and generative AI. Although scraping enables web searching, archival, and meaningful scientific research, scraping for AI can also be objectionable or even harmful to individuals and society.


Organizations are scraping at an escalating pace and scale, even though many privacy laws are seemingly incongruous with the practice. In this Article, we contend that scraping must undergo a serious reckoning with privacy law. Scraping violates nearly all of the key principles in privacy laws, including fairness; individual rights and control; transparency; consent; purpose specification and secondary use restrictions; data minimization; onward transfer; and data security. With scraping, data protection laws built around

these requirements are ignored.


Scraping has evaded a reckoning with privacy law largely because scrapers act as if all publicly available data were free for the taking. But the public availability of scraped data shouldn’t give scrapers a free pass. Privacy law regularly protects publicly available data, and privacy principles are implicated even when personal data is accessible to others.


This Article explores the fundamental tension between scraping and privacy law. With the zealous pursuit and astronomical growth of AI, we are in the midst of what we call the “great scrape.” There must now be a great reconciliation."