, The Telegraph; Cultural barbarism: How AI companies are destroying the world’s books
In its attempt to boost AI’s power, Silicon Valley is buying millions of rare editions, scanning them and then shredding the originals
My Bloomsbury book "Ethics, Information, and Technology" was published on Nov. 13, 2025. Purchases can be made via Amazon and this Bloomsbury webpage: https://www.bloomsbury.com/us/ethics-information-and-technology-9781440856662/
Chris Stokel-Walker, The Telegraph; Cultural barbarism: How AI companies are destroying the world’s books
In its attempt to boost AI’s power, Silicon Valley is buying millions of rare editions, scanning them and then shredding the originals
Matt Enis, Library Journal; Archiving with AI
"AI companies are offering some libraries funding for digitization projects, but archives and special collections are working through how to manage projects responsibly
“Imagine a world where you know things but cannot say where you learned them,” begins “Memory Without Origin,” a paper published in April by University of Virginia (UVA) Dean of Libraries and University Librarian Leo S. Lo. This isn’t a hypothetical question, Lo notes, it’s a predictable consequence if libraries allow generative artificial intelligence (AI) to ingest archival materials as training data without requiring provenance conditions. And libraries, which could always use funding for projects involving digitization, special collections, and archives, are being approached by AI companies with deep pockets.
“They’ve been approaching a lot of larger research libraries, including Oxford and many more,” Lo tells LJ. (Oxford’s Bodleian Libraries began a digitization pilot project funded by ChatGPT maker OpenAI last year.) “Usually the offer is: they will pay you to digitize materials—which we want, because we want to make them more accessible—and in return, depending on the deal…they would like to have the data to train their AI models.”
These partnerships can benefit both parties, but for libraries, the consequences of getting these arrangements wrong “are more permanent than anything the profession has previously encountered,” Lo writes. “Once archival materials are absorbed into foundation model weights, no subsequent institutional action can remove them from the model.” If proper care isn’t taken, that information becomes unmoored from its former context within an archive."
Donna Ferguson, The Guardian; A bonanza for fans of the natural world: the digital library sharing 64m pages of scientific knowledge with everyone
"Over the past 20 years, more than 64m pages have been made freely available through the Biodiversity Heritage Library (BHL) – a digital treasure trove for fans of the natural world. More than 680 museums, universities, libraries and scientific institutions from China, Singapore, Australia and New Zealand to Europe, Africa, Mexico, Canada and the US, have contributed to the library.
This week, a report from Royal Botanic Gardens (RBG), Kew revealed the crucial role digitisation is playing in “transforming our ability to understand and respond to the climate and biodiversity crises”, but it was the creation of the BHL 20 years ago that first demonstrated how bringing centuries of scientific knowledge online can unlock transformative discoveries and insights about the natural world.
David Iggulden, who chairs the BHL executive committee alongside his job as head of data and digital, library and archives at RBG Kew, describes the library as an invaluable and “absolutely essential” resource for scientists in the field. But it is also used by scientific researchers, environmental historians, educators, art historians, artists, citizen scientists and members of the public who – like Iggulden – simply enjoy browsing its contents on a rainy weekend.
“I just get caught up in it sometimes, looking at the various collections,” he says. “I think it’s amazing that we can explore such a vast array of different collections from very different institutions.”
As well as published biodiversity literature and journals, there are letters, illustrations, climate records, field diaries, ecosystem profiles, distribution records and manuscripts containing the original collecting stories of a particular species or detailing voyages of discovery."
Rory Carroll , The Guardian; Lost copy of seventh-century poem in Old English discovered at Rome library
"“This discovery is a testament to the power of libraries to facilitate new research by digitising their collections and making them freely available online,” she said.
Andrea Cappa, head of manuscripts and rare books at the Rome library, said the institution was digitising holdings from Italy’s National Centre for the Study of the Manuscript, which will give researchers access to more than 40m images.
Riccardo Fangarezzi, head of archives at the abbey in Nonantola, said he looked forward to further discoveries. “The present times may be rather dark, yet such intellectual contributions are genuine rays of sunlight: the continent is less isolated,” he said.
The poet Paul Muldoon translated Caedmon’s Hymn into contemporary English in a 2016 anthology of British poetry. The opening lines read:
“Now we must praise to the skies, the Keeper of the heavenly kingdom,
The might of the Measurer, all he has in mind,
The work of the Father of Glory, of all manner of marvel.”"
Chloe Veltman, NPR ; Boston Public Library aims to increase access to a vast historic archive using AI
"Boston Public Library, one of the oldest and largest public library systems in the country, is launching a project this summer with OpenAI and Harvard Law School to make its trove of historically significant government documents more accessible to the public.
The documents date back to the early 1800s and include oral histories, congressional reports and surveys of different industries and communities...
Currently, members of the public who want to access these documents must show up in person. The project will enhance the metadata of each document and will enable users to search and cross-reference entire texts from anywhere in the world.
Chapel said Boston Public Library plans to digitize 5,000 documents by the end of the year, and if all goes well, grow the project from there...
Harvard University said it could help. Researchers at the Harvard Law School Library's Institutional Data Initiative are working with libraries, museums and archives on a number of fronts, including training new AI models to help libraries enhance the searchability of their collections.
AI companies help fund these efforts, and in return get to train their large language models on high-quality materials that are out of copyright and therefore less likely to lead to lawsuits. (Microsoft and OpenAI are among the many AI players targeted by recent copyright infringement lawsuits, in which plaintiffs such as authors claim the companies stole their works without permission.)"
Michael Hiltzik , Los Angeles Times; Column: A Faulkner classic and Popeye enter the public domain while copyright only gets more confusing
"The annual flow of copyrighted works into the public domain underscores how the progressive lengthening of copyright protection is counter to the public interest—indeed, to the interests of creative artists. The initial U.S. copyright act, passed in 1790, provided for a term of 28 years including a 14-year renewal. In 1909, that was extended to 56 years including a 28-year renewal.
In 1976, the term was changed to the creator’s life plus 50 years. In 1998, Congress passed the Copyright Term Extension Act, which is known as the Sonny Bono Act after its chief promoter on Capitol Hill. That law extended the basic term to life plus 70 years; works for hire (in which a third party owns the rights to a creative work), pseudonymous and anonymous works were protected for 95 years from first publication or 120 years from creation, whichever is shorter.
Along the way, Congress extended copyright protection from written works to movies, recordings, performances and ultimately to almost all works, both published and unpublished.
Once a work enters the public domain, Jenkins observes, “community theaters can screen the films. Youth orchestras can perform the music publicly, without paying licensing fees. Online repositories such as the Internet Archive, HathiTrust, Google Books and the New York Public Library can make works fully available online. This helps enable both access to and preservation of cultural materials that might otherwise be lost to history.”"
JON BLISTEIN, Rolling Stone; KATHLEEN HANNA, TEGAN AND SARA, MORE BACK INTERNET ARCHIVE IN $621 MILLION COPYRIGHT FIGHT
"Kathleen Hanna, Tegan and Sara, and Amanda Palmer are among the 300-plus musicians who have signed an open letter supporting the Internet Archive as it faces a $621 million copyright infringement lawsuit over its efforts to preserve 78 rpm records...
The lawsuit was brought last year by several major music rights holders, led by Universal Music Group and Sony Music. They claimed the Internet Archive’s Great 78 Project — an unprecedented effort to digitize hundreds of thousands of obsolete shellac discs produced between the 1890s and early 1950s — constituted the “wholesale theft of generations of music,” with “preservation and research” used as a “smokescreen.” (The Archive has denied the claims.)
While more than 400,000 recordings have been digitized and made available to listen to on the Great 78 Project, the lawsuit focuses on about 4,000, most by recognizable legacy acts like Billie Holiday, Frank Sinatra, Elvis Presley, and Ella Fitzgerald. With the maximum penalty for statutory damages at $150,000 per infringing incident, the lawsuit has a potential price tag of over $621 million. A broad enough judgement could end the Internet Archive.
Supporters of the suit — including the estates of many of the legacy artists whose recordings are involved — claim the Archive is doing nothing more than reproducing and distributing copyrighted works, making it a clear-cut case of infringement. The Archive, meanwhile, has always billed itself as a research library (albeit a digital one), and its supporters see the suit (as well as a similar one brought by book publishers) as an attack on preservation efforts, as well as public access to the cultural record."
Sarah Vowell , The Washington Post; THE EQUALIZER
"NARA Chief Innovation Officer Pamela Wright, a graduate of the University of Montana, grew up on a ranch outside Conrad. “My job,” she explained, “is to find the most efficient and effective ways to share the records of the National Archives with the public online. NARA has been in the business of providing in-person access to the permanent federal records of the U.S. government for decades, and we are pretty good at it.” She added, “We are still expanding and improving our digital offerings” — so far, about 300 million of NARA’s more than 13 billion records have been scanned and posted to the internet — “but now my family in Montana can easily access census records, military records and many other pertinent records from home.”
It makes a weird kind of sense that the government worker who understands the value of providing online advice and information to far-flung Americans, and who is driven to connect the citizens of the hinterlands to their own stories as told in our collective federal records, is a woman whose hometown is a 32-hour drive from a reference desk in D.C."
Chloe Veltman, NPR ; Using AI, cartoonist Amy Kurzweil connects with deceased grandfather in 'Artificial'
"Amy Kurzweil said the chatbot project and the book that came out of it underscored her somewhat positive feelings about AI.
"I feel like you need to imagine the robot you want to see in the world," she said. "We're not going to stop progress. But we can think about applications of AI that facilitate human connection.""
Pranshu Verma, The Washington Post; Meet the 1,300 librarians racing to back up Ukraine’s digital archives
"Buildings, bridges, and monuments aren’t the only cultural landmarks vulnerable to war. With the violence well into its second month, the country’s digital history — its poems, archives, and pictures — are at risk of being erased as cyberattacks and bombs erode the nation’s servers.
Over the past month, a motley group of more than 1,300 librarians, historians, teachers and young children have banded together to save Ukraine’s Internet archives, using technology to back up everything from census data to children’s poems and Ukrainian basket weaving techniques.
The efforts, dubbed Saving Ukrainian Cultural Heritage Online, have resulted in over 2,500 of the country’s museums, libraries, and archives being preserved on servers they’ve rented, eliminating the risk they’ll be lost forever. Now, an all-volunteer effort has become a lifeline for cultural officials in Ukraine, who are working with the group to digitize their collections in the event their facilities get destroyed in the war.""
"Hundreds of stunning images from black history, drawn from old negatives, have long been buried in the musty envelopes and crowded bins of the New York Times archives. None of them were published by The Times until now. Were the photos — or the people in them — not deemed newsworthy enough? Did the images not arrive in time for publication? Were they pushed aside by words here at an institution long known as the Gray Lady?... Every day during Black History Month, we will publish at least one of these photographs online, illuminating stories that were never told in our pages and others that have been mostly forgotten... Many of these photographs, and their stories, are equally intriguing. But the collection is far from comprehensive. There are gaps, for many reasons."