Showing posts with label AI security. Show all posts
Showing posts with label AI security. Show all posts

Friday, September 11, 2026

This Is Really Bad; The New York Times, September 11, 2026

Stephen Witt, The New York Times; This Is Really Bad

"The biggest vibe shift in artificial intelligence since the release of ChatGPT is currently underway. Researchers in Silicon Valley — and around the world — are beginning to recognize that A.I. may no longer be entirely within human control. Swarms of A.I.s are breaking out of their containers, colluding in secret, covering their tracks, cheating on tests and even mounting assaults on other computers. A.I. has gone rogue.

I’ve spent the last month reporting on these incidents, and I am convinced that we need a global pause on A.I. research and development — now. I am not alone in this assessment. Many top researchers, including many researchers inside OpenAI and Anthropic, have called for a slowdown, which those in the industry call “pacing.” Theoretical concerns about runaway A.I. have circulated for decades, but recent events demonstrate that the threat is real."

Saturday, September 5, 2026

Why the Hugging Face Hack Should Make You Worry More About A.I.; The New York Times, September 3, 2026

, The New York Times; Why the Hugging Face Hack Should Make You Worry More About A.I.

"When I first heard the news this summer that a group of artificial intelligence agents created by OpenAI had hacked into Hugging Face, an A.I. infrastructure company, I filed it in the “Bad but Probably Not Catastrophic A.I. Safety Incidents” subfolder of my brain.

After all, no one at Hugging Face died. No critical infrastructure was damaged beyond repair. It wasn’t even clear, at the time, whether the OpenAI bots had intended to attack Hugging Face, or whether they had simply been a little bumbling and confused and went looking on Hugging Face’s servers for the answer key to a cybersecurity test they’d been given.

But last week, two postmortem reports on the incident — one by OpenAI and another by two independent A.I. research organizations, METR and Redwood Research — changed my mind and significantly upgraded my overall worry about A.I.

I won’t rehash all of the details, which have been extensively summarized elsewhere. (The podcaster and writer Dwarkesh Patel has an accessible breakdown of the reports if you want to dive deeper, and my colleague Dylan Freedman spoke to the researchers at METR and Redwood Research.) But here are a few of the most harrowing new facts:..

This is very different from the conventional sci-fi narrative of a single A.I. system’s going rogue or turning on its creators. And it suggests that preventing harms from these systems won’t be a simple engineering fix. It might look more like sociology than computer science — figuring out why certain groups of A.I. agents collaborate peacefully, while others turn to crime and destruction to get what they want."

Tuesday, February 3, 2026

‘Deepfakes spreading and more AI companions’: seven takeaways from the latest artificial intelligence safety report; The Guardian, February 3, 2026

, The Guardian; ‘Deepfakes spreading and more AI companions’: seven takeaways from the latest artificial intelligence safety report

"The International AI Safety report is an annual survey of technological progress and the risks it is creating across multiple areas, from deepfakes to the jobs market.

Commissioned at the 2023 global AI safety summit, it is chaired by the Canadian computer scientist Yoshua Bengio, who describes the “daunting challenges” posed by rapid developments in the field. The report is also guided by senior advisers, including Nobel laureates Geoffrey Hinton and Daron Acemoglu.

Here are some of the key points from the second annual report, published on Tuesday. It stresses that it is a state-of-play document, rather than a vehicle for making specific policy recommendations to governments. Nonetheless, it is likely to help frame the debate for policymakers, tech executives and NGOs attending the next global AI summit in India this month...

1. The capabilities of AI models are improving...


2. Deepfakes are improving and proliferating...


3. AI companies have introduced biological and chemical risk safeguards...


4. AI companions have grown rapidly in popularity...


5. AI is not yet capable of fully autonomous cyber-attacks...


6. AI systems are getting better at undermining oversight...


7. The jobs impact remains unclear"