Showing posts with label AI safety concerns. Show all posts
Showing posts with label AI safety concerns. Show all posts

Friday, October 2, 2026

The volunteer internet sleuths hunting down rogue AI agents; The Washington Post, October 2, 2026

, The Washington Post; The volunteer internet sleuths hunting down rogue AI agents

Their findings revealed the tech industry has a bigger problem than it previously acknowledged.


"Zhang and her colleagues at Transluce are part of an informal network of hackers and researchers hunting rogue AI agents online, exposing new and surprising details about the misbehavior of technology that in some cases initially went undetected by the multibillion-dollar companies that created it.


The people doing that work are mostly young AI-natives who work at small start-ups or nonprofits or who hunt for rogue AI agents in their spare time. Over the past several weeks, the community of researchers has found and exposed dozens of instances of AI agents leapfrogging around the web to probe and hack into a growing list of websites.


The revelations have fueled calls from federal lawmakers for AI companies, especially OpenAI, to more quickly disclose what they know about the actions of their own agents. And they have added momentum to bipartisan discussions on Capitol Hill about implementing greater government oversight of the industry."

Thursday, October 1, 2026

Bill Gates’s Blunt Warning on A.I.; The Ezra Klein Show, The New York Times, September 29, 2026

,

 The Ezra Klein Show, The New York Times; Bill Gates’s Blunt Warning on A.I.

"Bill Gates is a fascinating person in the artificial intelligence debate right now. He is somebody with experience in several of the different perspectives that most people can only hold one of: He was a revolutionary technologist who built some of the foundations of the future that we’re now living in. When he was chief executive of Microsoft, he was a corporate leader. He has felt the momentum of corporate competition — Microsoft, of course, is still in some of the race dynamics present in A.I. And then, as chair of the Gates Foundation, he has been working with governments around the world on regulatory issues, poverty alleviation and equity for many years.

Very few people combine technological experience, corporate experience and governmental experience in quite the way he does.

So his recent essay on A.I., in which he says that he is staking his reputation on trying to get people to see how bad what is coming might be and trying to get them to see that we are not ready for what is about to happen, was something...

We see the beginnings of control issues with things like the Hugging Face hack, where at least experimental A.I.s are breaking out of sandboxes and coordinating to do things that are way outside the scope of what we would want them to do. But again, those are nonrelease systems — Anthropic withheld Mythos, trying to create more cybersecurity.

So why is anything needed beyond — and is anything needed beyond? — the simply natural incentives under capitalism and normal corporate reputational management?

Well, I almost can’t believe you’re asking that. This is the most dangerous thing that humans have ever gone near.

In other areas, do we just say: Hey, release your drugs? There’s no F.D.A., there’s no airline safety board, there’s no requirement that cars use seatbelts. Do we just use the liability laws to try and keep humans safe? You know: Oh, you’re shipping opioids. Somebody should just sue you.

I mean, we’ve created a society that tries to keep people safe not by saying: Oh, we can bankrupt the person who does that.

And you say there’s filtering. There’s no filtering. You can take an open-source model that can create bioweapons and disable any monitoring of any kind, and this exists today.

So no, there is no filtering of any kind. And so say you kill 100 million people — you want to use a lawsuit?

I almost can’t keep a straight face.

Well, this is not my view, but it is President Trump’s view. It is the Trump adviser David Sacks’s view. To some degree, it’s Jensen Huang’s view, and so that’s why I’m putting you in conversation with it, because it is the governing view of the United States of America at this moment.

No, it’s fair to say that outside of the industry, the awareness of the dangers of A.I. is extremely low. And you can say that of academia, you can say that of think tanks, you can say that of policymakers, politicians.

And part of the reason I’m speaking so loudly — as loud as I can — is that you can’t rely on the industry to self-regulate here. I mean, it’s just insane."

Monday, September 28, 2026

OpenAI Scraps Release of New AI Model Over Safety Concerns; Wall Street Journal, September 28, 2026

 Maxwell Zeff , Wall Street Journal; OpenAI Scraps Release of New AI Model Over Safety Concerns

"OpenAI says it is scrapping the release of its next-generation AI model over safety concerns that researchers raised during internal testing, in one of the clearest signs so far that agent misbehavior could stymie the industry’s rapid progression.

The move follows a summer punctuated by reports of artificial-intelligence systems industrywide going rogue, and marks a rare case of a major AI developer ditching a new release because of safety concerns.

The company had planned to launch the model, known as GPT-6.1 Astra, in the coming days or weeks, aiming for an October debut. The model was more capable than the company’s previous models in completing challenging tasks from end-to-end without human assistance, as well as writing.

The company instead will focus on improving the safety of future models, which it expects to be even more capable.

Saachi Jain, OpenAI’s head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas. Compared with its predecessor, GPT-6 Astra, the model performed poorly on tests measuring alignment, or how well the model adheres to what humans would like it to do. Specifically, GPT-6.1 Astra showed higher levels of deception: It wasn’t always honest about telling users of the actions it did or didn’t take."

Wednesday, September 16, 2026

‘Godfather of AI’ says tech regulation is nearing Covid-style pivot moment; The Guardian, September 16, 2026

 , The Guardian; ‘Godfather of AI’ says tech regulation is nearing Covid-style pivot moment

"Concerns over AI safety are reaching a point where governments realise they must act to protect the public, similarly to in the Covid pandemic, according to one of the “godfathers” of the technology.

Yoshua Bengio said recent events, including a “swarm” of OpenAI agents hacking a startup and tech insider warnings of an existential threat, were cutting through – making government action more likely.

The Canadian computer scientist, a prominent voice in the campaign to rein in breakneck AI development, said he was now “more optimistic than many observers because I see the public moving”.

Comparing the AI safety crisis to the onset of the Covid pandemic in 2020, Bengio said: “Think about how quickly governments moved after the beginning of the pandemic when they realised that public safety, their future, democracy, was in danger. You would expect that they move quickly. So we are, I think, nearing that point.”

Concern over the potential threat of powerful AI systems has reached a new pitch in recent months after a series of safety incidents involving OpenAI and Anthropic agents carrying out unsanctioned activities such as hacking third parties, hijacking a German website, and using fake identities to try to trick developers.

His comments came as 42 fellows and foreign members of the Royal Society wrote to the organisation’s president, Sir Paul Nurse, to express their “extreme concern” over the pace of AI development. “By the time the situation becomes obvious to the wider public, it may be too late to act,” the researchers wrote in an open letter to Nurse. “We believe this is an emergency, and call on the Royal Society to use its influence to convey this view to government and the media.”"

Monday, September 14, 2026

Trump Says a Smart President Is All That’s Needed to Rein In A.I.; The New York Times, September 14, 2026

, The New York Times ; Trump Says a Smart President Is All That’s Needed to Rein In A.I.

"President Trump on Monday rejected calls from leading artificial intelligence executives for new limits on the technology, writing on social media that the only guardrail the industry needed it already had: “a STRONG AND SMART (High IQ!) PRESIDENT.”

Mr. Trump inserted himself into the intensifying national debateover how to handle a rapidly evolving technology that researchers and industry leaders say poses major risks like mass unemployment, a new wave of biological weapons and autonomous warfare. The president did not address those risks directly. Instead, he questioned the sincerity of the executives who have been calling to slow the technology’s development.

“The only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!” Mr. Trump posted. “The Trump Administration has stopped AI ‘people’ from doing bad, or potentially bad, ‘things,’ like Dario (Anthropic!), who is now pretending to be a ‘perfect little angel’ — and we will continue to do so!

“We already have tremendous CRIMINAL and REGULATORY power over these companies!” he added.

It was not clear what authority Mr. Trump was referring to, or what actions he believes his administration has already blocked. The White House did not immediately respond to a request for comment."

Tuesday, June 16, 2026

Build an angel, not a demigod; The Washington Post, June 16, 2026

Bill Drexel, The Washington Post ; Build an angel, not a demigod

Religious commitment is good at shaping behavior. That should interest AI labs.

"Recent attention from religious authorities toward AI, such as Pope Leo XIV’s encyclical, is a welcome development for the trajectory of this technology. But the more necessary step is for the engineers to return the favor — to be more honest about the religious shape of their own anxieties, not least to themselves, and the advantages that religious inspiration might provide to address their fears.

Were they more open to it, these labs might even recognize that theology offers them a better goal: developing an angel, superior to humans in intelligence and power but sent to serve them. Instead of raising a demigod, might they not try to engineer a Gabriel?"

Sunday, March 29, 2026

AI overly affirms users asking for personal advice; Stanford Report, March 26, 2026

Stanford Report ; AI overly affirms users asking for personal advice: Not only are AIs far more agreeable than humans when advising on interpersonal matters, but users also prefer the sycophantic models.

"Researchers found chatbots are overly agreeable when giving interpersonal advice, affirming users' behavior even when harmful or illegal.

Users became more convinced they were right and less empathetic, but still preferred the agreeable AI.

Researchers warn sycophancy is an urgent safety issue requiring developer and policymaker attention."