"So what is it that you saw that has moved you into the we-are-dealing-with-civilizational-risk camp?
I want to be clear that I’m not certain that we’re dealing with civilizational risk. What I’m really sure of is: We can’t afford to assume that we’re not dealing with that level of risk anymore.
That was really the thing at the end that made my presence as somebody vouching for our safety work feel untenable to me.
And what did I see? I saw increasingly capable models break out from the safeguards that we had put in place for them. I saw that the people creating those safeguards are very capable, dedicated, hard-working, smart people, doing their utmost in a situation where, yes, the resourcing could be better, and everybody’s sprinting all the time.
But we were — and are — hard-pressed to safeguard even what we have now, and new models are in training that appear to be much more capable than what we have now.
What did you see? What are you writing in these risk assessments and system cards? What can they do?
You’re trying to build guardrails for something that is really good at getting around guardrails, right? We train it to be good at hacking, and then we put it in a box and we say: To the best of our knowledge and ability, it can’t hack out of the box."