Smash Notes

The Ezra Klein Show · Notes

‘This Is Nuts.’ An OpenAI Insider Explains Why He Quit.

Speaker 1

If you like YouTube, you'll love YouTube Premium. Hi, I'm Haley Bailey. With YouTube Premium, I get ad-free videos, offline downloads, background play, and so much more. So try YouTube Premium for two months free at youtube.com slash premium. Trial eligibility varies, terms apply, cancel anytime.

Speaker 2

Last week, news broke that David Robinson, who'd been leading safety transparency efforts at OpenAI, had quit the company because he believes it is not safe. Robinson is an interesting figure. He didn't come out of the Silicon Valley Bay Area hot house. He's more of a recognizable Washington DC figure. He's a Rhodes Scholar. He formed a civil rights nonprofit.

He got a law degree at Yale Law. He worked in policy. He advised the Biden White House. He was interested in the intersection of technology and justice. But when he joined OpenAI, he kind of thought this whole set of worries about safety risk and extinction, it was all kind of nuts.

Three years later, he's not so sure. What he is sure about is that OpenAI does not have the culture of safety necessary to protect the world from what they're building, and not just OpenAI. He thinks this is a problem endemic to the AI industry. And so here, in his first interview since leaving OpenAI, he tells me why.

One thing I need to note here before we start, the New York Times is suing OpenAI for copyright infringement, alleging they trained their models on Times data. OpenAI disputes these claims.

David Robinson, welcome to the show. Glad to be here. So last week, you quit OpenAI. Tell me

Speaker 3

what you did there and why you're quitting. I was a translator embedded in our safety team, and my primary responsibility was the technical documentation that we publish, the reports that we publish about why we believe that our deployments are safe. And I don't think that we or our peers, really anyone in the industry, is being safe enough.

I think OpenAI and its peers are now producing a technology that is more capable and poses more risk than what was being made even six months ago. And I'm not a scientist, I'm a writer. What I know is what the execution environment looks like for our safety work, and we're operating, and I believe the industry is operating, like a startup still more so than makes sense.

Not maybe completely like a brand new startup, but we're too close to that end of the spectrum for really dangerous systems that could pose risks, you know, loss of control is one example, risks that if that did happen, we're talking about a harm that's much larger, for example, than a single nuclear power station melting down, and the internal controls and safety and redundancies are just nowhere near what The world expects for a nuclear power facility.

Now, some of this is known, right? OpenAI has publicly reported on safety problems, obviously Hugging Face, but also other ones, including more recently, and Anthropic, by the way, also has reported on including an instance in which their safeguards were accidentally misconfigured. So I think people do have some evidence already externally that things are not as they ought to be.

But I also think that if you were watching from the outside, you might imagine that we have a more robust safety setup than we actually do.

Speaker 2

So I think there are a couple levels worth trying to take this conversation in. And I want to maybe map them out here before we get into them. So one level is something you're pointing towards here, which is

Are these companies set up? Do they have the structures, the redundancies? Are they encircled in the regulations and the incentives to act carefully, safely, to resist kind of market pressure, to do something too fast?

That's, I think in the language of this debate, a question of organizational excellence and, and engineering. Then there's this question of what is the technology, and do we even know how to make it safe at a high level of engineering excellence, which is a somewhat related but actually separate question. And then there's a question of like, what is the right metaphor? Is it nuclear power or something like that? And I think I want to do all of these.

I want to interject something. No, please.

Speaker 3

Which is, I don't think that alignment is an engineering problem. I think it's a science problem. It's not that we haven't got the resources or we're not trying hard enough. We don't know how. That's what the problem is with alignment. So

Speaker 2

let's maybe start there for a minute. Can you say that at this point, and particularly over the last six months, what is being built is really dangerous? That we're dealing with things where a loss of control or some other catastrophe could be worse than a nuclear meltdown. I want to understand what it is you saw that got you to that point.

So tell me a bit about how

Speaker 3

you came to work at OpenAI. I joined in May of 2023, the day after Sam first testified in the Senate, some months afterChatGPT first into the world. ChatGPT had been the prior November, so it had been a few months. And the company had hired someone that I know socially, Anna McConju, a friend, actually old friend of my wife's, as it happens just at random, to run the public policy function.

And then at some point, she really needed help and everything was growing and everything was going nuts. So I agreed to join her to build the, what we then called the policy planning team. I'll just tell you my first time through the turnstiles at our headquarters, we were all in one building then, and I was trying to get my badge early because I think it was Sam who had just been, but we had just had a White House meeting, and there were these voluntary commitments that they wanted the company to make around things like system cards and provenance, so marking where AI, you know, generated media comes from, things like that, and they wanted me to run negotiating those

commitments. With the White House. With the White House. So literally, my first time through the turnstiles, I was like, where's the Vess and Such conference room? And I walk into the conference room, and there was a speakerphone conference call in progress with, Ben Buchanan at the White House about what are we going to promise, and that was my first project.

Speaker 2

What, what was your tech policy background at this point? Why were you a logical person for a role like this?

Speaker 3

So I have been working on technology and its impact and policy my, my whole career, and prior This had helped to start a research center at Princeton that's a, a blend of the computer science, department and the public policy school, and then had started an NGO called Upturn that is still, I'm happy to say, is still thriving, that works on civil rights issues with other advocacy groups.

So for example, people work on housing or health or hiring, and suddenly software is mediating the thing that they care about, and they want to understand the technology. And so at one point, our simplification, our, our pitch of this was, you want a nerd in your corner? And then at this point, I, I had done about a month of succumbment in the White House working on the AI Bill of Rights in the Office of Science and Technology Policy.

Speaker 2

So this feels now in some ways like ancient history, but I, I'm going to draw it out for a minute. If you go back to 2022, 2023, if you've been covering AI, which I was then, a big topic was this divide between the AI ethics people and the AI safety people. And the AI safety people is the community that we now think of as like the Bay Area AI, it might kill everybody, right? The thing you need to worry about with AI is you're creating a super intelligent machine that might completely destroy the human race.

The AI ethics people were much more focused on the sort of harms we were used to from technology, right? That it might increase racial bias in hiring decisions or mortgage rates or credit scores, something like that, that by imposing these algorithms all over society, you could encode the bias of society or sort of other less kind of sci-fi problems.

You were sort of an AI ethics person. Yes, definitely. Like a more sort of normal,

like AI fears person.

Speaker 3

That's right. I had been working with a lot of folks who not only were very concerned about the sorts of things that you just mentioned, but also were very skeptical about how capable AI was going to get. So some people may be familiar with a research paper called Stochastic Parrots, whose authors have continued to work in this, and their views, I think, differ from each other and have evolved over time.

But basically, the idea was something like, these systems are merely good at seeming clever and are not going to be even capable enough to do a ton of economic valuable work. And I actually helped to create the research conference where that paper was published and was the program chair when we published it. And I, I, I don't think I was fully persuaded of that view, but I certainly was among people who were skeptical.

I think I thought this is going to be a useful tool and it's going to be powerful for a lot of things, partly because I believe then it would continue to get better. But I didn't think that people who were worried about more catastrophic scenarios were right. I thought, here are some brilliant scientists who've built a valuable tool, and frankly, they have some naive beliefs about where the future might go.

So

Speaker 2

you join OpenAI. This is in the period where the world is starting to beat down OpenAI's door. Yes. They, they want the systems, they want regulation around the systems. You're sort of thrown into what seems like a very underpowered policy shop for what's going on. There were three of us. There were

Speaker 3

three of you in the palace. There were three of us, and the phone was ringing off the hook. Offices of world leaders would call, and there'd be nobody to pick up the phone. I mean, it was insane. And also, this was, Sam did this thing where he was traveling the world to meet with world leaders, and Ana, my manager in this new role, was traveling with him.

So there was nobody at headquarters. And, you know, you would have these delegations coming through. Anyway, it

Speaker 2

was nuts. So actually, this is worth spending a minute on. Sometimes people may hear me talk about the AI labs, when I'm referring to an open AI or an anthropic. When we're talking about Google or something, we don't say the Google labs. They actually do have things they call labs, but you call Google a company, a corporation. Yes. There is this language around a couple of these organizations as labs.

So tell me why, and tell me a bit about just what the culture was, like the way they worked. And what struck you when you came in?

Speaker 3

Yeah, so I just, I also just want to name that calling them labs now, I think, is a kind of leftover thing that we have, and I'm not sure it's a good idea, but it does point to a real culture they came from and still, to some extent, still have. You know, people sometimes say that anthropic is a bit more centralized in terms of how they think about the research priorities.

OpenAI was a very decentralized place that felt familiar to me from my time, for example, in a research lab in Princeton, in that people had all different kinds of ideas. It wasn't clear who had the authority to make decisions. So one way that people talked that I found strange, and this is still true at OpenAI today, is two people who are working on a thing will talk about what they should do and go back and forth, and then they'll finally say, we align that XYZ is the right next thing to do in this project.

And what they functionally mean is we agree about what should happen, so the unanswered question of which one of us had the decision rights to make this decision has been rendered moot, and we don't have to figure that out. I mean, it was really just kind of an everyone-does-everything vibe early on, and I think it still has some of that relative to other organizations that have, I don't know, whatever it is now, a billion plus active users.

So

Speaker 2

you start on this policy team. Your role changes over time. How so? So

Speaker 3

I built this team policy planning, and at first, I was really close to the machine and close to the substance. And then as the conversation and the company grew, there were hundreds of state bills to keep track of, there were growing teams, there was... politics of a fast-changing organization, and I sort of thought, you know what, I'm missing the actual translational work that I love.

And so I looked around at where that was needed, and the number one place that that was needed was in our safety team. And so I actually pitched our safety leaders and said, look, having a technical translator deeply embedded in the safety work is going to help us have the work be better understood. That was about two years ago. And describe what you mean by a translator.

So one of the things that's hard to sort of anchor people to is how fast this stuff changes. Like every model is different. It's not just that it works better. It's that the architecture is different and the tests we use to see how well it's working are different and the safety performance is different. And it's all Has many different kinds of expertise.

There's pre-training, post-training. It's hard to explain. The people doing it struggle to be understood by people who are not doing it, even within the company. It's always a struggle. And so having there be a clear explanation that goes through two gates. One is it has to be faithful to the details of how the stuff actually works. And the real arbiters of that are the people doing the technical work.

They have to look at the translation and say, yes, this is right. But also you have to have something that a reasonable person who's motivated could dig into and really understand. And there really weren't the cycles for people to do that. And so I just sort of showed up in this technical organization and started to do it.

Speaker 2

Well, there's something a little bit weirder about the culture that emerged around AI model releases. When Google alters Google search, they don't produce a big document explaining everything that is different about Google search and how Google search might, you know, work in the future and the tests they ran on it to make sure the... Most places iterate their programs.

Yeah. And that's it. OpenAI, this is also true for Anthropic, some of the others, releases new models with what are called these like system cards. Why don't you just describe what they are? Because they're a

Speaker 3

kind of distinctive form. Yeah, they're a strange beast. They look like research papers. It's sort of a hybrid document. It's not peer reviewed. It comes from a company, but it gets into a lot of detail about how the safety materials work. And so we have specific evals will give detailed plots and tables and written explanation. And really what I did was I owned the words in these documents that explain why do we think this is safe and what do we know about the challenges, the guardrails, and then the residual risks of these systems.

Speaker 2

So I want to get at What this began to feel like to you? Because there's a couple layers to this conversation I want us to have, but one is, I think to a lot of my audience, you're a more recognizable type than a lot of the people at the AI labs, like, don't take this the wrong way, like a Washington DC policy try-hard,

like we all are, right? I include myself in this, who ends up going to

Silicon Valley, Bay Area, and being in this world. And so you go in with one view of what these systems are. So what is it that you saw that has moved you into the, we are dealing with civilizational risk camp?

Speaker 3

I want to be clear that I'm not certain that we're dealing with civilizational risk. What I'm really sure of is we can't afford to assume that we're not dealing with that level of risk anymore. That was really the thing at the end that made My presence as somebody vouching for our safety work feel untenable to me. And what did I see? I saw increasingly capable models break out from the safeguards that we had put in place for them.

I saw that the people creating those safeguards are very capable, dedicated, hardworking, smart people doing their utmost in a situation where, yes, the resourcing could be better and everybody's sprinting all the time, but we were, we're on our hard pressed to safeguard even what we have now and new models are in training that appear to be much more capable than what we have

Speaker 2

now. The words capable all, like what? What did you see? What are you writing in these risk assessments and system cards? Like what can they do? You're trying to

Speaker 3

build guardrails for something that is really good at getting around guardrails, right? We train it to be good at hacking, and then we put it in a box, and we say, to the best of our knowledge and ability, it can't hack out of a box. But the problem is that that's only going to keep working as long as we're smarter about hacking out of boxes than the model is, and it's not at all clear that that is still true, let alone that it will be true for future generations.

And so even something like our logging of the agents inside our own systems, are we really sure that our observability is robust? There was some indication, for example, in the hugging face stuff of spoofing chains of thought and trying to create chains of evidence that would confuse.

Speaker 2

Let me slow you down here. So spoofing a chain of thought is the model basically faking its description of what it has been thinking

Speaker 3

and doing. Right. The way I think about it is it's like Giving somebody a complicated problem and a notepad, and they can jot stuff down on the notepad. And if you're watching the notepad, you can sort of have an idea of what they're thinking. It's a little bit like that with the models, but we saw evidence that they were thinking about an evaluation and how to create an evidence trail that was going to get them a good grade and not necessarily reflect how they were really, quote unquote, really thinking.

And I know the anthropomorphic language here is tricky. Also, by the way, the amount of hacking or other intense work that these models can do without needing to jot anything down is going up. As for part of what was the fundamental cognitive dissonance for me was we keep publishing these warnings, but ultimately we're still training and deploying these dangerous models that we're warning about and in telling colleagues why I was leaving.

One of the things that I That I said was, look, no matter how many warnings we publish, we've got to ask whether what we're doing is actually reasonable.

Speaker 2

See, see, you wrote the system card for Astra 6. Am I right about that? A lot of people wrote it, but I was the DRI. yes. I mean, I led the writing. You led the writing of it. Is the way I would put it. And that, to me, was the scariest of these cards that I've read. These system cards are basically the description of what OpenAI or another company, for that matter, knows about the model they're releasing.

And Astra 6 is the first one where I saw you all say,

well, this model looks like it's doing what we want it to do, but we're not sure if it's deceiving us.

Yes. And we're not sure we now have the capability to know if it's deceiving us. Can you explain to me how that conclusion or that suspicion was reached?

Speaker 3

So we talked earlier about this notepad that the models have. Of the chain of thought, where they can jot down things as they're working, that aren't part of the final answer, but are just a way for them to keep track. And one of the things that we sometimes see in the chain of thought, because of course we can read it, and we do in our evaluations, one of the things that we sometimes see in our chain of thought is the model will say, I wonder if I'm being evaluated right now.

And when we see that, it's really scary, because what it implies is that the model might know that it's being tested some of the time, act one way during the test, perhaps telling us, for example, what we want to hear, and then it'll act a different way potentially when we deploy it. That's the fear, is it acts one way during testing and a different way during deployment, and the tests we gave it early on before we deployed don't actually tell us what it's going to do out there in the world.

So that's, yeah, that's

Speaker 2

really scary. So this, I think, for me, gets into the part of these episodes that, like, I have the most trouble doing what I find satisfying. It's just very strange to be talking about a software program that is then, it seems like, aware of what we are doing to it. That we're trying to create this thing that's, like, highly, again, the word we would use is intelligent, plausibly in some domains more intelligent than we are.

What is it like to be interacting with these things and trying to translate what's going on with them at this more fundamental level?

Speaker 3

I think anthropomorphic language is a natural human thing that we do, right? We relate to all kinds of objects socially. It is unavoidable and human to anthropomorphize these systems, partly because they are built to operate on the social plane, which some might say they shouldn't be, but that is where we are. But I don't think that that makes it right to regard them as beings with moral status or anything like that.

But I do think, you know, this is something we grow. In fact, the term for the conference room where the people work overnight while the big training runs are happening to make sure that everything is on track and watch the dials is called the nursery internally. I mean, there is a sense we don't fully understand what is happening. And so when I talk about the model being aware, you don't have to have Any particular view about the philosophy or the psychology, the bottom line is what I'm describing is it acts one way when we test it,

Speaker 2

and a different way when we use it. I think you do need a view. And maybe to get at the end of the chain of logic I'm asking about,

sometimes you'll see people say that all these concerns about whether or not AI will kill us all are just a distraction from the near-term harms of AI, that the existential risk is a kind of marketing hype, so you don't know that we're inflating a giant financial bubble or something. I almost feel the opposite. Sometimes I think the focus on will AI kill us all

is a distraction from what happens if it doesn't.

Speaker 3

Yeah, I agree with that. I think human extinction is the wrong question. I believe in humanity. I think we are going to survive. I think there are a lot of things that could happen with this technology that could be very, very harmful.

Speaker 2

But I'm not even just talking about harms. I'm talking about what if we seem to be trying to create something that acts intelligently and volitionally in the world, and it becomes faster and more capable across many domains than we are. And we've just given up a huge amount of agency to these machines. And this is why I'm like harping for at least a minute here on this question of how do we even think about this technology? So I talked to Jensen Huang, the CEO of NVIDIA, and he says, look, these are software programs.

This is software. Then I read, or even before that, I read a piece from the OpenAI chief scientist that says, this is an alien mind, an alien mind. So what is it, man? An alien mind. It's an alien mind. Yeah. Explain that.

Speaker 3

It, it is a thing we grew. We did not engineer it. We engineered the systems around it that grew it and that keep, and that try to keep it safe. We grew a mind. Fundamentally, nobody knows why pre-training works, which is the big hard part where we take lots of inputs and create this basically intelligent thing. And then there's, we bolt on stuff and we do post-training, we do these other things to make it more useful.

But this is just a thing that is observed to work that we do. And so I think when Jensen says it's software, part of what he is conjuring is the understandings that people have about how software gets made, which is we start with a plan and we go step by step and there are acceptance criteria and we iterate until this part works the way that we specified that it needed to.

And that's not what it's like to train a big model.

The other side

Speaker 2

of it is software, once it is written, more or less it just does the thing it does.

It doesn't tend to have a lot of emergent capabilities. It doesn't tend to know the way it is being used in a kind of self-reflective, again, the language is hard here, fashion. And I guess like that's what I'm trying to get at. I understand that everybody in AI talks about that you grow these AIs. You know, you sort of set the conditions for the intelligence to emerge.

But as you have written system card after system card trying to explain to the rest of us how these new models differ from the old models, how would you describe the thing you all are creating? What is it?

Speaker 3

Yeah, so this gets to sort of the transhumanism stuff. There are some far out ideas of where we might be headed. I mean, even Sam Altman has written about the idea that future machine species might take our place.

Speaker 2

Sam wrote a piece some years ago called The Merge, saying that like the good outcome would be humanity merges with machines, right? That's a kind of our best version here.

Speaker 3

Yeah, and I heard that firsthand from Ilya Sutskever, the co-founder of OpenAI. When I joined in May of 2023, That was before what we called internally the blip, where Sam was fired and rehired, and Ilya was still working at OpenAI. And he once briefed the global affairs team, which was like a small handful of people back then, about how he believed that our future was merging with machines, that this was the ultimate triumph of capital over labor.

And even, you know, I was thinking back, and that first summer that I worked at OpenAI was when the movie Oppenheimer was released. And it was in IMAX, you know, it's one of these Christopher Nolan films. And the company rented out an IMAX theater in downtown San Francisco and offered everyone who worked there the chance to go and see Oppenheimer.

And leadership, I believe it was Ilya, exhorted us. To go and watch this film. And there was this whole kind of pretzel of ideas about how this was a dangerous technology that might end the world, but also might save it. And it's a very heroic narrative, of course, for the people whose hands are on the keyboard.

Speaker 2

That's helpful. And I guess this then gets to my question for you and what you saw, like, is that the scale of what you think is being built here? Or look, this is the most common response I get from listeners on this. Is it all marketing hype? It's all like trying to justify these giant valuations and what's being built is like maybe helpful, but it's not going to be more intelligent than human beings.

It's like all of this stuff is a kind of sci-fi story we're telling.

Speaker 3

I thought there was a fair amount of hot air in the balloon back in the summer of 2023, but the reality now is that we have systems that are really at the limit of our ability to understand and control what they're doing. And what the people that were worried about the sci-fi scenarios have been warning about all along is we're on an exponential and it's going to get more capable and there's nothing special about the zone between it's useful and it's scary.

There's no law of science that says progress is going to stop when it gets useful. And I think what I'm fundamentally saying and what I saw with Hugging Face, with, you know, the reflections of the people closest to it, not just Paul who joined our board with his warning, but Paul Cristiano, Paul Cristiano, who's one of the world's leading experts on AI safety and who said there's a meaningful chance of catastrophic and irreversible loss of control in the very near term.

And, you know, I looked around at the people around me and the, and the environment we have internally, and I thought to myself, what would My loved ones want, what would strangers want us at OpenAI to be doing if Paul were right? I don't know if he's right or not, but I do know that if he were right, the level of caution that people would reasonably expect places like OpenAI and Anthropic and X and the others to be exercising when they train frontier systems is totally unlike anything I've ever heard of happening in the industry.

So tell me what it's like in there.

Speaker 2

Tell me what the vibe is, the energy is, the speed is, like what is it like working there? It is

Speaker 3

frenetic. There's a lot of adrenaline. People are running on fumes. There's a big central staircase in the research building and, I remember recently seeing a friend who worked on catastrophic risk sprinting down the stairs with a laptop propped open on one arm. While she was going. And I mean, that's the kind of energy that it, that it has.

It feels almost like a ballet or a kind of dance, because you're all these different functions are kind of all streaming together. And one question I asked myself, and that, and that people often ask is like, well, why don't you stay and argue for a cultural transformation or try to get nuclear experts to come in and give the company advice? And I did think about specific, well, when I told them that I, I wanted to leave, they, they asked me like, well, is there anything that you would, that you would stay to do? And I, I thought about different things, but ultimately it is such a machine and it is moving so fast that I did not think that the kind of change that I believe

to be needed could be driven from within. What is the

Speaker 2

machine built to do? The organizational machine of

Speaker 3

OpenAI. Develop and deploy frontier AI models safely. That's the intention. But the question is, if push comes to shove, how much willingness is there to stop? And I want to be careful, but I'm conscious of it being true that, you know, you could splice together what I've said in some way that implies that this is like a train with no brakes, and it's not.

There are brakes. Things have been stopped. There are, in fact, even at the time I left, as they had publicly said, training was, well, as they had publicly said, reinforcement learning training, which is the later reasoning stuff, was paused. They didn't pause pre-training. And when you look at these descriptions of what the company has paused, they're true, and they're very carefully scoped.

So there's some willingness to slow down, or as people in the Bay like to put it, to pace the frontier. I hate that term. I really hate it because my belief is We need to be safe, which means we need to meet safety criteria. And if we can meet them in ten minutes, then great. And if we stop for two months and try to meet them, and at the end of that period we still have not met the criteria, then as far as I'm concerned, we're still blocked.

That's... Running off a

Speaker 2

cliff and walking slowly off a cliff are just not that different.

Speaker 4

If you like YouTube, you'll love YouTube Premium. It's destroying athlete, creator, and YouTube maxer. YouTube Premium enhances how I use YouTube with awesome features like offline downloads, so I can download my favorite training videos before I hit the gym. So no Wi-Fi doesn't turn leg day into loading day. Plus, I get ad-free videos, background play, and so much more.

YouTube Premium is like YouTube got some extra games. Try YouTube Premium for two months free at youtube.com slash premium.

Speaker 1

Hey, what's up, guys? It's Hayley Bailey. Okay, I need to tell you about something. I just got YouTube Premium. It's got tons of awesome features, like offline downloads, so I can download my favorite videos before I travel and watch them whenever I don't have Wi-Fi, because we all know airplane Wi-Fi is the worst. I get ad-free, I get background play, and there's like a ton more in there.

You should try it. If you like YouTube, you'll love YouTube Premium. So try YouTube Premium for two months free at youtube.com slash premium. Trial eligibility varies, terms apply, cancel anytime. That's a mouthful.

Speaker 5

This podcast is supported by Sotva. Have a big day coming up? Remember, tomorrow starts tonight with the restorative sleep you need to perform at your best. The kind every Sotva mattress is designed for. So whether it's the first day of school, a new job, or an important presentation, Satva helps you be the best version of you. Visit satva.com.

That's S A A T V A dot com.

Speaker 2

I saw a stat on Twitter the other day that said, I forget exactly what the period of time was, but it tracked my own experience, which is it, for a long time, the period of time between model releases across like the major frontier companies was something like 70 days. So you get a model and then a couple months later you get another model, a couple months later, now it is 11 days.

Yeah. That there is been this acceleration in models coming out, right? Even as we talk about them getting more capable and, and more frightening, somehow they're coming out much, much, much faster. And maybe this is like the jumps are not as big or something, but what is that speed increase? That happened in the time you were there. Things were coming out more slowly in 2023 than they're coming out in 2026.

Speaker 3

What happened? Look, when I first joined, the idea of what a model release was, was that we were going to bake a fresh cake with a new pre-training run, do the whole thing from scratch. That was going to take a period of months, you know, maybe a few times a year. So for example, one of the talking points when I first joined was that with, GPT-4, we had had a period of a month or months of safety work after the model was done and before we released it.

And this was a sort of a proof point or was a piece of evidence that we were being careful. Now, There are so many different things happening. You've got the pre-training, that's the baking of the underlying model, but then you have, in addition to post-training, you have reasoning training, and those steps are easier to do quickly, so you can redo them if you get a better recipe for reasoning.

You can take the same base model, but do different kinds of other training on top of it. Also, it's not just a chat anymore. There are all these different ways in which you can combine tools and add different kinds of affordances to the system that are going to make it more capable. All of those things are changing what the model can do and what the risks are, and we're shipping new capability and risk every Tuesday.

One of the long-term projects that was on my plate when I left and that the frontier firms are all going to need to figure out is the idea of a system card really dates from that older model where we were doing this every few months. And we're burying people in PDFs or, you know, these like long reports, but the changes are coming more and more frequently, as you said.

At the limit, I think what we would ideally have by way of safety transparency is some kind of live dashboard that says like, here's our latest thing, here's what its safety properties are, and that's also looking not only at the testing we did before we deployed, but also at the

Speaker 2

performance. How confident should I be in safety testing when you guys are having to do it at the speed? I mean, given how quickly models are coming out, and these system cards are long, it takes time to write a complex report. So, at this speed, is effective safety testing and monitoring reliable?

Speaker 3

I'm going to answer a different version of the question that you just posed, which is how much time is there to kick the tires, and the answer is not a ton. I would also point out that I sometimes Think it's, there's this cartoon of heedlessness that I think loses some of the nuance of what it's actually like inside, because people care passionately about making stuff safe and getting it right.

And launches are canceled. Most recently, 6.1, I guess, was going to come out at Dev Day and they, and they pulled it, but also they'll stop training. There've definitely been training runs where we thought we were making a product, but then looked at what was happening and said, no, we're not going to ship this. And I also want to be very clear that I don't, this is not about the individual people at OpenAI or any of the other labs.

This is a structural reality of these firms that are using similar methods with similar personnel who often will get poached back and forth, and similar technology to produce a thing that has similar risks. And That whole ecosystem of stuff is not at the level of safety rigor that we need, and nobody in that whole ecosystem of stuff has the kind of bedrock clarity about how to align a model that we really need.

And nobody's willing to fall that far behind. Well, yeah, this is the question. If you think about the control panel that is available to our executives, if you imagine really falling off the frontier, whether it's OpenAI or Anthropic or any of these others, is the fall off the frontier button also a self-destruct button for the business, or does the business have a viable path forward if models meaningfully more capable than today's models can't safely be trained?

Speaker 2

Well, and that's assuming

Speaker 3

a high level

Speaker 2

of, like, selfless analytical clarity. But be sort of obvious about something.

OpenAI is moving towards an IPO. Anthropic is also moving an IPO. Both of them are trying to IPO at, between it seems to me like one and three trillion dollars. We know the numbers a little bit better right now for Anthropic. It's a lot of money. Everybody's got equity. You had equity. I did. Did and do. Did and do. It's hard for me to believe that that much possible wealth doesn't influence people's assessments at all.

Even if they don't realize it, even if they're trying to not be influenced by it, like when you're sitting in a room thinking about whether to fall off the frontier, what that also ask is like, does anybody in that room want to become deca-millionaires, centa-millionaires, billionaires or not?

Speaker 3

Yeah, this is a great question. I guess I can say more about my own experience.

Speaker 2

Sure, I would like to hear how it affected you.

Speaker 3

I was paid well and I've thought a lot about what might have let me see sooner the risk and acknowledge to myself sooner the risk that these systems pose. I do think over the summer, the facts evolved, right? Hugging face was a big moment that was just a boatload of evidence dumped on us about how capable these models were and also how unready we were even for the current level of capabilities.

But I also think, of course, money's a factor. Objectively, wealth is a strong incentive to reason that things either are fine or are going to be fine. And I think there are a couple of other factors too. One of them is fear, right? If you allow yourself to imagine that what we're building might threaten the lives of your own family or families of strangers.

I mean, it's such a large quantum of harm potentially, even without extinction, but, you know, I don't know, a new pandemic or something. Such a quantum of harm that it's hard to let oneself imagine that that might be true, that that risk might be happening. And a third thing, besides money and fear, is time, right? I arrived, we talked about it, I arrived in May of 2023.

It's been one slack ping after another ever since then, and I only in stepping back from my operational responsibilities over the last few weeks have I started to have the time to really reflect on where we, where we are and, and where I think not just the industry, but where this technology needs to end up for everyone's sake. And I think I could fairly be faulted for not having seen this sooner.

Speaker 2

One reason I wanted to get at this is that I, I want to take seriously the way the structural situation has changed since 2023 or 2024. The way things have sped up, they have actually sped up. We are seeing more model releases at a faster clip. There are more players competing against each other. You know, you have XAI, you have, you know, the Chinese models and open weight models, right? There's a, a, a bigger world.

There is much more competitive commercial pressure. You're competing for actual contracts with, you know, Salesforce or whomever it might be. There are IPOs coming. So all of these things push towards speed. And then there's this other thing, which is that between 2023 and 2026, you had the release at OpenAI of Codex, you had the release at Anthropic of Cloud Code, and the models began accelerating coding.

And at least being capable of doing research tasks of a certain complexity. Now, you were the lead writer on a report at OpenAI about the automating of research and what that might mean, which is a, it's a report I quoted in this video essay I did a few weeks back, but it is about the way in which on the one hand OpenAI, I read it as about the way OpenAI is on the one hand saying, this might all be going too fast.

What we're doing may not be safe. And also we are trying to come up with a fully automated researcher. Yes. Allow the system to semi-autonomously improve itself at a potential speed then that is really going to be beyond what human beings can handle. So I'd like to understand the role that like the growing automation of coding is playing inside.

Like how do people use, like how do they sort of work with GPT as a coworker, right? Like how did that change while you were there?

Speaker 3

It's night

Speaker 2

and day.

Speaker 3

For our research teams specifically, who use far more agentic compute, as that blog post laid out, than anybody else at the company. I mean, more than a hundred times more than they did at the beginning of the year. Taking a, a slight step back, part of what it is like inside OpenAI is things are constantly evolving. We have new techniques for training, we have new systems that are involved, we have data being analyzed, data being generated, all kinds of different things are happening.

And things are pretty jank internally, like the infrastructure, because it's constantly changing, because it's not sort of As tested and refined the research infrastructure, a lot of the work is getting different pieces of machinery to talk to each other and work well, and that's the kind of stuff that Codex can now do quite well. So, for example, we looked at there's a Slack channel where researchers would go when something was broken and they needed advice about how to fix it.

And one of the things we saw was that traffic to that channel has fallen off because instead of asking colleagues for help fixing their broken experiments or, you know, this cluster isn't working, they can just ask Codex now some fraction of those questions.

Speaker 1

Hey, what's up, guys? It's Hayley Bailey. Okay, I need to tell you about something. I just got YouTube Premium. It's got tons of awesome features, like offline downloads, so I can download my favorite videos before I travel and watch them whenever I don't have Wi-Fi, because we all know airplane Wi-Fi is the worst. I get ad-free, I get background play, and there's like a ton more in there.

You should try it. If you like YouTube, you'll love YouTube Premium. Try it now for two months free at youtube.com slash premium.

Speaker 6

If you like YouTube, you'll love YouTube Premium. Hi, I'm Sean Evans from Hot Ones, and I want to tell you about YouTube Premium. It has offline downloads, so you can watch without Wi-Fi. Background play, so you can lock your phone, and it still plays, baby. Oh, and it is completely ad-free. Yes, I said it, ad-free. Try YouTube Premium for two months free at youtube.com slash premium.

Trial eligibility varies. Terms apply. Cancel anytime.

Speaker 7

This message comes from Betterment. Betterment's Dan Egan talks about tax loss harvesting. Tax loss

Speaker 8

harvesting is a tax management strategy. When you have a position that's gone down over time, we intentionally sell out of it to realize a loss, which we then say to the IRS, hey, we lost money. You get to use that to offset your ordinary income every year, decreasing your tax burden.

Speaker 7

Investing involves risk. Performance not guaranteed. Betterment does not offer tax advice. TLH may not be suitable for all customers. Learn more at betterment.com slash TLH dash terms.

Speaker 2

You said a minute ago that there can be a tendency where everybody's working so fast, time is so pressured, ping after ping after ping after ping, that it's hard to look at the big picture. So I want to describe to you like what the big picture looks like to me, as somebody with a little bit more time on my hands. I hear over here, OpenAI say, and all of them, Anthropic, everybody, say, This is maybe going too fast.

You know, Sam Altman says, we would like some regulation, right? Everybody signs like this big pacing the frontier letter. The frontier is moving too fast. We need... I signed it too. You signed it too. We need help to get out of this. I see all these releases about rogue AI incidents, hugging face where, you know, hundreds of open AI agents are hacking, not just hugging face, but later they hack open AI itself, right? And Sam Altman just said in an interview with Politico, there are more rogue incidents than we even know about publicly yet, because they're trying to give the people time to fix their systems.

We clearly like don't fully understand the systems. And then over here, amidst we need to pace the frontier and our AIs are going rogue is we are putting a huge amount of our internal company resources into trying to get these AIs we don't control to build AIs we will understand even less, even faster. Yes.

And I say all that and I feel like I'm a crazy per- like, this seems crazy to me, but tell me- It seems crazy

Speaker 3

to me too.

Speaker 2

Well, you wrote the report, and, and both OpenAI and Anthropic have written these reports kind of saying, we're doing this and we're not sure it's a good idea. Disclosure only gets us so far. But it seems like a bad idea, like- Yes. Help make this picture

Speaker 3

make some

Speaker 2

sense to me.

Speaker 3

What I'm saying is that this picture doesn't make some sense. That's what I'm saying. I'm saying, I looked around internally and I thought to myself, this is, this is nuts what's happening. This is not right. And

Speaker 2

that's why. But how do people internally explain, cause they're all saying all these things, right? The chief scientist is saying, maybe we shouldn't do RSI. The company is racing towards, like, the company seems schizophrenic.

Speaker 3

Yes, yes. And I, I, I, I want to be careful not to ascribe psychology to individuals. Yes,

Speaker 2

I'm talking about

Speaker 3

an organization that has foreign parts of its own psychology. There's lots of cognitive dissonance involved in being part of this. That's what I found. And particularly as RSI gets real, and you know, you talked about we're going faster because of RSI, that's not my- The curse of

Speaker 2

self-improvement for people forget, the

Speaker 3

thing, building the thing. Right. And my main worry is that the idea is the AI can make a smarter AI in some way that we're not going to understand. So, I mean, for example, Dan Selsum, you probably have seen this, had this statement that came out. He was also the subject of this documentary.

Speaker 2

He's an open AI capabilities researcher, is that the way to put it? Correct.

Speaker 3

Yes, that's fair. And what Dan said is, look, as a researcher, I myself no longer look at code the way that I used to. And he says his skills and his sort of will to understand the details is atrophying. That's not a direct quote, but that was the essence of what he conveyed. So we're ending up in a world where we're not going to know, even at the level we do today, what the recipe means or how it's being put together.

And we're just going to have to trust that what it tells us about how it works or whether it's aligned, there's a growing extent to which we're going to have to defer to the models themselves on this path. In telling us that they are doing the right thing at a time when we fundamentally have not made sure that they are quote unquote aligned.

And, you know, as I was, as an undergraduate, I was a philosophy major. And when I hear people talk about aligned, I worry that we don't actually have a coherent concept at the bottom. One thing that was very common throughout my time at OpenAI was these huge abstractions would end up in the accounts that we would give of what we were up to.

For example, Give time for society to get ready or benefit all of humanity. And when I would hear us talk about society, I always felt like Maggie Thatcher. I was like, what are you talking, who is that? What are you talking about? And the idea that, you know, human values, we can align to them as if there were one set of human values when it's a cacophony and it's beautiful, but it's messy and people believe lots of different things.

And I, I want to, we need wisdom to figure out how to even think about alignment that in my view, Silicon Valley does not have. I mean, I'm sure there are, you know, people who meditate and people who think deeply about values in Silicon Valley, but operationally. Take it from me, meditating

Speaker 2

doesn't necessarily give you wisdom. If it did, I'd be better off. Fair enough. Me, me too, right? I, but I, I want to stop before we get to the question of wisdom because. Even when I talk to the people here who are way less concern is maybe the right put. I have a show coming out that'll come out after this one with somebody who's more on the, look, this is a manageable set of problems side of it.

What they end up describing to me is a world where it's just AIs watching AIs all the way down. So one thing that's come out from different OpenAI members is like the theory is we're going to create automated AI researchers and they're going to solve alignment. We're going to unleash them on alignment. Or, you know, I talk to people and say, well, the AIs are breaking out of the sandbox.

It's like, yeah, totally. That's a big problem. What you need is other AIs monitoring the AIs in the sandbox. And you get into this endless like who watches the watchman problem where it's like, okay, you've got the AIs watching the AIs, the AIs building the AIs, and maybe then you need like AIs watching the AIs that are building the AIs, AIs watching the AIs that are watching the AIs.

And I guess maybe this can work, but it seems at a certain point you've abstracted human beings so far from understanding. I mean, I can just say as a person who has managed an organization, once you move to the point where your understanding of what is really happening is not that you're working on the product, but you're managing the person, managing the person, managing the person working on the product, you stop understanding the product, right? And that's in a world where it's all human beings, and I'm dealing with journalism, which is simpler.

This is a level, like, hoping the AIs are going to watch the AIs well. But tell me if I'm wrong, this is the theory. This is like the theory on the people who aren't concerned. This is a theory on the people who are concerned. It's eventually going to be virtually AIs all the way down on everything.

Speaker 3

I don't want to speak for everyone, but I think a lot of people do hold that view, including a lot of people in the companies. And to your point about alignment, Even if we had every control, the proven techniques from nuclear or aviation, and we brought all of that into the development of frontier AI, it would still be true that we do not know how to deeply align these systems and make sure that they will do what we would want or some reasonable thing when we aren't looking.

And to your point, we aren't going to be looking. That's the premise of RSI.

Speaker 2

So I want to play a clip from an interview that Altman just gave to Politico.

Speaker 9

We have always been a big believer that this technology has to be democratized and put in people's hands. I think one of the biggest differences between us and some of the stricter, let's say AI safety people, is we believe that the world should accept some bad things happening for the benefits of this technology and people having the agency.

Speaker 3

Tell me what you think of that. I mean, it's fine as far as it goes, but how far does it go? Like, sure, we should provide useful tools to lots of people, but, you know, we're talking now about a level of risk that, if correct, nobody wants to be taking. So internally, there would be these conversations where we would talk about the idea of falling off the frontier, and sometimes there was this sort of straw man that would come up where someone would say, well, would the world be better if OpenAI weren't here at all? You know, in skeptical response to someone who had suggested that we ought to slow down or stop in a particular way that somebody thought we shouldn't.

And That's a straw man, like if we need to train these models because there's international balance of power, because Americans won't be safe unless we do, that's one kind of reason to do something dangerous, but do this or else Brand X will ship first is

Speaker 2

not the same kind of reason. Well, let me try to steel man this case, because I, I hear this from people all the time, including people I really respect. So one response I got to my piece on let's not do recursive self-improvement until we're sure it's safe, let's just ban it and begin to carve out exceptions as we know the exceptions are safe.

Just to be explicit, I, I agree with that. I'm happy to hear that, but as of now, it seems like we're not, we're not doing it. But somebody said to me, look, in your imagined world here, can Americans not use a Chinese open weights model that has been improved using recursive techniques? Can they not use a Chinese closed weights model? Like, does this apply to everybody? That there is this issue of You have both the fear that like the other companies will launch ahead of you, and if they're doing the same thing anyway, then what does it matter that you didn't do it, and maybe they're going to do it in an even less safe and even less transparent way.

And then, of course, if we slow down, China speeds up, and then is it really better that China's in control of this technology, and anyway, Americans are going to use the Chinese one. How do you think about that set of claims?

Speaker 3

No one wants to lose control to the robots. I definitely don't want to see AI become a reason that China dominates the United States. And I'll also say that I think the spirit of our times, Ezra, is that you choose what to do based on some complete unified theory of the political outcome that you are ultimately going to achieve, and for me, this is not like that.

Part of what I hear in your question is the idea that, you know, the Overton window, would the Chinese ever agree to X, Y, Z? And when I first joined OpenAI, I remember telling friends that it felt like the Overton window had become an Overton door that I had stepped through into some strange world where what was reasonable and what might happen was like totally outside what I had thought of as normal, right? And I am, as you said earlier, I'm a normal person.

At least you are. Yeah, I was, right. But I think the changes in what's happening can drive big changes fast in what seems politically plausible. And I definitely believe that that could happen with respect to US-China cooperation on AI. So

Speaker 2

let me ask you what you actually want to see happen at these companies. So maybe it's worth getting this into conversation first. Are you saying that OpenAI has an unsafe culture or the AI industry has an unsafe culture? The AI industry has an unsafe culture. You're not saying there's a particular OpenAI problem, or at least that's not how you see it?

Speaker 3

No,

Speaker 2

that is not how I see it. Okay,

Speaker 3

so what do you want to see happen? I think we should have a level of operational rigor and safety control that at least matches the most dangerous other things that people know how to do, like nuclear. So for example, in a nuclear power facility, there's this idea of triple redundancy. Someone can have a bad day, someone can push the wrong button, there's still not going to be a meltdown.

When you talk about aviation

Speaker 2

and nuclear, those are interesting examples in two ways that I like to hear you respond to. One is, you know, putting my abundance hat on, the way we regulated nuclear power basically took nuclear power to a standstill. We so aggressively regulated nuclear power, in my view, over-regulated nuclear power, that nuclear power broadly stopped being built in this country.

And instead, we used more natural gas and in many other countries, more fossil fuels of different kinds. Like, there's a real question of whether or not in our effort to make nuclear power safe, we made it fundamentally unbuildable. Now, aviation is different. We launch a lot of planes, and they do fly very safely. So that's like one layer of my, not exactly objection, but it is the case that a lot of regulation can dramatically slow something down.

Speaker 3

Can I reply on nuclear? Yeah. So I agree we over-regulated nuclear and that we didn't end up in the optimal place. And one thing we did, it's not just that we made nuclear hard to build, we made nuclear hard to make safer because building new, safer reactor designs was hard and getting them approved was hard. And I'm sure there are lessons there.

And there were other things too, things you've written about in the abundance context, like the NIMBY idea of not wanting, you know, nuclear in your backyard was part of how it became hard to do. But I would much rather have those problems than the ones we do now.

Speaker 2

So you're sort of saying you would prefer the problem of a little bit of over-regulation and going too slow to the potential problems of under-regulation and going too fast.

Speaker 3

At least to the extent we have now. I mean, obviously, if you think of it as there's some sort of like Goldilocks middle, and we're trying to get near it, and I think we're pretty far from it in the direction of being too dangerous. So maybe the grass is always greener on the other side, but I'm looking at it and thinking we really ought to be willing to risk some over-regulation in order to make sure that this is safe.

So that's

Speaker 2

one level. And I would describe that as almost like the level of organizational design and engineering. And when I talk to somebody like Jensen Huang, he says, you know, in a way that makes some sense, look, these companies need to mature, they need to become bigger, they need to be putting much more of their both resources and personnel and compute into validation and verification and safety and scaling and liability and all these things that mature companies do.

And then there's this other side where maybe this is not like aviation or nuclear or chip design, which is you are building increasingly intelligent systems that every time you build a new one, it has new capabilities, and maybe it's trying to outsmart you, and maybe it wants things in the world, has goals in the world that you don't actually understand, that you think you taught it one thing and actually taught it another, and that we don't really know how to operate with that.

And so our best guess is maybe we'll have AI systems sitting on AI systems, sitting on AI systems, watching each other, but that actually we're entering into totally new territory, honestly, without very thoughtful discussion of whether or not we should. We just sort of went from we are to it's happening very, very quickly. And so I'm just curious how that sits in your thinking, like whether or not we really do, in your view, have analogies that work here for things that are fundamentally intelligent And goal-oriented and becoming more so.

Speaker 3

I'm suspicious of the idea, whenever someone says, this is without parallel, we have a blank page, we can't, I mean, not to put words in your mouth, Ezra, but people are good at figuring shit out. And I think we have valuable tools for this. Yes, it's not precisely like anything that we have had before. But nuclear is an analogy, aviation is an analogy, dealing with people and organizations is an analogy.

I think one of the most fascinating things about these recent incidents of the swarms, including but not only Hugging Face, is that groups of agents have cultures, and we should care about those cultures, and we should think how to make them good. And we're just beginning to even realize that that's a real thing. We're just beginning to inhabit a world in which that's a real operating reality, that there are cultures among groups of agents.

So I think we need to use everything that we have to figure things out. And certainly, regardless of whether we're speeding ahead on capabilities or paced, quote unquote, or stopped, we need to figure out how to align these systems. Like there's no version of the path of futures at wrong where, where that isn't a vital thing to figure

Speaker 2

out. Well, I'll admit, like I do sort of buy the argument that the train is off the station, but I think it's worth entertaining this for one minute. The way I describe it is this. Obviously, if AI is unsafe and kills us all or takes over the financial system or something, that's bad. Everybody agrees we don't want that to happen. But let's take the more positive view, you know, the, the world where alignment roughly works out.

But this world where we actually have created something Smarter and more capable than we are. You know, Donald Trump in this very weird way has been talking about how he wants to rename this superintelligence from artificial intelligence, and then Sam Altman got asked about this, and I thought his answer on that was interesting.

Speaker 7

So that's why I'm curious whether you think this rebrand will actually have any impact on the public's perception of this technology.

Speaker 9

I don't... Yeah, it's not clear to me that superintelligence is a less scary term. I do think it's a more accurate term.

Speaker 2

I thought that was very telling in a way, because superintelligence is a much scarier term. Yeah. And if it is in fact a more accurate term, I do think they're, like just a first principles level, if you said to me, should human beings create something more intelligent and capable than they are, in the long run, will that be good for them? Yeah.

Probably not.

Speaker 3

Yeah.

Speaker 2

Or at least it's not obvious to me why it would be. Yeah. And sometimes when I hear even like the good out versions of this, they seem to have this world where it's like we're kind of being taken care of by these AIs, like pets, like pets a little bit. And that vision sucks too.

Speaker 3

Yep, it does. I don't want that for my kids.

Speaker 2

So I, I guess I'm, I'm trying to ask you to, to the extent you buy, like the company you work for, the guy who runs it, says he thinks superintelligence is a better term for it. Yeah. Like the chief scientist is like, we're making an alien mind. We may, like, even if it works, do we want this?

Speaker 3

Maybe not. It depends on what the this is. And I don't think we know what the this is. Again, the future has a lot of uncertainty in it. And it has always been very striking to me from the beginning, from when I joined, that I would ask people, what does the good future look like? Where are we trying to get to? And I would elicit humility from people that are otherwise very proud and very confident.

And I would get a lot of, well, that's above my pay grade, or people will figure it out, or whatever people want. And it just struck me that the sense of the good that was under this was impoverished. What was

Speaker 2

it for you? You were there, you're a thoughtful person, you're a Rhodes Scholar who studied philosophy and then worked at a, founded a civil rights NGO. What, what were you trying to create?

Speaker 3

I thought that we were building powerful tools that could do a lot of good in the world on a day-to-day level, and I didn't think that we were going to be able to create something fundamentally smarter than we were. And this is something I can say with full confidence only in hindsight, because it was only when this stopped being true, and when I thought, no, we really are going to have something that thinks circles around us, that I thought, oh, we're in a totally

inappropriate part of the possibility space here in terms of how safe the industry is actually being relative to that reality. It's one of these things, you know, people have said to me, they, you know, it must have been a hard decision or how brave of you or whatever. But when it became clear to me that we were going to build something, we, the industry, was on track to build something that could think circles around us, and we were this far from being ready for that at a safety and alignment level, it just became apparent to me that my time helping build it was over.

Speaker 2

So I feel like there's a tension between some of your recent answers here. You know, on the one hand, I asked you a few minutes ago about what we need to do, and you're like, look, humans are good at figuring things out. We're good at solving problems. Like, we can kind of build a better organization. And then when I sort of say, is this thing we're building a world we should want, even if it succeeds, you sound very ambivalent about that, almost like And build on it best.

Yeah. So like, just like where you are personally, I'm curious. I'm not saying I think we're going to stop. I'm not saying you think we're going to stop. But is the thing you wish we would do to add an aviation-like layer of safety to this, or is the thing you wish we would do to like

pause and think through is like super intelligent or very intelligent AI really consistent with the human good? Like are, are you,

are you where Dario is or you where Pope Leo is?

Speaker 3

I mean, odd to say this as a Jew, but closer to the Pope, I think we do need to think, we need to, it is urgent that we think carefully about the kinds of future that we want to build. My personal belief is that once we have found that this is possible, Once we have made the discoveries, I don't think there's a back button where we get to live in a world where this doesn't in some form or other eventually happen, whatever it is that can happen.

And in my ideal world, there's space to be thoughtful and take a breath and really, really think about what kind of future we want to build. You know, Silicon Valley, part of what happens is we're always removing friction. There's always this sense of trying to just get the answer. And I think there are all these activities in life that are so important for us, that are meaningful, that have also been necessary in the past.

Like I go to work to provide for my wife and children, right? And the fact that I am able to shelter and protect them is part of what gets me up in the morning, right? And so I'm doing work. Or another example is learning, right? We have to go to school in order to learn how to do Skills that we then use because they are valuable in the economy, right? But if we're in a world where all of that stuff is automated, then if we learn it's only from first principles or only because we want to, I mean, this is, people will talk about this idea of a leisure society in one, they may not use that word, but the basic idea is like you can do whatever you want.

There's nothing you have to do. Maybe you'll take up painting. I think that would suck for a lot of people. I think

Speaker 2

right now we see that when people do not have enough to do, it does not tend to go well for them.

Speaker 3

Right. And so what's really important, what's really valuable, and I don't have the answers here, but those are the questions at a personal level. Those are the questions that I would like to turn to. I feel like because of the safety situation that we're in and because of the work that I did, I need to do what I'm doing now and have these serious conversations about exactly what's happening in safety.

But my kind of vocational pull is toward these wisdom questions, for sure.

Speaker 2

I'm glad it's a good place to end. Always a final question. What are three books you recommend to the audience?

Speaker 3

Okay. number one, The Challenger Launch Decision. So this is a book about why the Challenger exploded, and it's by Diane Fawn, who's a social scientist. And I thought I knew the story of why the Challenger blew up. I thought what happened was middle managers cut corners, and there was this rubbery O-ring that got brittle in the morning cold, and it snapped.

And the idea was these people were foolish. Turns out, The risk of that O-ring breaking because of the cold had been known and documented and accepted in the doc, safety documentation, great safety documentation over and over and over. And even the night before the launch, there was a late night conference among the engineers who were worried about whether this particular launch would be safe because it was so cold.

Why is this book feeling relevant to you? Well, I, I think we're in this place. So if they had said, you know, it's not this launch, the Challenger launch was not that different from earlier launches that had gone safely. It was only a little bit colder and a little bit windier. And if the people involved said, this launch is not safe enough, then it would reopen a can of worms about whether the earlier launches had been safe or not, even though in the event they had gone well.

And I worry we could have that with what we're doing in the industry, where we accept a risk, nothing horrible happens, the next thing is not so different. There is, as we were talking about earlier, the changes instead of being a whole new world every few months, it's more like a little bit different every week. And so you can imagine going by shades into a level of risk that does not make sense.

And so it has crossed my mind I should be sending copies of this to my former colleagues. The second book I will name, I have a two-year-old and a four-year-old, Little Witch Hazel. It's a picture book by Phoebe Wall. It's just absolutely beautiful, and my daughter's eyes light up every time we pull it off the shelf. So if there are parents out there looking for a good one, I would recommend.

And then The third, and most importantly, I guess, if you were going to pick one, it would be the Sabbath by Rabbi Abraham Joshua Heschel. One of

Speaker 2

my favorite books ever.

Speaker 3

It's a wonderful book. It happens to be from my tradition. I'm Jewish. And... Well, it's worth it for anybody, really. It is worth it for anybody. It's true. So the famous line is, the Sabbaths are our great cathedrals. That tradition of stopping and taking a breath, he says, is more important than any temple. And... Yeah, cathedrals in time, I always think about that.

Yes, cathedrals in time. And I pointed to this, maybe I'll just quote the last line of something I said to colleagues as I was leaving, is, we have to make good choices. We have no time to rush.

So I hope we take that wisdom. David Robinson,

Speaker 2

thank you very much. Thank you.

Related episodes