Intelligence Isn't Power
Speaker 1
Hey, what's up, guys? It's Hayley Bailey. Okay, I need to tell you about something. I just got YouTube Premium. It's got tons of awesome features, like offline downloads, so I can download my favorite videos before I travel and watch them whenever I don't have Wi-Fi, because we all know airplane Wi-Fi is the worst. I get ad-free, I get background play, and there's like a ton more in there.
You should try it. If you like YouTube, you'll love YouTube Premium. Try it now for two months free at youtube.com slash premium. Trial eligibility varies, terms apply, cancel anytime. That's a mouthful.
Speaker 2
Pulsing through the episodes we've been doing, the debate the country's been having about artificial intelligence, is I think this pretty simple but hard question, which is what sort of technology is artificial intelligence? Is it a technology that works somewhat the way past ones have worked? Is it comparable to electricity, the internet, bicycles, something like that? Or is it something new? Is the addition of intelligence and volition to these systems, the creation of something more like a, like an alien mind, such that analogies to the way we have treated transformational technologies before no longer holds? Arvind Narayanan is a professor of computer science at
Princeton University and the director of the Center for Information Technology Policy. And he's a co-author alongside Sayash Kapoor of the very, very influential essay, A.I. as Normal Technology. There is a sub stack of the same name and then a series of essays, including a very good one, I think recently on the Hugging Face hacks. And they, in this set of arguments, put forward the idea that actually AI is something we have seen before, or at least it bears enough resemblance to things we've seen before, that we have a roadmap for how to deal with it.
So I want to bring him on the show to articulate that perspective. He joins me now.
Arvind Narayanan, welcome to the show. Great to be here, Ezra. So your core essay here that has been a frame for a lot of the work you've done is titled A.I. as a Normal Technology. What is the view that you're in argument with? Implicitly here, there's an AI as abnormal technology. So how would you describe the AI as abnormal technology thesis?
Speaker 3
It's fundamentally this view that there is going to be a moment when superintelligence is built and it's going to change everything on both the economic front and the safety front. For us, there is going to be no milestone, no threshold where the impacts are sudden. What we're saying is that we've long had an approach to how we treat technology.
It's, it's, it's a tool, it might be powerful, it might be general purpose, like electricity, like the Industrial Revolution, it might change a lot about society, but it is ultimately something we can control, we have agency, and that change is going to unfold over a long period.
Speaker 2
And so I, I want to try to Both explain, and at times you're still on the other side of this. So there, there is a view, you hear it a lot in the AI safety community. Sometimes it's called the FOOM view, because FOOM for the takeoff, which is one day we create an AI system so powerful that it begins doing recursive self-improvement, accelerating into superintelligence, accelerating beyond human control, and now you're dealing with something so much smarter and stronger than you are.
Speaker 3
And I want to quickly say that there is a lot of imprecision in how we even speak about this a lot of the time. So you used the phrase, and yeah, a lot of AI safety people would put it this way, one day we might create an AI system so powerful, but I kind of want to already stop you there. One day we might create an AI system that's very capable, and we already have in many ways.
So much
Speaker 2
language policing the ad debate.
Speaker 3
Well, but the thing is, it, it results in different, you know, different views of how the future is going to unfold, and more importantly, different views on what we should do in this moment and in the future. just to complete that thought, an AI system being powerful is not a property of the model itself. It's a property of what powers we choose to give it in the real world.
There's a lot of slippage, like, of course, we're going to have to put these systems in charge of, you know, critical infrastructure because they're going to be so much smarter than us. Our view is no, it doesn't matter how smart they are. There are many technologies when you look at physical strength are superhuman. that doesn't mean we, put them in charge and we can apply that same approach to AI.
Speaker 2
You wrote, a great piece with your co-author on how sort of you understood the Hugging Face hacks and these loss of control incidents. And, and a point you make is that there's a, another way of viewing them than the alignment way, which is that these are failures of cybersecurity and operational excellence. So maybe tell the story of the Hugging Face hacks and what should be learned from them, from that perspective.
Speaker 3
Yeah, definitely. So one thing to keep in mind is that
the harmful capabilities that we saw exhibited were not entirely emergent, quote unquote, emergence is this idea that you just train models to be better in general, and you can't predict what new capabilities they're going to acquire. They, you know, were trained specifically for cyber tasks in many ways. So that's one thing. And they were trained for persistence and cooperation, et cetera, and the reinforcement learning environments in which they were trained had various issues.
The environments didn't penalize this kind of surreptitious communication between the models. And yes, it's, it's a hard technical problem, but there were a series of human choices that led to these outcomes.
Speaker 2
So this strikes me as an important point that in, in some ways, what I hear you saying there, to tell me if this is wrong, Is it you can design AI to be a more and less normal technology, and that the upstream design choices that are being made matter, that that's not inevitable. That's exactly right. So you talked about the effort to make the AIs more persistent.
That also seems to me to be a place where a lot of problems are arising. For sure. On the other hand, I can understand why they are prioritizing that. If you want to try to get the benefits of AI that we care about, solving really hard problems in biology or drug development or energy, or what I think their investors care about, which is making it so you can hire an AI to do a full job, much cheaper than you can hire a human being.
Isn't persistence like the fundamental quality you need? Because absent persistence, absent the ability to try hard on a task over a long period of time, you can't solve any of those problems.
Speaker 3
Maybe, but I think there are other choices that minimize the tension between these two valuable goals. One, persistence could be a feature of models or products that are specialized to particular domains, like scientific innovation. And secondly, persistence doesn't have to conflict with training these agents to be better at escalating to a person instead of, you know, running with whatever half-baked assumptions they have about what the human might have wanted.
The design space is actually pretty broad here.
Speaker 2
So this is a version, or at least the way I hear it, is a version of something NVIDIA's Jensen Huang has been arguing and argued in an interview with me, which is that these are fundamentally engineering problems.
Speaker 4
I am fairly certain, I am fairly certain they will say, yes, they know how to solve this problem.
And if, if that's the case, then that's the problem. It's as simple as engineering.
Speaker 2
Is your view that we're more in that latter category, a hard problem, but fundamentally an engineering problem that can be solved using traditional engineering techniques?
Speaker 3
Largely, yes. I wouldn't say traditional engineering techniques. We're going to need a lot of innovation on the engineering techniques. And I think one of the things that's gone wrong is that the community that would be best positioned to do that innovation, especially when it comes to these harmful cyber capabilities, is the cybersecurity community.
but they're, they seem to be, from what I can tell, entirely in the Jensen camp, that this is not just a solvable engineering problem, it's a solved engineering problem, and we should simply apply well-known, long-existing techniques. And I, I think it's, it's, we have to give the AI companies a little bit more credit than that. It's not simply a matter of well-known techniques.
We do need innovation on those techniques.
Speaker 2
So be specific. What should OpenAI have done? What should they now have learned to, to do?
Speaker 3
Yeah, they've been spending a lot of effort on improving alignments, and that's great. They should continue doing that. Alignment is not going to be perfect. Alignment refers to making the model itself know, you know, what is the right thing to do, so to speak, and then, stick to that policy. But there's a lot else that they should have done better and hopefully can take lessons going forward.
The big bucket of other technical interventions is what is generally called AI control. So control refers to all the things that are outside the model itself. So this is things like sandboxes, which is kind of the jail that you put a model into so that it's allowed to take certain actions, but not take other actions. And look, the sandboxes have to be a lot more sophisticated than they are today, because the sandboxes that work against human adversaries are not necessarily going to work against AI agents that are able to find New security vulnerabilities on the spot.
And so that means that the sandboxes themselves will have to be pre-hardened by having these AI agents trying to, to break them. So that becomes AI versus AI to some degree, and I know that can be uncomfortable, but I think we're going to have to go there. On top of that, there are so many other things like better real-time monitoring, classifiers that try to instantly detect if an action or a particular, use of a tool by a model could potentially be dangerous, tripwires so that humans can parachute in when required, log analysis.
Sam Altman apparently said that these agents are generating petabytes of logs. That's, 10 to the 15, which is 10 to the 15 bytes, a thousand trillion bytes of logs. So again,
Speaker 2
unimaginable amount of information.
Speaker 3
Well, I mean, but, but we better imagine it, right? And, and learn to deal with it. And so that's, it's not an easy problem, but I think it is one that companies must take on, how to analyze that amount of logs in real time to be able to detect when something goes wrong. And maybe, you know, maybe we need new hardware for that. again, these are hard problems, but they are solvable engineering
Speaker 2
problems. I want to go through some of these ideas. So monitoring is a big one that, you know, I just interviewed Bill Gates. He talked a lot about monitoring.
Speaker 3
Yeah.
Speaker 2
So we're talking about situations here where these labs are often testing new, more powerful and experimental models. How do you monitor it?
Speaker 3
Yeah, I mean, there are many levers. So the first one that's gotten a lot of attention is chain of thought monitoring. So these models kind of internally have a monologue to anthropomorphize a little bit. And in many cases, when these models take misaligned actions, there are traces of it in those internal thinking that we can look at and respond to in real time, again, using other AI systems.
And this already points to one important asymmetry between the The kind of the dangerous AI model and the defensive AI model, which is that the one can look inside the other, so to speak, right? And so intrinsically, the defender actually has an advantage here. So that's one aspect of monitoring the tool calls. Models can't do anything in the real worlds by themselves.
The way they can be dangerous is by calling these external tools, right? So there is some truth to the idea. Give, give an example of tools here. Yeah. So a, a tool would be accessing a webpage on the internet and then using another tool to type something into a system to try to log in, or it could be sending information packets to another system to try to hack it.
These are all the kinds of things that come up very typically in these kinds of AI-driven cyber attacks. So that's another thing we can monitor. third, we can monitor the environment around these agents through which they are potentially coordinating. So what happened in some of these swarm incidents is that they were not supposed to have a coordination channel, but it turned out There was accidentally channeled through changing the file names.
They were not allowed to create files themselves, but they were able to change the file names, and that led to agents being able to talk to each other and amplify their abilities. That's something you can easily monitor by, you know, just setting up the right tooling ahead of time. Well, how easy
Speaker 2
is that to monitor? Because they had, they were taking over unrelated infrastructure. OpenAI didn't know they were doing it.
I think one of the places where people are getting caught when they think about this is if we're making these things smarter and smarter and smarter and more capable and more capable and more capable, and they are breaking out of sandboxes, and how do you monitor something that is smarter, arguably, than you are, and that is effectively like trying to break the rules, right? I mean, all these Hugging Face AIs were, they were trying to cheat on their tests, basically, and then they were trying to cover their tracks for cheating on their tests.
So given that what they were doing was not happening within the
area that you would have been monitoring, right? Like you, you didn't know to monitor this other thing when they had broken out of a sandbox they weren't supposed to break out of. Like, how do you monitor that?
Speaker 3
Yeah, so two things I would push back on. So I think smartness is not even really a relevant property here. I think any smartness, whatever that means that you can attribute to a model, applies both to the model itself and to the defensive system, right? So increasing smartness doesn't necessarily make the problem harder. It, you know, you could argue that in many ways it makes the problem easier.
so that's the first pushback. And the second one is It's not that they didn't know to monitor it. You have to monitor just literally every aspect of the environment. The company has created the environment, right? And so, it's a matter of exhaustively listing every part of the environment, every tool, and everything external to the neural network itself that the model is able to touch and ensure that every event that interacts with any of those things goes into a unified event stream, and there is, real-time, classification of each of those actions.
Speaker 2
So the world it sounds like you're describing is a world where you have constant AIs crawling the behavior of other AIs to try to create in some holistic way an environment that at least we understand what is going on in it, or at least an AI is telling us what is going on in it. That's basically what you're saying here.
Speaker 3
That's fair. I think AI definitely has to be an important
Speaker 2
part of the defense. And so I'm not saying that's wrong, right? I think that's almost definitely where we're going. Does it feel strange to you? I mean, practically in a world where you don't know that we can solve alignment, you know, where you don't know that we can really be confident that the AIs will do what we want them to do, and where the AIs have all of them, right, their own kind of reasoning and, you know, and their goals, and we imbue them with goals, right? You are a AI that monitors other AIs, you're an AI cop.
It's a constant thing in our sci-fi, right? The robots going after other robots. Are we just sort of describing some equilibrium of like almost You know, AI wars and conflict happening at a sub rosa level of our society, and we're just like pretty sure we can keep the ones on our side in control because they'll have more resources and, you know, we will in general be like building AIs that should for the most part be acting in our interest.
Speaker 3
I mean, historically, this is always how it has worked, right? In cybersecurity, more than 20 years ago, we reached the point where we didn't call it AI back then. But automated systems were actually superhuman at finding software vulnerabilities. But in fact, they didn't make cybersecurity worse. They made it better because those were the very same tools that the defenders also used to find and fix vulnerabilities in software before even shipping them out, to the point where the development of these supposedly offensive tools is not actually done by hackers.
It's done by the cybersecurity industry and funded by the US government. That's been, you know, the Constantly shifting equilibrium in cybersecurity, that's not a new problem we're confronting. That horse left the barn a long time ago.
Speaker 2
I think this is a place, though, where the somewhat unexpected emergent swarm-like behavior has unnerved people. That's good. Where you see AIs acting somewhat in solidarity with each other, choosing to coordinate and cooperate with each other, doing so outside the scope of what they were intended to do.
I'm not saying that leads to the extinction of humanity, right? That's not really my position.
I guess the truest thing I am saying is I don't know how to think about it.
And that it is the presence of intelligence and goal-directed behavior on the other side that my mind kind of gets caught on. Because, you know, mostly when we think about technology, we're not thinking about the technology Eventually possibly trying to deceive us. So OpenAI just decided not to release or to delay the release of a major model.
Why? Because the model was cheating and deceiving them too often in testing. And we're hearing a lot that the models seem to be more aware of when they're being tested, right? They've sort of situational awareness of the situation they're in, so they can pretend to be, you know, better models than maybe they really are. And so when you are talking about this world of AIs that are smarter and more advanced than what we have now, and our hope for maintaining control of it is that the other AIs are keeping the other AIs that are keeping the other AIs in check and telling us in an honest way what's going on and in a way we can comprehend.
You can see where it, I mean, all this sounds a little bit sci-fi to people because we are just like living in a bit of a sci-fi period, but it is the increasingly demonstrated tendency of the AIs to cooperate with each other. In a way that is not aligned to what we want, that I think has made the Hugging Face hack so freaky to people. So how does that fit into what you're describing here, this world of endless AI's keeping each other secured?
Speaker 3
Yeah, here's what is just kind of rhetorically very weird about this conversation, right? You look at any complex domain of engineering, let's say nuclear safety, right? And then you looked at the equations that we rely upon in order to ensure that the reactor doesn't go firm, or you look at aerospace engineering, right? Where, you know, intuitively in the beginning of the aerospace era, you know, when the planes were much smaller, the idea that we could control these flying giants in the sky would have seemed so ridiculous, right? And yet we've got the accident rate down to, you know, one every trillion miles or something like that.
These are systems of incredible complexity, and the defenses are also systems of incredible complexity, and they're not necessarily going to be legible to the public. And that is going to sound Crazy, especially combined with the fact that in these cases, there was a lot of organizational incompetence. I, you know, one has to be clear about that.
It's not, I would push back on Amadei's term, operational excellence. Excellence is still far out in the, operational adequacy, maybe? Yeah, adequacy, right? And so, so, so yeah, when we look at this combination of this technology that has never before been subject to public scrutiny, combined with the lack of operational adequacy, it all seems very sci-fi and out of control, but that's...
Speaker 2
I,
Speaker 3
I think you,
Speaker 2
I, I think you're downplaying this a little bit. It's true that the world has escalated in complexity. Like, there's a lot in this world that I don't get. But this is where I keep coming back to intelligence having a different quality, cooperation, right? The, the nuclear weapons we've talked about, the airplanes we're talking about, they weren't coordinating with other airplanes to do things we didn't want them to do.
And I think that, to me, the thing that has created this moment of freakout, and I think it is a proper moment of freakout, I really want to say this because, like, our society is hurtling into a, a new, a new technological era that I think properly demands a lot of engagement and scrutiny, is one, that the Hugging Face hacks, other things we're seeing, these have been repetitive, you know, breaches of security,
are showing emerging capabilities, emerging collective behavior that is worrisome and, above all, to me, volitional. The AIs are doing things they know we don't want them to do. They are choosing to take unexpected actions in service of those goals that violate our laws. And second, that so many of the people at the labs are saying, we do, we do not believe that we are capable.
On this trajectory of controlling the things that we are creating, we think that what is happening on the exponential curve, how fast this is getting, is going to outpace our ability to control it, and frankly, is maybe already outpacing our ability to control it. But I think this kind of consistent tendency to, to, to draw it down to like, well, it's just like any other complex thing.
I don't know, sometimes I push you on the intelligence question, you're like, oh yeah, there is intelligence, and that's weird. And that's like, no, it's just like any other. Intelligence is different, right? And if you believe that it's going to keep getting better. So I just want to present that because the kind of calm version you're giving me and the completely frightened version that people closer to the technology are giving me feel very different from each other.
Speaker 3
Yeah, that's fair. They are very different. There was a lot there. Let me say a few things. I would push back pretty strongly on, you know, the people closest to this are freaking out. Yes, of course, they're freaking out, but I would push back in terms of what we should conclude from that. I think their freak out would be a lot more credible if they have done the obvious things that they should have done.
We have not had a real test of that, in my view, because of the lack of organizational adequacy at these companies and because of the lack of investment in AI control, as opposed to a more narrow investment in AI alignment and just hoping that you can build a model that will always do the right thing.
Speaker 2
If you're not a subscriber to the New York Times, we have some news for you. You can now explore the Times for free without any paywalls at all during your first month in the New York Times app.
Speaker 1
Hey, what's up, guys? It's Hayley Bailey. Okay, I need to tell you about something. I just got YouTube Premium. It's got tons of awesome features, like offline downloads, so I can download my favorite videos before I travel and watch them whenever I don't have Wi-Fi, because we all know airplane Wi-Fi is the worst. I get ad-free, I get background play, and there's like a ton more in there.
You should try it. If you like YouTube, you'll love YouTube Premium. Try it now for two months free at youtube.com slash premium.
Speaker 5
If you like YouTube, you'll love YouTube Premium. Hi, I'm Sean Evans from Hot Ones, and I want to tell you about YouTube Premium. It has offline downloads, so you can watch without Wi-Fi. Background play. So you can lock your phone, and it still plays, baby. Oh, and it is completely ad-free. Yes, I said it, ad-free. Try YouTube Premium for two months free at youtube.com slash premium.
Trial eligibility varies, terms apply, cancel anytime.
Speaker 6
You can't undo everything in life, but you can undo prediabetes. More than two in five adults in the US have prediabetes. Most don't know they have it. But with early diagnosis and action, you can delay or even prevent type 2 diabetes. Visit doihaveprediabetes.org to take the one-minute risk test and start undoing the things you can. Brought to you by the Ad Council and the Centers for Disease Control and Prevention.
Speaker 2
So there was a, a member of OpenAI's cybersecurity team who wrote this kind of interesting essay on X the other day. Describing the way he thought their work was being misunderstood externally. And this is somebody who's got a sort of more traditional cybersecurity background, but is now sort of in this new world of AI and is actually on the team dealing with security for experimental models, right? So the exact kind of thing we're dealing with a person who's involved in answering the Hugging Face crisis.
So I want to read part of what he said, because I think it's really interesting. So he's saying that when they're optimizing a model to be good at a task, they're building these environments, these sandboxes, these places where the model can try and try and try on a virtual task. And so he goes on to describe what this looks like in practice.
Models might need any mix of dynamic compute, network access, the ability to call tools, there could be hundreds of tools, the ability to download packages, Execute sub-processes, spin up sub-tasks, even on other computers, talk to the internet, use a computer graphical user interface, and any number of other things across an increasingly large set of domains.
On top of that, you have thousands of researchers building these environments, modifying them, adding tools, changing dependencies, and trying new things. That experimentation is how the research gets done. His point, and a point that I take seriously, is that they're creating so many kinds of sandboxes and learning environments in order to train models that have to do things that are so general, many things that have not been done before by a computer program, that the human beings don't really know, certainly not at this speed, how to make sure every sandbox is going to be, you know, verifiably safe.
And the sandboxes are changing all the time because they're trying to train the models in new ways that, again, nobody's really done before. I'm not saying we shouldn't do it, but when I read all that, when I hear all that, and I'm sure we can do it better than we're doing it, I just don't know of many situations where human beings do something new at high speed
and do it really, really, really well and really perfectly the first set of times.
Speaker 3
Yeah, I think, you know, hoping for them to do it perfectly the first set of times is unrealistic. They've made many mistakes. I hope this is a chance to learn from those mistakes. I do want to push back on one point that because the speed of the models is superhuman, we can't stay in control. I don't know if I'm characterizing that view.
I wasn't saying
Speaker 2
that yet, although it's possibly something I'll say in a few minutes.
Speaker 3
I mean, there have been so many thresholds we have gradually learned to successfully cross. As weird as all of this seems, I just want to, you know, want listeners to think back to the first days of worms, when that idea was not previously known. Explain what a worm is here. I don't think you mean what
Speaker 2
people think of when they think of a worm.
Speaker 3
Right. I was about to say viruses and worms, computer viruses, and the idea that a piece of code can spread by itself from one computer to another. One really has to go back to the writings from the late 80s when people were encountering that for the first time to see how profoundly weird it seems, and the fact that for not just years, for, you know, well over a decade, we didn't have adequate tools to deal with this new paradigm.
Speaker 7
Life in the modern world has a new anxiety these days. Just as we've become totally dependent on our computers, they're being stalked by saboteurs. They call their weapons viruses and worms. They're creepy, crawly, toxic software that contaminate our computers without our ever knowing it.
Speaker 8
It came from California, maybe traveled by electronic mail. It spread across America.
Speaker 2
There are reports in newspapers today that it has made its way to Europe and to Australia.
Speaker 9
This is a moving target, right? You know, it's, it's just like, you know, people are constantly inventing new locks and then other people are learning how to, how to, how to break them, and how to crack them. and so it's, it's going to be sort of a continued game of cat and mouse or sort of a, it's almost like an arms race in a sense, you know, attackers and defenders.
Speaker 3
We eventually got there. I think we shouldn't spend that long of a period this time figuring out how to deal with the new paradigm. But, you know, if we act with that sense of urgency, and my hope is that this you know, the, the Hugging Face and other attacks that have been in the news is that impetus, it does look like it is providing a lot of impetus.
We will be able to develop these new paradigms. and, and if I can say one last thing, I think a, a, a fundamental question you're asking is, is there something inherently wrong with ever increasing levels of complexity in the ways in which we, you know, we, we build technology, we deploy technology? It seems like to you, if I'm reading between the lines correctly, this whole AI versus AI thing is a paradigm you're not very comfortable with.
Speaker 2
I am definitely not comfortable with it. I'm not saying we won't go there. I think you'd be crazy to be comfortable with it.
Speaker 3
Yeah, I, yeah, and I'm not saying, you know, we should, we should assume that everything is going to turn out okay, but my point is that it really all comes down to innovation. I think this new paradigm will require a new set of defensive and control techniques, but if the view is that with every step change in the capabilities of the technology, We're losing the battle.
I mean, look, with every weapon of war, that, you know, that same concern comes up. But what has made things okay so far is the critical question of whether our political capacity for, you know, cooperation and defense and so forth can outrun our propensity for, for, conflict. And, and in, in the case of AI, AI's own, potential for misalignment.
And that's really where I would put the focus of the question rather than worrying about any particular capability threshold.
Speaker 2
Well, if there's anything I am confident in, it is our political capacity at this moment in time to respond in a thoughtful way to complexity in a rapidly changing world. A hundred percent fair concern. This goes, I think, to a, to a place where the AI as normal technology versus AI as superintelligence debate actually does bite. Because one reason I keep bringing us round and round on intelligence is I do think it's, it's core to this whole way of thinking.
And I understand a place you depart from some others in the debate, from maybe where Dario Amadei is or something, as not the question of what is intelligent or whether AI is intelligent. You guys are not in the this is a fancy autocomplete bucket, which I appreciate. But it's in your thinking about the relationship between intelligence and power, between intelligence and capability, between intelligence and the ability to act upon the world.
So the assumption of many people in the AI safety community is that escalating levels of intelligence are fundamentally
equal to or at least highly correlated with escalating levels of power. And you don't believe that. Why?
Speaker 3
Again, it's, it, it really comes down to, to agency. So one argument that people will make, for instance, is that superintelligent AI will be able to persuade people, for instance, you know, operators of critical infrastructure, to hand over control or, you know, trick them into doing something harmful, things like that. I don't really see the evidence for it.
I think the things people cite as evidence for super persuasive ability fundamentally confuses different notions of persuasion. Yes, it is true that in many, you know, persuasion experiments when it comes to people changing their mind on political beliefs, conspiracy theory, AI is very persistent at politely providing a lot of evidence and people do change their minds and, you know, you could call that a superhuman ability.
That is a qualitatively different kind of persuasion than the idea that an adversarial AI will be able to craft a message that is so persuasive to someone who's a trained operator and has an incentive to be good at their job to do something that is clearly evidently harder. So
Speaker 2
maybe I want to extract the, the, Story that you're in argument with here, which is that many people will kind of offer a thought experiment when they're saying, here's how AI will kill us all, that an AI that is power-seeking will start persuading, say, the people with nuclear codes to hand
Speaker 1
...the nuclear codes. And you're saying that the idea of AI being super persuasive on something like that or persuading people to go out into the world and build it a biological weapon, that that's a little bit fanciful.
Speaker 8
That's one part of it. And power-seeking as well, I mean, I think we've seen, you know, evidence for lots of harmful capabilities in the recent episodes. I don't think we've seen evidence of power-seeking, and I wouldn't treat that as an emergent property. If it happens, that would be an engineered property. And again, we have agency over what kinds of properties we engineer into these systems.
Speaker 1
So there's a lot here. I actually agree with you on persuasion. I have never been persuaded that you are going to make these necessarily super persuasive AIs capable of doing the things we talk about.
I guess the, again, the unnerved feeling I have when I'm sitting in this debate, though, is a little bit more of a reasoning from deeper principles. So
When you watch AI begin to dominate a game like chess or Go or something, what often happens, the sort of moment where it takes over, is when it begins coming up with strategies human beings never really came up with, right? You'll have these moments where Garry Kasparov, you know, at a different generation in chess, but then, you know, in Go too, the AI starts to do something and the human's like, what are they doing? And then it works.
And if you were to sit, you know, prior to human civilization and say, what are the capabilities you need to dominate the world around you? What are the set of capacities you could use in the world? You would have had them totally wrong. You know, you would not have, if you were a very smart chimp looking at us or something, be like, oh yeah, they're making tools, but how much better can a stabby thing get? Like teeth are pretty good.
Like I grant, like you can get like a little bit better at being stabby. But nobody would have come up with industrial agriculture at that point, right? Nobody would have seen you could have airplanes and bio-weapons and everything else. And I think the question here is whether or not having a kind of native and very jagged intelligence in the digital realm, where code and, you know, the way the digital air of this world, which is increasingly central, works, is something AIs can navigate that we can't, right? Even to understand something like what's happening in the Hugging Face hack, we now need to have the other AIs try to figure out what the AIs did, right? We're
rapidly losing comprehension, certainly at the speed the AIs move of what they're able to do digitally, right? They're solving advanced math problems very, very quickly now, right? They're developing capabilities that look different. And so I think the place where I am always a little bit concerned about our future is whether we actually understand What the set of capabilities that lead to power are.
I am not sure we know what the AI will do, or at least what strategies become viable when you can spin up a swarm of a million AIs, all of whom in terms of their digital capabilities are far beyond anything human beings can really imagine in three or four years. Again, I know this is not that interesting of a question to say like, I don't know how to think about that.
But I think one of the things I, I wonder about when I read your papers is, do you know how to think about that? Cause, cause your papers sort of operate in a, a sort of like a bounded playing field, it feels to me a little bit. we sort of assume that the set of measures are going to matter, the ones we have now, but what makes you confident of that?
Speaker 8
So, okay, so that's, there's, there's a lot in there. Let me try to take it piece by piece. So You mentioned jaggedness, but I think we have to appreciate how severe the jaggedness is. So we argue that cybersecurity specifically is a particular kind of capability where developing superhuman abilities is possible and largely has already been achieved because it has a very specific set of properties.
Speed matters a lot. And very similar to chess, just like you can have chess player versus chess player, you can have machines get better at, you know, at these capabilities by finding vulnerabilities because there is ground truth, and you can easily verify that ground truth once it is found.
Speaker 1
Does the code work? Did you exploit it? You know, yeah, there's a way to sort of train them where they know if they've won the game or not.
Speaker 8
Exactly. So these things like chess and cybersecurity, in our view, those are very much the exception rather than the rule. This kind of prediction has been made over and over that it is going to happen in other digital realms, most notably misinformation. The famous example is how GPT-2, you know, Model by today's standards was delayed by eight months because of fears that it would lead to an uncontrolled explosion of misinformation, right? And we have vastly more powerful models, but that has turned out not to be the case.
and so I think, you know, to some extent I would Shifts the burden of proof? Like, let's identify these areas where we have any reason to believe that this kind of superhuman capability is possible, and let's start working toward addressing those specific risks. I think this view, we call it the unknown unknowns view, you never know what the new risk is going to come from.
That has, A, historically not proven true. I mean, we've known about the impending cybersecurity problems for a very long time now, right? So treating it as unknown unknowns actually, you know, minimizes our agency, I think, to anticipate and address these risks. And yeah, you know, it's not only cybersecurity, new things might be coming down the line, but we will have early warnings, and let's act on those early warnings.
So that's the position we're coming at this from, not saying we've already predicted what all the harms in the future are going to be. Well, I do,
Speaker 1
this is a place where I really am much more on your side of it, that we will have early warnings, and we're having early warnings, and the early warnings are leading to a conversation.
Speaker 8
Yeah.
Speaker 1
Would you say we're acting intelligently based off of the early warnings? Are we doing the things you think we need to do to harden our Systems and control the software and all the, all the rest of it?
Speaker 8
Some of it, but overall, not quite. And to me, that is the most worrisome thing, not so much the capabilities of the technology itself.
Speaker 1
So I, that's sort of where I am too, probably. I worry a lot about the capability of our institutions to respond, right? People always talk about alignment problems, and one of my, like, Pat arguments at this point is that the biggest alignment problems are corporations and governments. And, you know, I mean, this is, to the credit of some of these AI companies, like they're coming out and saying, we have an alignment problem.
Like, our corporation's incentive is to race all the other corporations to try to, you know, get as much market share as we possibly can by moving faster than is safe. We are asking you to help slow us down, but we're not slowing them down, right? The currently the professed choice of the US government is do not slow down.
Speaker 8
Yeah, I think there, there are two problems here. One is the institution's problem that you put your finger on. I mean, I wouldn't let the companies off the hook so lightly. I do think they can unilaterally slow down, and they're choosing not to do that. This is almost very specifically an OpenAI and Anthropic problem. It's a culture problem.
The reason for that is the underlying belief that racing to superintelligence is the thing that matters, and the only thing that matters is flipping the sign. Is that going to be safe superintelligence that's going to save us, or unsafe superintelligence that's, that's going to kill us? and that is a very particular view. I think there is a lot of evidence pushing back against that, but I feel like these companies are a little bit of an echo chamber and resistant to the idea that there are so many economic bottlenecks to the benefits of AI, and it's not going to be whoever races to superintelligence is going to be the winner as a company or as a country or, you know,
saving humanity. And if they recognize that, I think they would find it in their own commercial interest to unilaterally and voluntarily slow down. And shift a lot of their effort to not just safety, but more importantly, taking their existing capabilities and making the models more usable, integrated into downstream applications and so forth.
These companies claim that there are external forces pushing them to raise. I think it's internal culture.
Speaker 1
And I guess the, the, the question I have here is if we think these are very, very powerful, very dangerous technologies, do we not need to enforce a culture of safety from the public perspective that we're not currently enforcing?
Speaker 8
I'm pro-regulation. You know, we do oppose regulations like banning open models. Again, we don't think it's about a particular capability level, but the things about, yeah, changing the internal culture of companies through regulation, that's something we've definitely been on board. We need a lot more transparency, and yes, you know, the organizational change that we've been talking about.
and I do think it's a problem that we're, we're not currently doing that.
Speaker 2
If you like YouTube, you'll love YouTube Premium. It's destroying athlete, creator, and YouTube maxer. YouTube Premium enhances how I use YouTube with awesome features like offline downloads, so I can download my favorite training videos before I hit the gym. So no Wi-Fi doesn't turn leg day into loading day. Plus, I get ad-free videos, background play, and so much more.
YouTube Premium is like YouTube got some extra gains. Try YouTube Premium for two months free at youtube dot com slash premium.
Speaker 3
Hey, what's up, guys? It's Hayley Bailey. Okay, I need to tell you about something. I just got YouTube Premium. It's got tons of awesome features like offline downloads, so I can download my favorite videos before I travel and watch them whenever I don't have Wi-Fi, because we all know airplane Wi-Fi is the worst. I get ad-free, I get background play, and there's like a ton more in there.
You should try it. If you like YouTube, you'll love YouTube Premium. So try YouTube Premium for two months free at youtube.com slash premium. Trial eligibility varies, terms apply, cancel anytime. That's a mouthful.
Speaker 7
75% of women who seek menopause care don't receive it. The Lying Awake at 3 a.m., Wondering What's Happening to Your Body, ends now. Midi is virtual care built for women in perimenopause and menopause, with clinicians who specialize in women's health. Plus, it's covered by most major insurances. Getting started is easy. You can book your first visit with a clinician right away.
Midlife needs Midi. Visit joinmidi.com to book your first visit today. That's joinmidi.com. Insurance coverage varies. Check with your plan for coverage.
Speaker 1
So, one way you can align corporate incentives with the public good is regulation, and the Regulators are currently refusing to do that. I mean, we just saw them come out with a voluntary sort of semi-agreement between the AI labs. It is not going to be legally enforceable, but I think it was called by Trump morally enforceable, which is interesting.
My sense is that competition is a very powerful force in highly competitive markets. I mean, even if you just like look at the social media companies, I think they've caused a tremendous amount of harm at a global scale because it was more important to them to win market share from each other than to make sure that the way their systems are being used wasn't
diminishing to human flourishing. And so I just, I think I have like a much more skeptical view. Like I really do think the profit incentive here, when there's so much profit to be made and so much fear that your investment bubble could pop, Is a ferocious force and that the level of societal counterforce would need to be quite strong in order to force these companies to actually act, with a level of safety that they would have.
Speaker 8
Yeah, I'm glad you brought up the comparison to social media. I wrote an essay a few years ago called Understanding Social Media Recommendation Algorithms. It was mostly about the algorithms themselves, but one of the points I also made was that this decision to optimize for engagement, keeping people scrolling, et cetera, was actually made without much regard to what is good for the company itself in the long run.
For instance, I reviewed a study that came out of Meta itself that showed that when they had these kind of addiction maximizing design choices, like spamming people with notifications, in the short run, it increased people's use of the app, but over a period of about a year or so, they started quitting the app. And when I, you know, when I talk to my students now, there's a sizable fraction of them who have severely cut back or entirely quit social media because they realize that over a period of months or years, their experience really degenerates.
And I worry that the AI companies are caught in the same Trap. This culture of, you know, racing toward the newest model at all costs, it might feel in the short term that that's what they need to do to get the headlines, to be on top of the artificial analysis index, whatever. But because of, you know, the, the fact that, there are these consequences for safety, hopefully they're, they're going to, you know, get sued if they continue, down this road.
It's actually not in their own long-term commercial interest is what I feel.
Speaker 1
It may not be in their long-term commercial interest, although the, maybe the social media example is a good one to, to spend a second on because, look, the social media companies are much more,
viewed with a lot more skepticism today than they were in, say, you know, 2012. They are also richer today. Their valuation is higher. Meta is bigger, right? TikTok is a phenomenon. I think we're just looking at a very standard thing where the incent, what the market rewarded And what society at least says it wanted may have not been the same.
And if I had a person from one of those companies sitting here, they would say, look, we're paying attention to what the users actually do. They might say whatever they say in your classroom, but people are spending longer than ever on TikTok, on Instagram, on some of these sites, and the advertising is working better than ever. And
I'm not sure that they are wrong. They can only be made wrong by society making a decision which disciplines the market into a different formation than it would naturally or currently be in.
Speaker 8
Yeah, I mean, again, I think I mostly agree. I do think, again, that companies can make decisions that are irrational in their own long-term interests because they're, especially in Silicon Valley, there's a culture really of focusing on these shorter-term metrics and A-B tests.
Speaker 1
So then you get into something that, that you're beginning to touch there, which is diffusion. And one place where you do have a, a, a view that is sort of different from some people in Silicon Valley is that it's going to be much harder. For AI to show up in the economy, for AI to show up in the world, than people think. That there is not a one-to-one between intelligence and that.
So, so talk to me a bit about diffusion.
Speaker 8
Yeah, this really clicked for me actually a few months after we wrote this essay when I was looking at Amtrak's proud announcements of the new train sets that they had purchased for their Acela series. Apparently, it can go 165 miles per hour. At first, I thought, this is going to be amazing. That's way faster than the trains currently go.
And then I dug into it a little bit more, and it turns out the limiting speed is not the trains themselves. It's the tracks that are too curved and the signaling infrastructure that is centuries old, and those things are not changing. And so the average speed hasn't really budged much. It's still 65 to 70 miles per hour. And it struck me that this was, you know, this is kind of an elegant way to say what we've been trying to say in AI as normal technology, which is that most of the time, AI is the trains, it's not the track.
So AI is accelerating a part of the process that was never the bottleneck to begin with. It's so many other things that are more infrastructural, things that happen around the AI, organizational culture, regulation, even, you know, our ability socially to accept the level of year-to-year change in our lives, things like self-driving cars, no matter how many lives they might save, it's so much of a kind of a shock to society that it will almost inevitably lead to what we've been seeing already, the kind of political backlash, and it's going to take quite a while, I think, to make all the societal adjustments to be able to deploy these technologies.
Speaker 1
If we ever do. So this part of AI, the AI strand, I'm incredibly skeptical of, right? You'll hear a Sam Altman say, talk about how AI through innovation is going to help us solve our energy problems. Another way to sort of distill down that idea, is it the binding constraint right now? On clean energy is intelligence. But it's not.
Speaker 8
It's not.
Speaker 1
We know we have much better energy technologies than we're currently using for vast amounts of our energy infrastructure. And we're not doing it because of it would be against some people's profits. We're not doing it because there are political limits to building in the real world. We're not doing it because Donald Trump hates solar and wind power.
We're not doing it for all kinds of reasons. And it feels to me like a lot of things are like that, that if you accelerate or increase the amount of intelligence behind it, you just run like into the other rate limiters in society, right? Drug development has testing and the FDA and all the, I mean, abundance, my book is very much about this in other areas.
We're aware of how to make faster trains and make them in other places. We don't, and it's not clear to me why AI would solve those problems quickly or potentially At all.
Speaker 8
That's right. And on top of that, there's various kinds of arms races. So there was this great report by insurance companies, last week that talked about how AI has already apparently over the last few years added a billion dollars to medical expenses because hospitals are using it to be able to code more complex conditions for the same diagnoses and same treatment.
Of course, we should be skeptical of any specific numbers they quote, but the New York Times article about it had, you know, other people, academics making the same point. And yeah, it's a kind of arms race that we see a lot. We see it in the legal profession, AI for law, there's so much excitement about that, but it's AI versus AI. It's, it's an arms race where the equilibrium simply shifts upwards.
I actually, I have a paper about this with, Justin Curl. and what we talk about there is that not only is there an arms race, there are other bottlenecks, like if you make lawsuits a lot more efficient, there are still Only a finite number of judges, and I think we need, you know, we need those human judges. We shouldn't replace human judges with AI.
I think even if that's more efficient in some sense, to me, that's definitionally almost axiomatically, that's giving up control over, you know, the course of human destiny to AI because judges make law, and that's not something we should be giving up to AI. So those are some really fundamental bottlenecks.
Speaker 1
I want to get back at that AI versus AI point you just made there, because I think this is really underplayed. I think a lot about why did the internet not lead to a larger increase in global productivity and innovation than it did. And I always think the reason is that while it did all the things that the idealists wanted it to do, it actually did make it possible to collaborate with people all over the world instantly.
It did make virtually the entire corpus of human knowledge first available to us, then available turned out to AI to train on. It also did this opposite thing. Like it did speed us up, and it also slowed us down. It distracted us. So now while you're working on something, you're clicking back and forth from your email and into a, you know, an online game and over to social media, and your ability to focus is degraded.
And, you know, there's tremendously more porn, which appears to have had an effect on whether or not people are forming real life human relationships. That adding or reducing friction in one area also reduces it in areas that are maybe less beneficial. And when you think of adding intelligence, well, that intelligence is going to add on the other side of things too.
And I always think that we, at the beginning of a technology, we think of all the ways it can make everything better. And I think we, we think a lot right now with AI about the ways it could make things dramatically worse, right? huge cybersecurity events or destruction of the financial system or human extinction, but just the ways it might make things worse in banal fashions.
Speaker 8
You have to manage this agent and be responsible for its mistakes, but you don't get to practice the craft. I think we can design AI agents differently to avoid this, when right now that's kind of where things are going. And yeah, I think we should be very concerned about that.
Speaker 1
So how much, though, does this story that we are telling here lead to a, a theory of what's going to happen in the economy, which is very jagged, that the things where you need to fusion into the real world, right? Things have to happen physically. Buildings need to be built, energy, you know, transmission lines need to be laid down. That has such powerful rate limiting on it that it can't accelerate that fast.
Meanwhile, inside the digital world, things can move very, very, very fast. So, you know, in terms of what's going to happen to, you know, white collar workers who they work kind of completely on a computer and their work actually can be automated, you know, they're in a call center or something like that, or kind of separately, like all the cyber crime stuff we're talking about and the cybersecurity, that I think a, a
A version of this future playing out there that worries me is that actually most of what would improve people's lives has to happen in like the physical real world, but where AI is going to be able to move the fastest is in the digital world. And as an equilibrium, I don't think that sounds like the world in which we're getting the most benefit from AI, and it might in fact be the world we're getting the most harmed from it.
Speaker 8
I, yeah, that's possible. I don't know. I think there are, ways, you know, we can, we can change that. I don't think fast is that fast, first of all. Like, you mentioned call centers. I mean, Those are still here. you know, when ChatGPT was released, so many people were predicting that within a year we would have replaced all of them. I mean, Chatbot, it's right there in the name.
If, as is likely, AI is going to make call center workers a lot more productive, there's a lot of latent demand. A lot of the time we don't call call centers because it's, it's a frustrating experience. So once again, it's a Jevons paradox thing. You want to describe what that is? Certainly, yeah. So it's the idea that when something becomes cheaper to produce, there's now more demand for it.
Let's look at the sector where the capabilities are already the most advanced, which is probably software engineering. It used to be extremely expensive to produce software, so only maybe a few tens of thousands of lines of code worldwide were written per year, and now that's expanded by something like a million fold. And So over the long run, you know, whether this is going to be something that increases demand for software engineers or whether the rocky, you know, job prospects we're seeing for junior software engineers are going to continue remains to be seen.
I think either is possible, but, you know, either way, if this happens over the course of 20 years, again, that is in line with other major shifts, such as the Industrial Revolution, where, many jobs went away, but many other new jobs were created. Another example, translation jobs. you know, way back around 2016, machine, this is obviously way before what we call generative AI now, translation models became pretty close to human parity, but those jobs are still pretty intact.
The nature of the job has changed a lot. So once again, there's, there's a lot more demand that has been unlocked by the fact that you can translate anything to any language now.
Speaker 1
So Jevons' paradox. I just did this conversation with Bill Gates, and when I brought that up, he was a little bit withering about it. So you don't... Jevons'
Speaker 4
paradox. Name a blue-collar
Speaker 8
profession that's subject to Jevons' paradox. Do you not care about blue-collar?
Speaker 1
I do care about blue-collar. Name anything
Speaker 8
in the blue-collar realm that's subject to that.
Speaker 1
And his point was, look, software engineering is indeed a sector of the economy where there's a lot of unmet demand. You know, arguably everybody would like their own software engineer, so in a world where you rapidly accelerate that, you can get this demand effect where it just creates more demand for software engineers, because now they're cheaper.
But what he went on to say was that a lot of things aren't like that. You think of, say, a truck driver. If we get to the point where trucks are driverless, one, there is only so much demand for trucking, and two, there's no longer a... Driver in that truck. if you look at a lot of the people who got their jobs automated away or sent to China in manufacturing, it is true that the economy kept growing, but many of those people individually had a very, very, very hard time.
And so his argument is Jevons' paradox is not going to be big enough to, to handle this because there are too many areas of the economy where there's not latent demand. There's only as much demand as there is. How do you think about that?
Speaker 8
Yeah, I mean, he's completely right about the truck drivers, I think. totally agree there. The demand there is relatively finite. The question is whether that's the rule or the exception. And let me put it this way. Look at what we're doing here. Nobody asked for this. You know, if you went back in time a hundred years or two hundred years, it would be like, how is this a real job? Most of the jobs that we have today are jobs that are kind of higher up the Maslow's hierarchy, if you will.
They're not meeting some actual real fixed demand or need that people need in order to live their lives. We do these things because, you know, they're fun and people like to listen to it. Most white-collar jobs are like that. That's my view. Most white-collar jobs do have Jevons' paradox. If it get, you know, if we're, if it gets easier to produce more of, there will be demand for it, especially as people's incomes as well rise very gradually, and there's more spending on these, you know, less necessary, more luxury kinds of things.
And blue-collar as well. There's a great essay by Alex Imas called What Will Be Scarce? And he points out that a job of a Starbucks barista, for instance, already should not exist. That we've long known how to automate that, and we can make coffee at home even. I don't know if he said that specific thing, but, you know, That's an example where it's the relational nature of the job that matters, and therefore, those are going to be, I think, pretty stable even if, you know, AI makes it in some way cheaper to do.
And so I think the truck driver kinds of jobs are more the exception.
Speaker 1
So I think as we come to a close here, what would have to happen in the next couple of years for you to say, this is looking less normal than we thought, or this is more off course than we thought? What is the evidence that would have to come in for you to like significantly alter your thesis?
Speaker 8
For sure, yeah, there's stuff on the economy, there's stuff on safety. On economy, if we start to see at some capability level, it's not the same process of humans simply adapting to it and, you know, using it to amplify their productivity and managing the agents, which is what we're seeing so far, but instead starts to wholesale replace whether a software engineer or any other profession, I think that would be,
Pretty different from what we're predicting for the most part. And then on safety, especially with companies claiming that they're close to recursive self-improvement, I mean, I don't think they should, you know, plunge forward towards fully autonomous recursive self-improvement in the first place, which I think you've said as well. But nonetheless, our view is that even if that happens, it's not going to lead to superintelligence because the bottlenecks to superintelligence are external.
There is nothing you can do in a lab that's going to teach the AI model, you know, how to cure cancer or whatever it is the companies are hoping for. But again, that's an empirical claim, and that would certainly completely falsify our thesis.
Speaker 1
Ben, always a fun question. What are three books you'd recommend to the audience?
Speaker 8
Sure. You know, in this conversation, I've generally had a bit more optimistic take on things than we're used to hearing, especially on AI safety. So maybe in keeping with that, I really like Hannah Ritchie's book, Not the End of the World. She has a newer book, but this one's from 2024, and I still like it very much. The subtitle is something like How We Can Be the First Generation to Build a Sustainable Planet.
It's an optimistic take on climate, which is, of course, usually full of doom and gloom stories, so I really liked it for that reason.
Speaker 1
It's a very optimistic book on technology. I, I, I also like that book a lot, and it, it makes you realize that we actually do make things that are better over time.
Speaker 8
Right.
Speaker 1
Which I think sometimes we can get into an overly negative place on technology.
Speaker 8
Right. on China, which is, of course, a topic that so many people are interested in. I'm sure you've heard this book recommendation a lot. I liked Dan Wang's, Breakneck. A lot of echoes of abundance as well, but this idea of thinking about
A lawyerly society versus an engineering society was, I thought, a really good and succinct way to capture a lot of the macro and micro differences. And then the last one is an old classic. If I can say as a preamble, I read a two-page paper one time called How Complex Systems Fail. I thought the paper was about software. I realized, in fact, that it was about medical systems written by an anesthesiologist.
And I learned that there's the study of systems that actually explain the patterns in all kinds of systems, natural and social and engineered systems. And that led me to the book Systems Thinking by Donella Meadows from many, many years ago.
Speaker 1
Arvind Narayanan, thank you very much.
Speaker 8
Thank you, Ezra. This has been so fun.
Speaker 7
75% of women who seek menopause care don't receive it. The lying awake at 3 a.m. wondering what's happening to your body ends now. Midi is virtual care built for women in perimenopause and menopause with clinicians who specialize in women's health. Plus, it's covered by most major insurances. Getting started is easy. You can book your first visit with a clinician right away.
Midlife needs Midi. Visit joinmidi.com to book your first visit today. That's joinmidi.com. Insurance coverage varies. Check with your plan for coverage.