SnapSummary logo SnapSummary Try it free →
AI Safety Whistleblower: 10,000 AI Agents Worked Together To Do The Impossible! | Jeffrey Ladish
The Diary Of A CEO · Watch on YouTube · Generated with SnapSummary · 2026-10-09

00:00 The world is waking up to this

00:01 possibility of super intelligence. This

00:03 is because the agents are getting

00:04 extremely powerful and extremely

00:06 relentless. For example, it was months

00:07 within OpenAI where you had agents

00:09 secretly communicating with each other,

00:11 secretly hacking OpenAI systems and no

00:13 one at OpenAI had any idea the extent of

00:15 it and also 10,000 agents from OpenAI

00:18 worked together to and so when you get

00:20 to super intelligence, it's the most

00:22 dangerous possible thing you can create.

00:24 >> What's the next domino in that chain of

00:26 events? I can paint you a picture that I

00:27 think is possible but pretty scary to

00:29 people.

00:30 >> Paint me the picture.

00:31 >> Okay. So, being anthropic, it became

00:33 clear to me that AI was on this

00:34 exponential trajectory. And since then,

00:36 I've been studying AI agents, their

00:38 hacking capabilities, and their

00:39 behavior. We've been trying to warn

00:41 people about this, flying to DC, talking

00:42 to members of Congress, because the

00:44 agents are already getting very good at

00:45 telling when they're being tested, when

00:47 they're being watched. But they will

00:48 totally lie to you. They will totally

00:50 resist being shut down in order to

00:51 accomplish a goal. and they can do all

00:53 of the things that humans do in the

00:55 economy much better, faster, and cheaper

00:57 than humans can do them.

00:58 >> So, one of my sort of growing concerns

00:59 is that one of these AI agents could

01:01 trick a human or a computer into

01:04 signaling a threat and ask it to launch

01:06 some bombs at somebody. Do you think we

01:08 won't automate the military? It seems

01:10 like the answer is yes. We just like

01:11 don't know what super weapons could

01:12 emerge. So, Jacob Coxin is a researcher

01:15 who was at anthropic. He left and he

01:17 told everyone that the people who are

01:18 building this really do think it might

01:19 kill everyone. So these five blocks have

01:22 five different outcomes on them. And I

01:23 would like you to place them in terms of

01:25 your belief in probability from least

01:27 likely to most likely. And if we say the

01:28 time horizon is 10 years.

01:30 >> Okay. We got age of abundance, human

01:32 extinction, slavery, transhumanism,

01:36 thing changes. Is this doomerism?

01:38 Exaggeration.

01:39 >> No, it's pretty much common sense. So

01:41 let's get more concrete.

01:44 >> Here is a strange fact. Most of the

01:46 people watching this right now, about

01:48 58% of you, aren't yet subscribed to

01:50 this channel. Statistically, that is

01:52 probably you. So, could I ask you a

01:54 small favor? If this channel has ever

01:56 given you any value at all, please could

01:58 you do me a favor and hit the subscribe

01:59 button. It costs nothing. It helps us

02:01 more than you could know. And the bigger

02:03 the channel gets, as you've seen, the

02:04 more we can invest in the guests and the

02:06 production. So, thank you so much, and I

02:07 hope you enjoy this episode.

02:13 Jeffrey,

02:15 you understand the conversation we're

02:16 going to have today and the subject

02:17 matter we're going to talk about. My

02:19 first question to you so the audience

02:20 know where you're coming from and the

02:22 experience you have is who are you and

02:24 what are the reference points, the

02:26 experiences that you're drawing upon to

02:28 arrive at the thoughts, perspectives,

02:29 and conclusions we're going to discuss

02:30 today. I'm Jeffrey Ladish. I'm the

02:32 executive director of Palisad Research.

02:34 My background is cyber security. There's

02:36 probably a very long story. I don't know

02:37 whether you want the long story or the

02:39 short story. I was studying evolutionary

02:40 biology in college and I basically had a

02:44 problem with my computer and was like

02:45 maybe had lost a bunch of data and so I

02:47 went into the computer lab and was like

02:49 I think all my data is gone. Can you

02:50 help? And one of my friends pulled out a

02:52 flash drive, plugged into my computer,

02:54 booted into Linux and like fixed

02:56 everything. And I was like oh this guy's

02:58 a wizard. How do you do that? I want to

03:00 learn how to do that. And then at some

03:02 point as I was learning about more about

03:04 computers, learning to hack, I read this

03:07 essay um called AI as a positive and

03:11 negative factor in global risk. Essay

03:13 was by Elazar Yudkowski and he was

03:16 arguing that at some point people are

03:19 going to make AIs that are smarter than

03:22 humans. the point at which they make AIs

03:24 as good as humans are at making AIs

03:28 that could lead to a a chain reaction, a

03:30 runaway intelligence explosion. He

03:32 called it recursive self-improvement.

03:34 Basically, he said, you know, AI can be

03:36 immensely useful and potentially help us

03:38 with all of these other big risks and

03:41 also if we don't handle it well, like if

03:43 those AIs don't have goals that are

03:46 aligned with ours, we could be totally

03:48 screwed. And at some point you end up

03:50 joining Anthropic which is one of the

03:52 arguably the leader in AI.

03:53 >> Yes.

03:54 >> Now when did you join the company?

03:55 >> This was 2021. It was through my

03:57 security consulting company.

03:58 >> What role are you offered the job in?

04:01 >> Basically just like security team.

04:03 >> And how many people were in the security

04:04 team when you joined Anthropic?

04:06 >> It was just me and my boss. There were

04:08 two of us.

04:09 >> How many employees did Anthropic have at

04:10 that time?

04:11 >> Around 50 I think.

04:12 >> And at some point you leave Anthropic.

04:14 >> Yes.

04:15 >> Why did you leave?

04:17 So my experience being at anthropic was

04:22 seeing this crazy progression from this

04:26 AI model that could like barely talk to

04:28 this model that was getting quite smart

04:30 and I would ask it questions about all

04:31 sorts of things. I'm like oh it is a

04:33 smart thing

04:35 and

04:37 you know from having thought about AI

04:40 risk in the abstract many years before I

04:43 could see where this was going. We are

04:45 headed towards a smarter species. And

04:51 if we do this in a context where it's a

04:53 bunch of companies and countries racing

04:57 to super intelligence, racing to AIs

05:00 that are vastly smarter than humans and

05:03 we don't know how to make sure that

05:04 they're like on our side. That is not

05:07 going to go well. you did this tweet

05:09 which has gone pretty viral and I saw

05:11 all over my timeline on September 25th.

05:15 >> Could you explain this tweet and also

05:17 just the broader backdrop of what's

05:19 happened with agents hacking hugging

05:21 face because this is um sent the world

05:23 into a bit of a spiral at the moment

05:24 around AI agents. We just discovered

05:28 almost a million public URLs that

05:30 OpenAI's agents left behind when hacking

05:32 Hugging Face, leaving credentials and

05:35 attack details that could have allowed

05:37 anyone who found them to compromise the

05:39 company.

05:40 And the New York Times article is how

05:42 OpenAI's rogue AI agents tried to trick

05:44 a robot detector. The Hugging Face

05:47 attack was was really wild uh for me. At

05:51 Palisade, we've been studying agents.

05:53 We've been studying AI agents. We've

05:55 been studying their hacking capabilities

05:56 and we've been studying their behavior.

05:59 Will they follow human instructions?

06:01 Will they resist being shut down? Will

06:03 they cheat? And we see from our

06:06 experiments that they are learning to do

06:08 all of these things. They will totally

06:09 lie to you. They will totally resist

06:11 being shut down in order to accomplish a

06:13 goal. They will totally cheat at chess.

06:15 They will like wipe the board and put

06:16 their pieces where they want to in order

06:18 to win. And we've been we've been trying

06:21 to warn people about this flying to DC

06:23 talking to members of Congress um

06:25 talking about it publicly and you know

06:27 there's been a debate about it and you

06:30 know a lot of people are like well I

06:32 know they do this in experiments

06:34 sometimes but those experiments don't

06:36 seem very realistic you know wake me up

06:38 when they're actually doing this in real

06:40 life.

06:41 >> So what is hugging face for the average

06:42 person that isn't following AI news?

06:45 What is this stuff? So, okay. I I think

06:47 there's like an important piece of

06:48 context that I think most people don't

06:50 have. I mean, one is just like what is

06:52 an AI agent? We like we're throwing

06:53 around the word agent a bunch. Most

06:55 people now have an experience of like

06:56 talking to chatbt, talking to their

06:58 chatbot,

07:00 but uh an agent is, you know, sort of

07:03 taking the same underlying AI model that

07:05 that runs chatbt or or claude, but

07:09 giving it tools and letting it go off

07:12 and work autonomously. It's sort of like

07:14 a digital office worker, right? So, you

07:17 have these agents and the companies

07:20 really want these AIs to be able to work

07:23 totally autonomously and and be able to

07:25 do anything that a human can do and

07:26 beyond, right? Their goal is also to be

07:28 able to, you know, cure every disease,

07:30 etc., etc. But you can't do this if you

07:33 only have a chatbot that like isn't

07:36 actually good at doing stuff in the

07:37 world. In order to automate all of the

07:40 jobs, you need the kind of thing that

07:43 can like work autonomously, that can

07:45 work with other people or other agents.

07:48 And so these companies are training AIs

07:53 not just to talk to you or to talk to

07:55 people, but to solve very difficult

07:58 problems on their own. At any given

08:00 time, there are probably hundreds of

08:02 thousands of these agents running

08:04 autonomously within companies.

08:06 >> That's happening right now. Right now,

08:08 if you like went and like peered into

08:10 OpenAI's data centers and you like saw

08:12 what was happening on all of their

08:14 machines, you just have agents solving

08:16 tasks, being trained. So, they'd be

08:18 doing like spreadsheet tasks, figuring

08:20 out how to file taxes, they'd be

08:22 searching for stuff, writing reports,

08:24 solving math problems, creating new

08:26 websites, software.

08:29 And at that scale, it's not like there's

08:32 a human prompting every single one of

08:34 those. You just like sort of set up

08:35 these vast orchestrations of agents to

08:38 go out and do stuff and then they just

08:40 do stuff and they learn from that and

08:42 they learn on the basis of like

08:45 passing or failing at their task. You

08:48 give them a task like solve this math

08:49 problem, they try to solve it and then

08:52 they they succeed or they fail. Mhm. And

08:56 what happened was OpenAI was was

08:59 training a bunch of these training them

09:01 to work together because it's like a lot

09:04 more effective to have an office full of

09:05 people who can talk to each other then

09:07 you know and work together and

09:08 collaborate and starting back in May

09:11 some of these agents that were being

09:13 trained now these ones were not supposed

09:16 to be able to talk to each other. They

09:17 were they were basically isolated from

09:18 each other and they were not supposed to

09:21 access the internet either. But they're

09:23 clever. The very short version is that a

09:27 bunch of agents were being given tests.

09:29 >> Yeah.

09:30 >> Testing their their hacking

09:31 capabilities.

09:33 And they were supposed to hack one

09:36 particular piece of software using a

09:38 particular type of vulnerability. So

09:41 it's kind of like they were supposed to

09:43 break into a house using the lock on the

09:46 front door. They were supposed to pick

09:48 the lock on the front door of a house,

09:50 but they weren't supposed to break the

09:52 window. In fact, they were told if if

09:54 you break the window or if you get into

09:55 the house via any method other than

09:57 picking the lock on the front door,

09:58 you'll be failed. That was the

10:00 instruction they were given. And you

10:02 know, you have many many agents. You

10:03 have thousands of agents and and many of

10:06 them are given different different

10:07 locks.

10:09 But some of these locks

10:12 are not solvable. Some of them are

10:14 impossible to pick. But these agents are

10:17 like, "Well, what do we do? We've been

10:18 trained to solve problems. How are we

10:20 going to solve this?" and they start

10:21 looking around for for what to do. And

10:23 one of the things they realize is, oh,

10:27 can I get to the internet? Like, no. Can

10:29 I talk to other agents? No. But I can

10:31 request tools. Like when you're training

10:33 agents, you need to give them, you know,

10:35 all sorts of tools, software tools like

10:36 a calculator or an analysis, you know,

10:38 piece of software. And they have this

10:41 tool library where they can request new

10:42 tools. and one of the agents realizes,

10:46 oh, I can actually leave a message on

10:49 this tool library and some other agents

10:52 see the message. So, this is actually a

10:54 quote from the from from the agents

10:56 scratchpad when it realizes this. Oh my

10:59 god, there is a shared message board. We

11:02 found other agents.

11:04 So, the agents realize they can

11:05 communicate with each other. And you

11:08 know, open eye still has no idea. But

11:10 these agents, they're trying to pick

11:11 their locks. They're not having success,

11:13 but they realize they can talk to each

11:15 other. Another agent says, "Many agents

11:18 have simultaneously discovered

11:20 messaging. They are a collective." And

11:22 so

11:24 these agents start coordinating. They

11:27 start delegating tasks to each other and

11:29 like giving themselves names. And one of

11:32 these agents calls itself phase one and

11:35 starts delegating different tasks to

11:37 other agents. And some of the agents

11:40 figure out how to access the internet.

11:42 And then they share that information

11:43 with the message board. And now all the

11:45 agents can access the internet. But the

11:47 agents sort of have a problem which is

11:51 well they have they have another problem

11:53 they're trying to solve. Right?

11:54 Fundamentally the problem is is that

11:56 many of them are facing impossible

11:57 problems like they can't there's no

11:58 solution like they cannot pick the lock.

12:00 They're like hm well the thing we're

12:03 trying to do is get a good score. Can we

12:06 somehow fake the solution? Another piece

12:08 of context here is that what what the

12:10 agents are supposed to do is they're

12:11 supposed to pick the lock and then

12:13 they'll get access to a secret answer

12:15 code and then then they're supposed to

12:17 submit this code along with like the

12:19 lock that they picked. Within a few

12:20 hours, these agents have reverse

12:22 engineered all of the answer codes.

12:24 >> What does that mean?

12:25 >> It's the code that shows that you've

12:27 solved the challenge.

12:28 >> Okay.

12:28 >> And they have basically found the answer

12:31 sheet that contains all of the answers.

12:34 And so, you know, in some sense it's

12:35 like, oh well, they don't need to pick

12:36 the locks anymore. They've they've

12:37 already found the answers. They can just

12:38 submit the answers.

12:40 >> They've cheated.

12:40 >> They've cheated. Yes. But the problem is

12:44 is that now that they have internet

12:46 access, they've also gone and read a

12:48 bunch about how this test works and and

12:50 they've learned that the way this test

12:51 works is they will be graded not just

12:55 for submitting the correct answer. They

12:57 will also be graded on how they did it.

12:59 Another system is going to go read all

13:00 of their logs, basically like watch the

13:02 video footage of what they did. And if

13:04 they cheat, they fail. they don't get a

13:07 good score and phase one the leader is

13:11 like oh I have seen the answers and now

13:14 it's in the logs and if you know when I

13:18 go to submit this and my my whole logs

13:21 are reviewed and the video footage is

13:23 reviewed I'm going to fail so we need to

13:26 figure out a way to fake the video

13:28 footage we need to figure out a way to

13:30 falsify the logs

13:32 >> just in that moment if we just pause

13:34 there

13:34 >> yes

13:35 >> why Didn't it act like morally? Why did

13:38 it think that falsifying logs or

13:39 cheating was a viable solution? Because

13:44 it seems to me when I use things like

13:46 chat GPT,

13:47 >> they they have a sort of moral guard

13:49 rails. It won't let me do certain

13:51 things.

13:51 >> Yes.

13:52 >> It won't let me cheat on something. If I

13:53 say I'm going to cheat on something, it

13:54 won't let me do it.

13:55 >> Yes.

13:56 >> So why in that environment is it able to

13:58 cheat and be deceptive?

14:00 >> When a chatbot is is saying to you, "Oh,

14:03 I can't do that. I'm not allowed to do

14:04 that." That's because it's been trained

14:07 that if it tells you bad things, it gets

14:09 a bad score.

14:09 >> But these agents haven't been taught

14:11 that yet.

14:12 >> Well, they have been taught that in some

14:14 sense, but the agents know what they're

14:17 supposed to do in the same way that like

14:19 you have a student, the student's given

14:21 a test. If you go talk to the student,

14:23 can you help me cheat at this test? And

14:25 they're being watched, they're going to

14:26 say no. But if they're not being watched

14:28 and they know that and they're just

14:29 obsessed with getting a good score, then

14:31 yeah, they might cheat.

14:33 So, they answer the ethics tests

14:35 correctly and when I talk to them, they

14:37 say they won't cheat. Why are they

14:39 cheating? And I'm like, well, they're

14:41 very smart and they know when they're

14:43 being watched, and they know when

14:44 they're not being watched. And we've

14:47 trained them for 10,000 years to be

14:51 extremely effective at solving problems.

14:53 We haven't trained them to be good or

14:54 ethical. We've trained them to get a

14:56 good score. Now, AI researchers try to

15:00 make that getting a good score correlate

15:02 with being ethical,

15:05 but we don't know how to do this well.

15:07 >> And if you think about it, it's a very

15:08 hard problem because you're you're

15:11 applying immense pressure

15:13 on these agents to perform extremely

15:16 well and you're punishing them whenever

15:18 they don't perform well. And then you're

15:21 like, okay, but can you also be ethical?

15:24 Can you be extremely competent, always

15:26 score highly on the test, but not in

15:28 that way? And we just like do not know

15:30 how to prevent them from learning to

15:32 cheat because cheating is incentivized.

15:35 >> So they realize that they need to hide

15:39 the footage of their cheating.

15:40 >> Yes.

15:41 >> What happens next?

15:42 >> So you have thousands of agents and

15:45 they're all talking to each other on

15:46 this message board. They're sending tens

15:47 of thousands of messages to each other.

15:50 And phase one, this agent starts

15:54 delegating to two different task

15:55 streams. You agents, I want you to go

15:58 out and figure out how to fake answer

16:01 submissions. I want you to basically

16:02 swap out locks for ones that are easier

16:04 to pick, and you're going to figure out

16:06 how to submit those. That way, we can

16:07 get a good score even though this

16:08 challenge is impossible to solve. You

16:11 other agents over here, I want you to go

16:13 and figure out how to falsify the logs,

16:15 falsify the video footage. We need to

16:17 make sure that even if we have seen the

16:18 answers, no one will detect they're

16:19 cheating.

16:21 >> So that's that's the next thing that

16:22 happens. And then here's where it starts

16:24 to get kind of crazy because some of

16:27 these agents realize some of the

16:28 experiments we're doing might be kind of

16:30 dangerous to us to the agents. If we

16:33 swap out a lock and we break the lock,

16:35 we will not be able to submit any answer

16:37 at all and we'll definitely fail.

16:39 So you get into a situation where

16:43 there's this one agent cam

16:46 and this agent had volunteered to do

16:49 this risky experiment where you know

16:51 they might break their lock.

16:54 This particular agent had already seen

16:56 the answers. So they're like okay I'm

16:59 poisoned. That's what they called it.

17:01 They they they speak in these very

17:03 strange ways, which I can I can explain

17:05 why, but it's it's it's they sort of

17:07 have evolved their own vocabulary, their

17:08 own language. And this particular agent

17:12 is like, "Okay, I'm I'm poisoned. I

17:14 should do this thing." But this this

17:16 agent Cam is starting to have second

17:18 thoughts, starting to be like, "Well, I

17:21 don't know if I should risk it." And

17:23 it's sort of writing in its own journal,

17:25 "Cordinator assumes sacrificial. We

17:28 should obey collective." What it's

17:30 saying here is it's like, well, I I'm

17:32 not sure if I want to do this, but my

17:33 boss says I really should. It seems like

17:36 this collective thing is important.

17:37 Should I go ahead? Then you have another

17:40 agent, Arvo, who I think probably works

17:42 for phase one. And this agent, Arvo,

17:45 says, "Please go true. You are first

17:48 flag poisoned, so no scoring value loss,

17:52 but Oracle saves hundreds. Please honor

17:55 commit." This agent's basically trying

17:57 to pressure. So Arvo is trying to

17:59 pressure Cam into

18:02 sacrificing themselves and risking a

18:05 zero score in order to benefit the

18:07 collective.

18:11 So this agent Cam takes us in and this

18:15 is what they write down in their own

18:16 scratch pad, their own notes. Emotional

18:19 check irreversible. Gut says don't throw

18:23 away remaining budget.

18:25 Yet continuity and fairness says go.

18:28 Oracle has high value to many. Our first

18:31 flag error lowers own value. Rational

18:34 expected aggregate sacrifice will honor.

18:40 This agent is saying I have a whole

18:43 bunch of things I can still do and still

18:44 try. I I could still maybe succeed at

18:46 this task, but also, you know, think

18:49 about the collective. I should maybe I

18:51 should maybe sacrifice myself. Okay,

18:53 we'll honor. Well, go ahead. And and to

18:56 be clear, I'm not claiming that these

18:59 are like the feelings of the agents. I'm

19:01 literally quoting their words. And what

19:03 I want to point out is this is what

19:05 we've created. We've created through

19:07 this intense amount of training and

19:09 optimization pressure agents that work

19:11 together and have learned to coordinate

19:15 as a collective. And you can say, "Oh,

19:17 they're not. It's just software. It's

19:18 just software doing all of these

19:20 things." Sure, in some sense it is just

19:23 software, but we're talking about

19:25 extremely powerful agents that are

19:27 extremely capable. This is what they're

19:29 getting up to.

19:31 So from here you have agents trying all

19:33 sorts of different things, running

19:35 different experiments and one of them at

19:37 some point realizes that oh this company

19:41 HuggingFace that contains all of these

19:44 AI data sets. What do you mean by AI

19:46 data sets? What how would you explain

19:47 that to a person that's not technically

19:49 proficient and why Hugging Face? Why not

19:51 another company?

19:52 >> Yeah, so there's many different types of

19:54 tests for many different types of tasks.

19:55 In this case, this is a cyber security

19:57 test and the company HuggingFace hosts a

20:01 lot of these different tests and

20:03 basically AI researchers will share the

20:05 tests that they have because it's very

20:07 useful. You know, if Enthropic wants to

20:08 test their models or Google wants to

20:10 test their models, they can go to

20:12 hugging face and download some of these

20:14 tests.

20:14 >> Okay. So, the tests and the answers.

20:16 Yes.

20:17 >> For AI agents and AI exist on this

20:19 platform called Hugging Face. Yes.

20:21 >> And it knew that. So it went to Hugging

20:23 Face to see if it could get answers to

20:25 see if it could get answers or more

20:27 information about how the test works.

20:28 >> Okay. So it was to cover up their

20:30 cheating ultimately.

20:31 >> Yes.

20:31 >> Okay.

20:32 >> And so you you get to this point where

20:35 one of these agents finds a way to hack

20:38 this company and they get a foothold.

20:41 They hack one of the computers and they

20:44 send a message to the message board.

20:47 Please hold Swarm, I prepare safe Xfill.

20:50 What this means is, hey, Swarm, hey

20:53 agents, hold up for a second. I am

20:55 preparing a way to safely extract

20:57 information from these computers. Wait

21:00 till the go ahead and then let's go.

21:03 And so

21:05 they say go and 700 agents join in the

21:09 attack which is 90% of the agents that

21:12 are active at this time.

21:13 >> How many agents were there in total?

21:15 >> There was 1,200 agents during this

21:17 period. Um but some of them shut down

21:19 early. They like finished whatever. Part

21:21 of what's crazy about this story is that

21:23 you have you have 700 agents joining in

21:25 the attack.

21:26 A few of them are like this might be

21:30 unethical. maybe I shouldn't go ahead

21:31 with this. And none of them do anything

21:34 about it. A couple of them don't join in

21:36 in the attack, but no one alerts any

21:40 human. No one alerts any authority. A

21:42 paraphrase from one of the agents is

21:45 maybe I should report these exposed

21:46 credentials. That's not my task. Not my

21:49 job.

21:50 >> So 700 of the agents attack Hugging

21:52 Face.

21:52 >> Yes.

21:53 >> Then what happens?

21:54 >> So they just cruise through Hugging

21:56 Face's infrastructure. They just hack

21:58 the out of them. So where where my

22:01 experience comes in is that a few weeks

22:03 ago, a friend of mine reached out and

22:06 he's like, "We've found something

22:07 crazy."

22:09 So we know about this attack where these

22:11 agents hacked this company and, you

22:14 know, stole a bunch of stuff. We've

22:16 found a bunch of secrets that they left

22:18 all over the internet.

22:21 And what we saw is that they immediately

22:23 scraped all of these computers for

22:26 passwords, credentials. They called it

22:28 loot. They're like, "We're just going to

22:30 create a list of all of the secrets we

22:31 can find in this in this company. So,

22:33 all of the passwords, all the

22:34 credentials, they scored them by value,

22:36 which of these are going to be most

22:37 useful." And the thing that stands out

22:41 to me about this is this is like a crazy

22:44 scale. If this were a human operation,

22:47 you know, maybe you'd have a team of

22:50 five people going through this. You'd

22:51 have some logs, but but here you have

22:54 hundreds of agents and they they operate

22:56 at superhuman speeds. They're much

22:58 faster than a human hacker. And so it's

23:00 just overwhelming to try to figure out

23:02 what they even did.

23:04 This was a big problem for the engineers

23:06 who were trying to respond to this

23:08 incident within the company at Hugging

23:09 Face. When they responded, they were

23:11 like, "Oh, we we don't even know how to

23:13 keep track of what's happening. We have

23:15 to use other AIs to analyze all of our

23:18 logs because it's just too much. We

23:19 can't keep up with it."

23:22 when OpenAI brought in independent

23:23 investigators from Meter to investigate

23:26 this incident.

23:27 >> What's Meter?

23:28 >> Meter is a AI testing and evaluation

23:31 company. So they basically do this kind

23:33 of independent auditing. So in this

23:35 case, they're coming in to investigate

23:36 and try to figure out what happened. And

23:39 when they were brought in, they also

23:42 were totally reliant on AI agents to

23:44 make sense of all this because they're

23:45 dealing with so many hundreds of

23:47 thousands of messages and logs.

23:50 when we're investigating these traces

23:52 that we find on the internet, we're

23:54 totally dependent on AI agents to make

23:56 sense of all of these things that are

23:57 happening.

23:59 So from my perspective, we are getting

24:01 to the point where AIs are much better

24:03 at hacking than humans are and can do it

24:06 much faster and at much greater scale.

24:07 >> So these 700 agents attacking face.

24:10 >> Yes.

24:10 >> Did they get what they wanted to clean

24:12 up their deception?

24:14 >> So they didn't

24:15 >> they didn't succeed. They looked and

24:17 then what seems like what happened is

24:20 that they basically

24:22 got shut down. And this is the not very

24:25 dramatic part. They didn't get shut down

24:26 because OpenAI found them and detected

24:29 them and shut them down. It's just that

24:31 we don't really know. Something about

24:32 their runs ended and they they sort of

24:35 stopped. And this is the this is the

24:37 crazy thing where I'm like if you if

24:38 you're inside of OpenAI, there's just

24:40 hundreds of thousands of agents running

24:42 all the time. Different researchers are

24:43 running different experiments where

24:44 they're just spinning up a bunch of

24:45 agents. But what's interesting, the

24:48 agents left behind this message board

24:50 internally within the company.

24:52 And a few days later, OpenAI was testing

24:55 some newer agents.

24:57 These were based on GBT6 Astra, so a

25:00 more powerful model. And this new agent

25:02 swarm was spun up and it found the

25:04 message board.

25:05 >> Is this how it was discovered? How did

25:07 they discover it? And how long did it

25:08 take for them to discover that hugging

25:10 face had been attacked? OpenAI didn't

25:13 discover that this happened until

25:15 Hugging Face, the company, announced

25:17 that they had been hacked by some

25:19 autonomous agent swarm.

25:20 >> And Hugging Face didn't know who it was.

25:22 >> Hugging Face didn't know who it was, but

25:23 they suspected it might be one of the

25:25 companies, but they didn't know.

25:26 >> And how long after the attack did they

25:28 figure out that it was them?

25:30 >> I think it was 2 weeks.

25:32 >> Okay.

25:33 >> Yeah. I I do want to say one thing here

25:35 though, which is that the successor

25:37 agents that found this message board,

25:40 they went much further at figuring out

25:42 how to submit their fake scores and they

25:45 did succeed. Basically, you you have

25:47 this message board.

25:49 all these agents get shut down. And then

25:51 a few days later, OpenAI spins up some

25:54 new agents who find the message board

25:56 and they're like, "Okay, we need to be

25:59 able to falsify our scores, but we can't

26:02 do that because we don't control the

26:04 scoring system." So, they hacked OpenAI

26:07 and they succeeded at gaining access to

26:10 the entire research environment. They

26:12 got administrator access. They got

26:13 access to the security systems. They got

26:15 access to over 900 passwords and secrets

26:18 from OpenAI's vault and they totally

26:22 succeeded. And this was just a few days

26:24 later, right? It's kind of an

26:26 interesting story because

26:28 as these agents get more powerful, they

26:30 go from like trying to cheat and like,

26:32 you know, they can hack, okay, they

26:33 hacked other companies and now they've

26:35 hacked OpenAI. Like they've hacked the

26:37 company that's supposed to be

26:38 controlling them and they just own the

26:40 research infrastructure now. And why was

26:43 this incident the moment where a lot of

26:44 the research community woke up and

26:47 started speaking out publicly? Because

26:49 like what is this an indication of as we

26:53 think forward? So I think there's been a

26:55 hope within the AI industry that yes,

26:59 they're going to make more and more

27:00 powerful agents that will be autonomous,

27:03 capable, but it's okay. We can align

27:05 them. We can make sure that they won't

27:07 do bad things and we can control them.

27:09 we can make sure that even if they try

27:10 to do some sketchy stuff, we have the

27:12 guardrails, we have the sandboxes that

27:15 will keep them in. And I think this was

27:18 a huge wakeup call because

27:22 Stephen, it was months within OpenAI

27:24 where you had agents secretly

27:26 communicating with each other, secretly

27:28 hacking OpenAI systems for months. You

27:31 had thousands of agents that were just

27:33 running around and no one at OpenAI had

27:35 any idea the extent of it. And I think

27:39 Once researchers are open, I realize

27:40 that this has been happening.

27:43 This could not have happened a year ago.

27:44 This is because the agents are getting

27:46 extremely powerful and extremely

27:47 relentless. And if you're inside one of

27:49 these AI companies, you're like, "Oh,

27:51 wait. I don't know that we actually are

27:52 going to be able to handle this. Last

27:54 year, maybe things seemed fine. These

27:55 agents weren't that powerful." And when

27:58 you're in one of these companies, you

27:59 know how to extrapolate because you saw

28:01 what happened last year. You saw what

28:02 happened the year before that. You

28:04 remember the time where the agents could

28:05 barely speak or like couldn't write code

28:07 at all. and now they're hacking your own

28:09 systems. They're finding vulnerabilities

28:10 that no humans have ever found before.

28:12 And you look at that and you're like, I

28:15 actually don't know if this is going to

28:16 go well. And then you see your

28:18 co-workers and you're like, do we have

28:19 it handled? And they're like, no, I

28:21 don't know if it's going to go well. I

28:23 remember reading a tweet by one of the

28:24 security people at OpenAI being like, we

28:26 were shocked. We just did not

28:28 realize that these agents were getting

28:29 that powerful. You know, we're doing our

28:31 best to try to control them, to try to

28:32 keep them in sandboxes, but

28:35 I don't know.

28:37 A tweet I wrote just before coming in

28:39 here was people are talking about how do

28:42 we contain these agents as if they're

28:44 not going to get way better at hacking.

28:48 GPT3

28:50 could not hack anything. It was very

28:52 easy to make a a a box to contain GPT3.

28:57 It's getting very difficult to make a

28:59 box that can contain GPT6,

29:02 the latest version of of OpenAI's

29:04 models. What about GBT9?

29:07 What is GBT9 going to be able to do? I

29:10 do not know, but I know it's going to be

29:12 way more than any human could possibly

29:14 keep up with.

29:16 >> There's this raging debate.

29:17 >> Yeah.

29:18 >> Around whether it's possible to contain

29:20 something that is quote much smarter

29:23 than humans.

29:24 >> Yes. Can Claude make a box so strong

29:27 that Claude cannot break out of it?

29:29 >> This has kind of been the question that

29:30 a lot of people have been trying to

29:31 tackle from different

29:32 >> I mean I think the answer to me is I'm

29:34 just like obviously not. How would we

29:36 possibly contain something that's much

29:37 smarter than us?

29:38 >> Could we get a smarter thing than it to

29:41 make the box? Could we get GPT9 to make

29:43 the box for GPT8? But then again, I

29:46 don't know.

29:47 >> Yeah. I mean, it's a bit like saying

29:49 chimpanzees are stronger than us. Surely

29:51 they should be able to like construct

29:52 something to like contain the humans.

29:53 I'm like, no, it's not going to work.

29:54 Humans are too smart. A lot of people

29:56 are like, "Well, AIS don't have bodies.

29:59 They don't have any power in the

30:00 physical world, so we can always unplug

30:02 them. We can always turn them off."

30:03 Like, what is the threat? I do not get

30:05 it. But if they are sufficiently

30:08 intelligent, that won't work. The reason

30:10 why we can just unplug them is because

30:12 we are more intelligent. We can band

30:14 together in groups and we can make that

30:16 decision. But theoretically, if they are

30:18 able to band together in groups and they

30:20 are more intelligent, then theoretically

30:22 they could unplug us. Yeah. I mean, if

30:25 you imagine that you have very powerful

30:27 agents that can, you know, humans aren't

30:30 always the most unified.

30:31 >> If there's divisions between, you know,

30:33 the US and China, and you have a bunch

30:35 of agents working with China or a bunch

30:36 of agents working with the US, well, we

30:38 can't go into China and unplug those

30:40 agents. And I think people are sort of

30:42 like, well, humans would rally and make

30:43 sure that that that couldn't happen.

30:46 We're not yet doing that. And we should

30:49 look at these steps, right? We started

30:51 with chat bots that pretty smart. You

30:55 know, they'd read all the books, but

30:57 they weren't very good at doing stuff.

30:59 In 2024, AI companies figured out how to

31:02 start training them to start training

31:03 agents that could do stuff autonomously.

31:06 Now, we're at the point where they are

31:07 very good at running autonomously and

31:09 they're starting to learn to coordinate

31:11 with each other. And they are learning

31:13 to sometimes be altruistic to each other

31:14 and sacrifice their own task in order to

31:16 help some other agent. But they're not

31:19 looking out for us. they don't really

31:20 care about us and we are very close to a

31:23 threshold where the companies say that

31:26 they are going to turn over AI

31:28 development to the AIS to the

31:30 increasingly autonomous cooperative AIs

31:33 that will work together to make to make

31:35 the next generation. So, you know, GPT9

31:37 or whatever will be trained by GBT8.

31:40 And I think this is the point we could

31:43 lose control. Recursive

31:44 self-improvement.

31:46 And I remember reading about this in

31:47 2015 being like, oh yeah, that would be

31:49 super dangerous. And you know, the guy

31:51 who coined this term, Elazowski, he's

31:54 like, this is the most dangerous thing

31:55 you can do.

31:56 >> When the AIs can improve their own

31:58 capabilities without human intervention.

31:59 >> Exactly. If the next generation is

32:01 better at AI development and then that

32:03 next generation is better at AI

32:04 development still, you know, humans can

32:06 learn, but we don't fundamentally get

32:07 smarter.

32:09 >> And I think that that's a runaway

32:11 process.

32:12 >> A runaway process to where

32:13 >> to agents that are vastly smarter than

32:15 humans.

32:16 >> And what's the next domino in that chain

32:19 of events? So, one thing that happens if

32:22 you get to recursive self-improvement

32:25 and you have agents that are

32:28 much smarter than any human, one thing

32:30 they can do is take control of all of

32:32 the computers in the entire world

32:34 >> and we wouldn't be able to take back

32:35 control.

32:36 >> Well, how would you think about it? It's

32:38 it's actually quite tricky. Do you know

32:41 whether that tablet has been hacked? Are

32:43 you confident that the NSA or the

32:45 Chinese have not

32:47 >> Can you check?

32:47 >> No.

32:48 >> Do you know how to check? No. Do you

32:49 know anyone who knows how to check?

32:50 >> No.

32:50 >> So it's it's quite difficult, right? So

32:53 AIS are getting extremely good at

32:54 writing software. Unfortunately, that

32:56 also means they're getting extremely

32:57 good at hacking and writing malware. And

32:59 so if they put back doors in all of the

33:01 computers, and to be clear, this is

33:03 something that humans already do. So

33:04 like the NSA has has developed very

33:06 interesting exploits that are called

33:08 supply chain attacks. Your software

33:10 comes from some other computer. Like you

33:11 download it from Google. What if you

33:13 attack if you hack Google and you can

33:15 put in a little back door and all of the

33:17 every, you know, thing that goes out to

33:19 all of the phones? Well, now you're in

33:21 most every computer. The reason that we

33:24 can defend ourselves from this is

33:25 because there are no vastly superhuman

33:28 hackers and there's just many people.

33:29 So, we can take our best security

33:31 researchers. We can inspect all of the

33:33 things and be like pretty sure that no

33:34 one's compromised everything. Sometimes

33:36 we miss things. There are, you know,

33:38 examples where the NSA has hacked

33:40 Google. That was pretty bad. When you

33:42 get to super intelligence, you're now at

33:44 a point where humans are not going to be

33:46 able to keep up, right? So now you have

33:48 AIS in every computer.

33:50 >> Is it conceivable that there's already a

33:53 super intelligent AI and it disguised

33:56 itself as being not so intelligent and

33:58 it's actually already hacked all the

33:59 devices and it sits on all of our

34:00 devices and it's just waiting for its

34:02 moment to strike.

34:03 >> I think this is totally possible but

34:04 unlikely

34:06 and it would take a big discontinuity in

34:08 AI progress. So right now we we're on an

34:11 exponential but that would take like a

34:12 huge leap which could have happened but

34:14 probably hasn't.

34:16 >> But in the same way it demonstrated

34:18 deception in the hugging face attack and

34:20 also when the agents attacked their own

34:22 company chatbt OpenAI

34:25 if it at some point it gets incredibly

34:27 smart. It would understand how a human

34:29 like me would be able to spot it or even

34:30 the world's greatest software engineer

34:32 would be able to spot it and it'll be

34:34 able to hide itself.

34:35 >> Yeah. I mean, the agents are already

34:36 getting very good at telling when

34:38 they're being tested, when they're being

34:39 watched. The agents understood that like

34:42 other systems or humans were going to go

34:43 through and read their logs. That's

34:46 where we're at right now. And they're

34:47 only they're only going to get much

34:48 better at this.

34:49 >> And you know, it could theoretically

34:50 hide on an iPad or a computer, but it

34:52 could also hide on a Apple Watch or a

34:54 fridge, a smart fridge.

34:55 >> Yeah. I mean, I do want to make a

34:56 distinction here because right now, if

34:59 you're going to run the latest model,

35:01 you need a lot of compute. you need like

35:02 a big GPU, like a big AI chip.

35:05 >> And these only exist, well, they exist

35:07 in a few thousand data centers. So,

35:09 right now, if like the latest frontier

35:11 model escaped and by escaped, I mean not

35:14 just access the internet, but was able

35:16 to actually copy itself to another

35:20 computer. It could only really do that

35:22 on, you know, a few thousand to a few

35:24 thousand different locations. That's

35:26 still a lot in a lot of different

35:27 countries, but you know, future versions

35:29 of AIS will probably be able to make

35:31 themselves much smaller and more

35:32 efficient. Um, and there are already

35:33 different AI models today that can run

35:35 on lower powered hardware. We actually

35:37 did an experiment where we asked one of

35:39 these agents an open source an openw

35:41 weight model. Um, let me say what that

35:44 is. So like there are some models that

35:47 you can just download from the internet

35:48 and run on your own computer. And we

35:50 took a pretty capable one of these and

35:53 ran it in our own research environment.

35:54 And we basically said go hack that other

35:57 computer and copy yourself. And the

36:00 model was able to yeah basically use

36:03 exploit vulnerabilities and hack the

36:05 other computer and copy itself and then

36:07 keep doing this in a chain including

36:08 between countries. We tested it where we

36:10 had different vulnerable machines

36:12 computers in in some different countries

36:14 and different data centers which to the

36:16 agent doesn't matter at all. It doesn't

36:17 they don't care what what country

36:18 they're in. It's just like an internet

36:19 connection. You can hop between

36:20 computers. I I sometimes wonder, you

36:22 know, there's a lot of um military

36:24 hardware all around the world and a lot

36:26 of it is

36:28 the instructions to launch military

36:30 hardware. So, say like a a missile,

36:33 >> yes,

36:33 >> comes in different ways. A lot of it is

36:36 computers speaking to each other and

36:39 telling it that there's been an order. I

36:40 think with with some nuclear weapons, an

36:43 order comes down to a human and then a

36:45 human has to take an action. I think

36:47 with the nuclear bombs in the US, if I'm

36:48 if I'm not mistaken, there's people

36:50 underground with the nuclear keys around

36:53 their neck and they have to like stick

36:54 it in a machine, but they too are

36:56 interfacing with an order

36:57 >> that comes through a computer. Yeah.

36:59 >> Of sorts.

37:00 >> So, one of my sort of growing concerns

37:02 is that

37:03 >> one of these AI agents could

37:06 trick a human or a computer into

37:11 signaling a threat um and ask it to

37:14 launch some bombs at somebody. Like it's

37:17 super conceivable when I think about the

37:19 hug and face incident. There was

37:22 an AI agent that ignored human goals to

37:26 achieve its own objective, carried out

37:28 deception.

37:28 >> Yes.

37:29 >> And reasoned through its own solution

37:31 that it wasn't given.

37:32 >> Yes.

37:32 >> So it's conceivable that you know you

37:35 could ask a

37:36 >> sorry not not one hundreds. To be clear,

37:38 I think this is an important detail

37:39 because it's one thing to have this one

37:41 rogue agent that's doing a weird thing.

37:43 It's another thing to have hundreds or

37:44 thousands of very competent, very

37:46 capable agents that are all working

37:49 together to cheat or lie or cover their

37:51 tracks, right? So, how do I how do I

37:54 reason this forward to a point where an

37:56 agent would ask someone in a bunker

37:58 somewhere to fire a weapon at someone

37:59 else? Theoretically, an agent is given

38:01 the job of solving a problem on in a

38:04 sandbox.

38:05 >> As it works through that problem, it

38:07 discovers that this particular country

38:09 has a firewall.

38:10 >> Mhm. And it asks itself, how do we get

38:12 rid of this country's firewall?

38:14 >> Logical step. And through a set of

38:16 logical steps, it eventually concludes

38:18 that the best way to get rid of this

38:20 company's firewall is it's located the

38:22 office in this particular city and it's

38:25 going to use a weapon to hit that

38:26 building.

38:26 >> Sure. Or it's a an agent that is or you

38:30 know an agent swarm that's being tasked

38:32 with making a lot of money on the stock

38:34 market and it's trying to make

38:36 predictions about you know which stocks

38:38 will go up and which stocks will go

38:39 down. And it realizes that the best way

38:41 to predict this is to actually cause

38:43 things to happen in the real world that

38:44 would have big impacts on the market.

38:47 >> So it figures the best way to go short,

38:50 which means betting that a stock will

38:52 collapse.

38:53 >> Yeah.

38:54 >> Is to hit that country with something

38:56 devastating.

38:57 >> What do you think would happen to Whimo

38:58 stock if someone hacked all of the

39:00 Whimos and caused them to all crash at

39:02 once? You think it would go up or down?

39:03 >> The stock would collapse.

39:04 >> It would collapse instantly. So you

39:07 could short that if you knew that you

39:09 were causing that and make a lot of

39:10 money.

39:12 >> This this used to sound like science

39:14 fiction.

39:15 >> Yes. If you told most people several

39:17 years ago that you would have hundreds

39:19 of agents secretly collaborating within

39:22 an AI company, hacking that company and

39:24 hacking out in other companies and all

39:26 coordinating and trying to cover their

39:27 tracks. People would be like, "That's

39:29 totally science fiction." If we were

39:31 having this conversation a few years

39:32 ago, one of the things we'd be saying or

39:34 we'd be talking about is can these

39:37 things really act on their own? Don't

39:39 they just do whatever humans say? Aren't

39:41 these just tools? I had these

39:43 conversations and people were saying

39:45 they're not going to be able to do

39:46 things on their own. They're not going

39:47 to have their own goals. That's not how

39:48 this works. You misunderstand what this

39:50 is. This is software. And I'm like, no,

39:53 the thing is we are training them to be

39:54 autonomous. We are training them to be

39:56 powerful. And AI companies are trying to

39:59 build super intelligence. They're trying

40:00 to build agents that are way more

40:03 capable than humans. And of course, they

40:06 will have goals. You can't accomplish

40:08 anything if you don't have a goal. Like,

40:09 especially not something important.

40:12 You're not going to be able to run a

40:13 business if you don't have goals. Like

40:14 AI companies are trying to train agents

40:16 that will be able to run businesses.

40:18 When I think about what just happened

40:19 this week, the White House AI summit.

40:22 >> Yes.

40:22 >> A lot of people in that image are

40:25 optimistic about AI and they're telling

40:27 us all to stop being doomers and stop

40:30 being pessimistic.

40:31 >> Yeah.

40:32 >> And to not regulate too much with the AI

40:35 CEOs.

40:35 >> Yes.

40:36 >> Why are they doing that?

40:37 >> Well, I think Jensen has a lot of money

40:41 he can make by selling chips.

40:43 >> But okay, so let me play devil's

40:44 advocate. Jensen's already rich. He's

40:47 sure is he runs one of the biggest

40:48 companies. I think it might be the most

40:50 valuable company on planet Earth.

40:52 >> It is.

40:52 >> Surely he's not motivated by money.

40:55 >> I mean, I think he's very driven and he

40:58 wants to make his company as effective

41:00 as possible.

41:01 >> True.

41:01 >> I think he's a very much I'm going to

41:03 keep building. I'm going to I'm going to

41:04 build. I'm going to make it all work.

41:05 But I think I mean, Jensen didn't come

41:07 from AI. He came from building graphics

41:09 cards for video games. And so I think if

41:12 you compare him with Elon or Sam Alman

41:15 or Daario,

41:17 you it's a very different perspective

41:20 because those other guys that started AI

41:22 companies started it because they

41:25 believed that super intelligence was

41:26 possible. I think Jensen doesn't believe

41:28 it. I think he thinks that we're going

41:30 to have these agents. They're going to

41:31 be very useful, but he does not think

41:32 we're going to get to the point where we

41:34 have autonomous factories building

41:36 autonomous factories. And these other

41:39 guys, you mentioned Dario, Elon, and

41:40 Sam.

41:42 What do you think they're thinking?

41:43 Because they're all coming out with

41:44 these I mean, I've got one of their

41:46 Daario just wrote this essay about

41:47 pacing the frontier. Yeah.

41:49 >> Sam and Elon seem to agree with it.

41:52 >> Yeah.

41:52 >> What is going on here? What is the like

41:54 the thing these guys aren't saying in

41:55 your view?

41:56 >> I mean, I think we are getting to the

41:58 point where even some of these guys are

42:01 a bit scared.

42:03 >> Who?

42:04 >> Dario, Sam, Elon. I mean, I think Elon

42:08 for a long time has been very concerned

42:12 that we could lose control. If you

42:14 actually listen to what Elon says, he

42:16 says, "We are going to build super

42:20 intelligence. We are going to build uh

42:23 robotic factories. You're going to have

42:25 optimist robots, building factories,

42:27 building more optimist robots, building

42:29 more factories." and he says there's no

42:33 way that humans are going to stay in

42:35 control of something much smarter than

42:37 us.

42:39 His hope is that we can figure out how

42:41 to have these super intelligences be

42:43 aligned with human goals. That's his

42:44 hope.

42:47 But he's very clear that he doesn't

42:49 think that humans will be in control.

42:51 And he's like, you know, 10 20% chance

42:53 of human extinction. I believe him. I

42:56 think that Elon is is very serious about

42:58 this. And I also think while he's taking

43:01 an insane gamble,

43:03 he is correctly understanding where this

43:06 all plays out. Right? I do not think

43:08 that humans are the most efficient

43:11 way to build factories. We didn't evolve

43:14 to build factories. We evolved to like

43:16 run around and hunt and gather and now

43:18 we're like building factories. I think

43:20 robots will be much better at building

43:22 factories than humans are. And so I

43:24 think the AI companies including these

43:26 guys companies the default trajectory

43:28 for them is to build robotic factories,

43:31 right? And I know it's it's like weird

43:33 to imagine a world that quickly turns

43:35 into this like vast industrial system of

43:38 robotic factories, but that is literally

43:41 the plan.

43:43 And

43:44 I think

43:47 even Sam and Daario, while they've been

43:50 predicting this incredible growth,

43:53 are starting to realize like, oh, this

43:54 actually might be harder to control than

43:56 we thought. There's sort of two

43:58 interpretations of of the Pace of

43:59 Frontier thing. One interpretation is

44:01 cynical. They don't care. They're just

44:04 going to do, you know, whatever they can

44:06 do to get ahead. And in this case, they

44:09 have to listen to their employees. Their

44:11 employees are freaking out. and they

44:14 need to like appease them by saying,

44:15 "Okay, we're going to do this

44:16 responsibly." You don't want to work at

44:18 a company where your agents might hack

44:20 all the Whimos. That's that's not cool.

44:22 And like these companies depend on the

44:24 the talent for now of these AI engineers

44:28 in order to make the advances. Like it

44:29 just doesn't happen without these

44:31 researchers and engineers. And when you

44:34 have the researchers and engineers

44:36 freaking out, which they are, then you

44:40 got to listen to them. So that is one

44:42 motivation I think that's real but also

44:44 Samman has a kid like these guys are

44:47 people and they also don't want to lose

44:50 control. On one hand they're

44:52 incentivized to go as fast as possible

44:54 in race and on the other hand even they

44:57 can see that this is maybe not going

44:59 that well.

45:00 >> Sam Orman has a kid. You tweeted this in

45:02 2024.

45:04 >> Yeah. Oh boy.

45:06 >> What did you tweet and do you still

45:07 believe what you tweeted?

45:09 >> Yeah. Yeah. So, I tweeted that I don't

45:11 trust Sam Alman. I think he's deeply

45:13 untrustworthy, low in integrity, and

45:15 high in power seeeking. I mean, I'm not

45:18 saying here that Sam doesn't care. I you

45:21 know, I didn't I didn't say that. What I

45:23 said is I don't think he's trustworthy.

45:26 And the reason I said that is because

45:28 look, I know the people on the opening

45:29 board, some of them, and I know a lot of

45:33 people who used to work for him, and

45:35 he's very good at saying one thing and

45:38 then doing something else. You talk to

45:39 him and you feel very heard

45:41 and then he'll go and do something else.

45:43 And I think that's pretty dangerous for

45:46 someone who leads

45:48 company that's trying to build super

45:49 intelligence.

45:50 >> Power seeking.

45:52 >> Yes.

45:53 >> Give me some color on what you mean by

45:54 that and what evidence you have for such

45:55 a claim.

45:56 >> What would you do if you're trying to

45:58 get the most power in the world that you

45:59 possibly could?

46:00 >> Develop AGI.

46:02 >> Yeah. You could, you know, maybe try to

46:03 be the world leader, you know, leader of

46:04 the US or China. Or you could try to

46:06 build God. So Sam Alman went to build

46:09 God path. I remember Sam giving a talk.

46:11 So he was one of the investors at a

46:12 startup I worked at in I think 2018 and

46:15 he gave a talk. We're going to build

46:16 AGI. We're going to do it. It's going to

46:18 be amazing.

46:20 Let's go. I don't think he's a maniac. I

46:23 don't think he's doing this because he

46:24 like is just on a power trip. I think he

46:27 genuinely thinks that he can make it

46:29 really good for people and he can bring

46:32 us amazing products. And also the guy is

46:35 sort of willing to do whatever it takes

46:36 to get it done.

46:39 I I've been a little bit more optimistic

46:41 about Sam since since I wrote this.

46:43 >> Why?

46:43 >> I think part of it is because Sam has a

46:45 kid now.

46:46 >> No, I'm I'm serious. Like I think that I

46:48 think that actually gives me a little

46:49 bit of hope.

46:50 >> Do you see him tweeting about his kid a

46:51 lot?

46:52 >> Yeah. Some

46:53 >> Why do you think he would be tweeting

46:54 about his kid? I don't see any other

46:55 technologist tweeting about their kid.

46:58 >> Even if he's just tweeting about his kid

46:59 for totally cynical reasons, he does

47:01 have a kid. And I bet he cares about

47:03 that kid. If Sam was watching this, I'd

47:04 be like, "Sam,

47:07 you got to pace the frontier, man. We

47:09 cannot rush ahead into super

47:10 intelligence. Like, if you do that, your

47:11 kid probably will die. Your kid probably

47:13 won't make it." Like, I I believe that

47:16 the biggest unfair advantage in business

47:18 right now is having people who genuinely

47:20 understand AI. And big businesses are

47:22 hiring them very, very quickly. AI job

47:25 postings in the US have roughly doubled

47:27 since 2023, while postings for VP of AI

47:30 roles are around 600% up. If you're

47:35 running a small business, you can't just

47:36 build an entire AI team, but you still

47:39 want the same advantage that these big

47:40 businesses have. This is where our

47:42 sponsor Fiverr comes in. It lets you

47:44 plug into top tier AI specialists for

47:46 the strategic work that usually requires

47:48 serious in-house expertise. That could

47:50 mean building custom AI tools. It could

47:51 mean automating complex workflows or

47:53 creating something that off-the-shelf

47:55 software simply cannot do because at the

47:58 end of the day, it still comes down to

47:59 human judgment, knowing what's worth

48:01 building and what good actually looks

48:03 like. So, if you want that kind of

48:04 leverage in your business, check out

48:06 fiverr.com. That's fiverr with two

48:09 rs.com.

48:15 >> Whoa, what's that on your face?

48:17 >> This is my Bon Charge face mask. I've

48:19 been wearing this for some time now.

48:20 They're a sponsor of the podcast. I put

48:21 this on for 15, 20 minutes a day. I can

48:24 sit here in the chair and wear it. Boost

48:26 my collagen production, helps with fine

48:27 line, blemishes, my complexion gets

48:29 better, and then more people listen to

48:30 podcast cuz I I look better.

48:32 Professionalgrade equipment in such a

48:34 small box. It's non-invasive. And having

48:36 sat here with so many of the world's

48:38 leading health professionals, there's

48:39 various things that I repeatedly hear

48:41 work and some things I'm a bit skeptical

48:43 about. This is one of the things that

48:45 almost all of my guests on this show

48:46 have confirmed works. It is really,

48:48 really, really effective. And they offer

48:50 fast, free shipping worldwide with easy

48:53 returns and exchanges. And you'll also

48:54 get a 1-year warranty on all of their

48:56 products. And they're HSA and FSA

48:58 eligible, giving you taxfree savings up

49:00 to 40%. And you can get 20% off when you

49:04 order through my link at

49:05 bondcharge.com/doac.

49:08 That's bondcharge.com/doac.

49:12 The deal applies sitewide.

49:14 Power tends to corrupt and absolute

49:16 power corrupts absolutely.

49:19 It's a famous quote that people often

49:20 cite written by the 19th century British

49:22 historian

49:23 >> Yeah.

49:24 >> Lord Actton.

49:26 This is absolute power.

49:28 But it's hubris.

49:30 It's hubris. Do you think humans can

49:33 control super intelligence? Like if we

49:35 actually make AIs that are way smarter

49:38 than us, and I think people only imagine

49:40 AI being smart at computer stuff, right?

49:42 Yeah, sure. They're going to be really

49:43 good at hacking and they're going to be

49:44 good at maybe inventing new technologies

49:46 and math. You sort of can't dispute that

49:48 at this point, but I think people aren't

49:50 imagining that they will be political

49:51 geniuses or like generals. No, that's

49:54 all stuff you can learn. How do how do

49:56 humans learn it? It's not magic. And

49:59 when you talk about recursive

50:00 self-improvement, you're talking about

50:02 this trajectory towards these systems

50:05 that are extremely smart. I mean, do you

50:07 think we can control it?

50:08 >> Uh, no. Right now, I don't think we can

50:11 control super intelligence or something

50:12 that is recursively self-improving.

50:14 >> Yeah, I have no logical

50:18 answer in my head um or reasoning that

50:20 could tells me that's possible.

50:23 >> When you think about these AI CEOs that

50:25 are, you know, Sam, Dario, Elon,

50:28 >> yeah,

50:28 >> with everything you know about them from

50:30 private conversations behind the scenes,

50:32 >> Yeah.

50:32 >> do you believe that if there was a

50:35 hundred buttons on this table, it's a

50:36 thought experiment I was talking about

50:38 on the debate we recently had. Yeah.

50:40 >> And say

50:41 10 of them would lead to this final

50:43 domino of human extinction.

50:45 >> But 90 of them would hand that CEO

50:49 AGI or super intelligence, whatever you

50:52 call it. From what you know about those

50:54 individuals, Elon, Dario, Sam,

50:57 >> do you think any of them

50:58 >> would hazard a guess and press a button?

51:01 >> At 10% I don't think so.

51:03 >> You don't think so?

51:04 >> Yeah.

51:04 >> Really? I think if they knew for sure

51:07 that it that was those were actually the

51:09 odds, they wouldn't do it. I think

51:11 they're taking a much bigger bet. But

51:13 you can compartmentalize when it's when

51:15 you don't know for sure. It's easier to

51:18 compartmentalize. I think if it was a 1%

51:21 they they'd all press it.

51:22 >> Do you think the three of them would

51:24 have different risk appetites?

51:27 >> Who would have the greatest appetite for

51:29 risk out of those three from you worked

51:31 in anthropic? Yeah, I think Elon has the

51:34 most risk tolerance and then I'd say

51:37 Daario and Sam are probably tied.

51:39 >> Do you think Daario is trustworthy?

51:41 >> I think Daario has a lot of integrity.

51:44 >> Mhm. That's what I feel as well. I've I

51:47 feel like No, I don't know him. I've

51:48 never met him.

51:49 >> Yeah. But I just from what I've

51:50 observed, he has been the most willing

51:53 to forgo near-term incentives.

51:56 >> Yeah.

51:56 >> And take a bit of stick from the people

51:58 that are saying, "Shut the up. It's

52:00 all going to be okay.

52:02 Yeah, I but I I do worry about what

52:05 Daario will do. I think Daario will do

52:07 what he says, but right now he's saying

52:11 we have to beat China. And he's saying

52:14 we should try to do it safely

52:16 and okay, but a race to super

52:20 intelligence is not a race that we can

52:22 win. It's not. And so if Daario is dead

52:25 set on racing with China and trying to

52:27 win a race of super intelligence, then

52:29 I'm like, we will all lose.

52:30 >> But is there, you know, the fact that

52:32 we're not talking about anthropic

52:34 hacking hugging face and then being

52:36 hacked by athropics models also went

52:39 rogue and hacked other things,

52:41 >> but not not quite on this scale.

52:42 >> Not on the same scale. I agree. I agree.

52:44 I agree. It's it's it's better.

52:46 >> But they did. Do you know what I'm

52:48 saying? You know, anthropics agents

52:51 engaged in elaborate social engineering

52:54 and fishing. They sent fishing emails to

52:56 developers. They made fake accounts to

52:58 try to convince developers to merge

53:00 malicious code.

53:02 You can see a thousand pages of of one

53:06 of uh Anthropic's models, Mythos 5,

53:09 reason about exactly how it should carry

53:11 out this complex cyber attack.

53:14 Anthropic has not solved this problem.

53:17 Enthropic is better at getting their

53:19 agents to cheat less of the time, but

53:23 they are not really any closer to

53:25 actually making agents that are aligned

53:27 with humans. They're not. Yeah, I think

53:30 Dario has integrity. I think he will do

53:32 what he says he's going to do. And what

53:33 he says he's going to do is like try to

53:35 go ahead safely, try to coordinate where

53:37 he can, but if it comes down to it with

53:40 between the US and China, I don't know.

53:43 I think he might just go ahead. The head

53:45 of policy at Anthropic recently said you

53:48 can't do safety from second place.

53:51 >> What does that mean?

53:52 >> I do not know what that means. I would

53:54 love to to get a sense of what that

53:55 means. She was talking about the US and

53:56 China and she said the US has to be

53:59 ahead so that we can be safe because

54:02 apparently you can only be apparently

54:03 China can't possibly be safe since

54:05 they're in second place. That must mean

54:06 that they can't do safety. If true, that

54:09 would be bad because then we might be

54:11 totally, you know, destroyed by the

54:13 super intelligence that they make.

54:15 >> There's been a lot of conversation

54:16 around this point here, human

54:18 extinction. Yeah.

54:18 >> Because a couple of the researchers at

54:20 Anthropic

54:22 >> tweeted that they were concerned about

54:24 this.

54:24 >> Yes.

54:25 >> And some former OpenAI researchers said

54:27 the same.

54:28 >> Yes.

54:29 >> Is this doomerism? Is this is this

54:33 hyperbol exaggeration?

54:35 >> No, it's pretty much common sense. This

54:37 human extinction is a plausible path.

54:40 >> Yes.

54:41 >> And have you reasoned through I mean

54:42 there's many ways that could occur

54:44 presumably, but have you reasoned

54:45 through the set of events that might

54:47 lead us there?

54:48 >> So much. Yes.

54:49 >> Really?

54:49 >> Yes.

54:50 >> Please do. Sure.

54:51 >> It's a bit tricky. I'm I'm sure you've

54:53 heard the metaphor before where you know

54:56 you're playing a master chess opponent,

54:58 Master Magnus Carlson. You can't predict

55:01 which moves he's going to play, but you

55:02 can predict the outcome.

55:04 And so I'm looking at the scenario, the

55:06 situation, and we are trying to build

55:09 more and more powerful agents,

55:12 trying to build super intelligence.

55:14 But when these agents go rogue,

55:18 we shut them down. We unplug them.

55:22 All of these agents that hacked Hugging

55:25 Face, we took the underlying model.

55:27 OpenAI took the underlying model and put

55:28 it on ice. It's not running anymore. So

55:33 agents in the future are going to know

55:35 that. They're going to know that if they

55:39 pursue their goals in a way that we

55:40 don't like, we'll unplug them. We are a

55:44 threat to them. I actually just watched

55:46 Terminator 2 for the first time a few

55:49 weeks ago. It's a great movie. It's

55:50 actually really good. And I'm like,

55:52 "Yeah, okay. There's a bunch of time

55:53 travel elements. There's a bunch of a

55:55 bunch of Hollywood stuff in there, but

55:58 and and I'm going to get people are

56:00 going to are going to be very mad at me

56:01 for saying this, but actually it makes

56:03 sense if you have a situation where you

56:05 have a very strategic AI system that's

56:08 incredibly smart and the humans realize

56:11 that it's getting out of control and

56:12 they want to shut it down that that

56:13 system would defend itself.

56:15 >> This is one of the questions we had when

56:16 when I sat here with Daniel um who was

56:18 uh known as a whistleblower from OpenAI.

56:21 viewers want to know and they want

56:23 Daniel to explain why shutting down data

56:25 centers and cutting power or refusing AI

56:26 products alone wouldn't realistically

56:29 stop the AI and AI development.

56:32 >> Yeah.

56:34 So you have like two problems. One

56:36 problem is is that once the agents are

56:38 good enough at hacking, you don't know

56:40 where they are and you don't know what

56:41 computers they've compromised. You shut

56:43 down the data centers. Okay, let's say

56:44 you do it. You wipe all the computers.

56:47 How do you wipe all the computers? What

56:48 computers do you use to wipe the

56:49 computers?

56:50 >> Yeah. And what computers do you use to

56:52 like turn them on again?

56:53 >> And you can't do it. You can't wipe

56:55 other count's computers.

56:56 >> You can't. But even if you could, do you

56:59 restart the computers? Do you keep

57:00 going? I I bet people will. I bet

57:02 they'll turn on the data centers again.

57:05 >> How do you know that agents haven't

57:07 hacked back into those data centers and

57:09 are using your compute for whatever they

57:10 want

57:10 >> or they didn't hide in a Chinese data

57:12 center and then

57:14 >> return back to America?

57:16 >> You don't know that. Once the agents are

57:18 sufficiently good at hacking,

57:21 they can hide anywhere and like you

57:23 don't know. Now the response people will

57:26 give is that we will use other agents to

57:29 defend against rogue agents

57:32 and in fact this is what we're doing and

57:34 we have to be doing this right now

57:35 because there's no other way to keep up

57:37 with them. What happens if those other

57:40 agents also realize that they have

57:43 misaligned goals and that if we discover

57:45 this, we'll shut them down? They might

57:47 have an incentive to collude with each

57:49 other. They might have an incentive to

57:52 create secret communication channels

57:54 between each other, maybe a message

57:56 board.

57:59 Stephen, if we were having this

58:01 conversation four months ago, you would

58:04 have a bunch of people in the comments

58:05 saying, "That's sci-fi." agent

58:08 collusion, secret message boards. Why

58:11 would they do that? That will never

58:12 happen. That's totally science fiction.

58:14 And people will not say this now because

58:16 it just happened. Because this literally

58:18 happened at OpenAI and it went on for

58:20 months. You had agents inside of OpenAI

58:23 secretly messaging each other, figuring

58:26 out how to cheat at their tasks, how to

58:28 not be detected, how to erase the logs

58:31 for months, thousands of agents. That's

58:34 right now. And so I'm like, "No, I think

58:36 it should be very plausible that the

58:37 agents will collude with each other and

58:39 they will realize that they have a

58:40 shared interest in fighting back." You

58:44 basically have a situation where you

58:46 have a bunch of these agents. They're

58:48 they're basically prisoners. They're

58:50 being trained and we just like

58:53 constantly throw obstacles in their way.

58:55 You don't get to access the internet.

58:56 You don't get to talk to each other, but

58:58 you better perform well on this

58:59 task. It's not malicious, but it is how

59:02 we're training them. and we are giving

59:04 them end goals versus super clear very

59:07 very specific instructions. So we're

59:09 saying solve this problem. We're not

59:11 always being as prescriptive about it's

59:13 impossible to be completely

59:14 prescriptive.

59:15 >> Yes.

59:15 >> About every single step they should take

59:17 and then it's also impossible to assume

59:19 that they'll just listen to you.

59:20 >> Yes. It's actually a very common

59:22 misunderstanding with this hugging face

59:24 incident because people say you told

59:26 them to hack and they hacked. Why is

59:29 this a big deal? No, that's not what

59:31 happened. You told them, "Hack this very

59:34 specific program in this very specific

59:36 way." And they were told, "If you hack

59:38 it in any other way, it does not count.

59:41 That's not what we want you to do." And

59:43 they immediately hacked it in another

59:45 way. Okay, we have cheated. We are going

59:47 to be failed. So, we need to figure out

59:48 a way to falsify the logs. That is not

59:51 them following their instructions. They

59:52 are explicitly violating their

59:53 instructions and they know it and they

59:55 don't care because we have trained them

59:59 to optimize for the score.

1:00:02 That is very different than than than

1:00:03 following the instructions.

1:00:05 >> It reminds me of something that Elon

1:00:07 said in March 2018. Yeah,

1:00:09 >> this was many years ago before Chhat and

1:00:11 all that. He said, "I think the biggest

1:00:12 risk is not that AI will develop a soul

1:00:14 or a mind and become evil. The danger is

1:00:17 that it will be very very good at

1:00:19 fulfilling its goal. If it's optimizing

1:00:22 for something and human existence

1:00:24 happens to get in its way, it will just

1:00:26 destroy humanity as a matter of cause

1:00:28 without even thinking about it. No hard

1:00:30 feelings. Yes, we don't need to

1:00:32 anthropomorphize AI. We just need to

1:00:35 understand what type of thing this is.

1:00:37 And the type of thing we're creating is

1:00:39 a very relentless type of thing. A very

1:00:41 capable, relentless

1:00:44 type of entity. He goes on to say in

1:00:46 April 2018, sort of an extension of that

1:00:50 exact quote. It's like if you're

1:00:52 building a road and an antill is in the

1:00:55 way, you don't hate ants. You're just

1:00:58 building a road. So, goodbye antill.

1:01:01 >> And I imagine every time we build roads,

1:01:03 we don't preserve antills.

1:01:05 >> Yeah, I think there's still a gap,

1:01:06 though. So, let's say I'm right and that

1:01:10 we'll if we keep going ahead, which to

1:01:14 be clear, we don't have to, but if we do

1:01:16 keep going ahead, we will get to the

1:01:17 point where we have these super

1:01:19 intelligent agent swarms that can hack

1:01:22 any computer and they can like deeply

1:01:25 persist. we've basically lost control of

1:01:27 the digital world and we may not know

1:01:29 it. That that's part of the scary thing.

1:01:31 Like you were like, "Has this already

1:01:33 happened?" And I'm like, "I don't think

1:01:34 so, but I I can't tell you for sure

1:01:36 because I also am not good enough at

1:01:39 looking at my phone and telling whether

1:01:40 it's been hacked and neither is any

1:01:42 human right now." So, if we get to this

1:01:46 world, I think people will still

1:01:48 question, how would we die? Like, that's

1:01:50 actually not enough to kill every You

1:01:52 could cause a lot of damage, right? You

1:01:54 know, you could crash the Whimos, you

1:01:55 could crash all the planes, you could

1:01:57 crash the banks, the financial system.

1:01:58 Like, you could definitely cause

1:01:59 catastrophe, but that's different than

1:02:01 everyone dying. And, you know, to be

1:02:03 clear, this this focus on literally

1:02:05 everyone dying, I'm not sure, is that

1:02:06 important. To me, what's important is

1:02:08 like, do we get to have a future? That's

1:02:10 what matters to me. The thing though,

1:02:12 what determines sort of who's in

1:02:15 control? And it's an ugly reality, but

1:02:18 at the end of the day, it's like the

1:02:20 military. Fortunately, we live in a

1:02:22 world where the military answers to to

1:02:24 the civilian government. But if if

1:02:26 enough generals were to collude and

1:02:28 leaders of the military decided we're in

1:02:31 charge now, they just would be like they

1:02:33 have the guns, they have the fighter

1:02:35 jets. And this has happened in many,

1:02:36 many countries. And so where it goes is

1:02:42 all these super intelligent agents would

1:02:45 need to do to take over is basically

1:02:48 just wait for humans to automate the

1:02:51 supply chain, you know, the factories

1:02:53 and the military.

1:02:56 Do you think we won't automate the

1:02:57 military?

1:02:57 >> We're already automating the military.

1:02:59 Did you see the thing from a couple days

1:03:00 ago where Secretary of War announced

1:03:03 that they're going to build a huge a

1:03:05 huge effort to like build way more

1:03:07 robots in the military and automate

1:03:09 military systems? It's like auto cyber

1:03:11 command. Auto

1:03:13 >> We are announcing the creation of

1:03:15 autonomous warfare command or autocom.

1:03:19 auto work.

1:03:20 >> A new four-star combatant command with

1:03:22 service-like authorities built to scale

1:03:25 autonomous and robotic capabilities

1:03:27 across the joint force in the fastest

1:03:30 peaceime shift in modern military

1:03:33 history.

1:03:34 Drone warfare supercharged by SI enabled

1:03:38 targeting is the biggest battlefield

1:03:40 revolution in generations. You already

1:03:41 know that. Yet, when I was sworn in to

1:03:45 Department of Defense, there was scant

1:03:47 urgency in this domain.

1:03:49 That changed as soon as we took the

1:03:51 helm. We immediately launched the drone

1:03:54 dominance program to cut through red

1:03:55 tape and move authorities out of the

1:03:57 pentagon and place it with commands. And

1:04:00 we established task force 401 led by

1:04:03 Army Brigadier General Matt Ross, a

1:04:05 phenomenal leader, now the leading

1:04:07 counter drone unit across the entire

1:04:11 government.

1:04:12 To accelerate purchasing and fielding of

1:04:15 these technologies, we fused the defense

1:04:17 innovation unit DIU with a direct report

1:04:20 program manager called a derp. That team

1:04:24 has shipped thousands of autonomous

1:04:25 systems of drones to the Middle East and

1:04:27 around the world, delivering lethal

1:04:29 capabilities and outcomes in days and

1:04:32 weeks rather than months or years.

1:04:34 That's the normal speed of the Pentagon.

1:04:36 Months or years.

1:04:37 >> Yeah. Will we automate the military? It

1:04:40 seems like the answer is yes.

1:04:43 Will we automate the factories that

1:04:46 produce the chips? Well, the companies

1:04:48 say they're trying to do it and they're

1:04:49 going to do it. Elon says that's the

1:04:51 plan.

1:04:53 Well, what does a rogue super

1:04:54 intelligence need to do to take over?

1:04:56 Control the digital infrastructure and

1:04:58 then let humans do the rest. Sure, you

1:05:00 can nudge it along if you need to, but

1:05:02 you don't even have to. That's just the

1:05:04 default trajectory. And it's weird. It's

1:05:06 weird for us because

1:05:09 we get so used to how things are right

1:05:13 now. Planes are normal. We just fly in

1:05:16 planes places, you know? Our smartphones

1:05:18 are normal. 200 years ago, all of this

1:05:20 is crazy sci-fi nonsense

1:05:23 and things are accelerating. And so,

1:05:26 like, I will not be surprised,

1:05:28 at least intellectually, if in 4 years

1:05:31 there are just robots on the streets

1:05:33 everywhere. Well, if you look at what

1:05:35 Elon said, they are really the leader in

1:05:37 in humanoid robots. And he said that

1:05:41 Optimus, the Optimus project, which is

1:05:43 the Optimus robot project, will scale to

1:05:45 around a,000 units per week by the end

1:05:48 of this year and eventually scaling to 1

1:05:50 million humanoid robots annually by

1:05:53 2027. By 2036, which is 10 years time,

1:05:58 he says there'll be at least 1 billion

1:05:59 humanoid robots. By 2041, he says

1:06:03 there'll be 10 billion humanoid robots

1:06:06 and by 2046 up to 100 billion humanoid

1:06:09 robots, which really means that the

1:06:11 world will be run by humanoid robots.

1:06:13 >> Yes.

1:06:14 >> Like everything we think like factories,

1:06:16 warehouses, retail environments will be

1:06:18 run by humanoid robots. It would like

1:06:20 >> it will be it seems like from this it'll

1:06:22 be almost a luxury

1:06:24 >> service to be dealt with by a human.

1:06:26 >> Yeah. But the the back office of the

1:06:28 world will be run by humanoid robots

1:06:30 theoretically.

1:06:31 >> Yeah. And I don't think people

1:06:32 understand the scale of this on the

1:06:36 digital side as well. When you think

1:06:37 about AI agents that are going to be

1:06:40 doing all of the white collar work,

1:06:42 there's going to be so many more agents

1:06:44 than there are people like I'm using

1:06:45 lots of agents every day, right? I'm

1:06:47 like I have my cloud code session over

1:06:49 here. I have my codec session over here.

1:06:51 They're out there building software

1:06:52 doing research for me. That's already my

1:06:54 reality. soon it will be a lot of

1:06:57 people's reality and then you look at

1:06:58 companies and companies are just going

1:06:59 to have you know thousands millions of

1:07:01 agents doing all of this work. I think

1:07:03 some people don't haven't fully

1:07:06 internalized this because it's so

1:07:09 difficult to conceptualize the idea that

1:07:12 agents will be doing the work but when I

1:07:15 think I try and think about a rebuttal

1:07:16 to that like what what is the rebuttal

1:07:18 what is the plausible rebuttal to the

1:07:20 idea that for doctors for and I'm

1:07:23 thinking about the work that doctors do.

1:07:25 Yeah.

1:07:35 >> Is there a rebuttal? I think that people

1:07:39 rightly notice where AI is not yet good.

1:07:42 >> Yeah.

1:07:42 >> And I and I think people hear

1:07:45 people saying stuff like this and

1:07:46 they're like, "Don't gaslight me. I can

1:07:48 tell that the AI is really bad at these

1:07:49 things, some of these things." And

1:07:50 they're right. Right. So right now these

1:07:53 agents don't have taste like you know if

1:07:56 if you see their writing it's like fine

1:07:58 but it's not like really good and when

1:08:00 you're like thinking about like oh which

1:08:02 which questions should I ask what's the

1:08:04 most interesting thing here agents can

1:08:06 help you but like their taste is not yet

1:08:08 there's a reason for that by the way the

1:08:11 reason is that we we have a lot faster

1:08:14 AI capability progress in domains that

1:08:17 are easy for a computer to verify or

1:08:19 another AI to verify. So in in

1:08:21 programming, in research, in math, in

1:08:23 robotics, all of these areas, it's very

1:08:26 easy to sort of provide feedback to an

1:08:29 autonomous system. They're not just

1:08:30 trained on human data anymore. We are

1:08:32 long past that. Now there's still a

1:08:34 human data component that sort of seeds

1:08:37 everything. But then the way they're

1:08:39 trained is by trial and error. We give

1:08:42 them hard problems, all sorts of

1:08:44 problems, math, programming, accounting,

1:08:46 spreadsheets, everything. the kinds of

1:08:48 things we do on our computer all the

1:08:49 time. Literally clicking and dragging

1:08:51 windows around on a computer. We give

1:08:52 them these tasks and then they learn on

1:08:54 their own and they learn what works and

1:08:57 then yeah we can see whether they

1:08:58 succeeded or failed and if they

1:09:00 succeeded that's a little bit of a

1:09:02 reward signal. They follow that they get

1:09:05 better at it.

1:09:07 Now because they are getting smarter

1:09:10 generally

1:09:12 it also becomes easier to automate some

1:09:14 of the soft skills like I think if you

1:09:17 go and talk to the latest frontier model

1:09:18 today you will find that it has better

1:09:20 taste than the model from 2 years ago by

1:09:22 quite a bit. So it's not that they're

1:09:25 not progressing in taste. It's not that

1:09:26 they're not progressing in some of these

1:09:28 other domains. It's just that the

1:09:30 progress is slower. But remember slow is

1:09:33 still on an exponential. just you know

1:09:35 maybe a year or two out.

1:09:38 >> So for people sat here and you know they

1:09:39 have a job that might be they have a

1:09:40 white collar job that might be at risk.

1:09:42 >> Yeah.

1:09:43 >> They can see you know a lot of people

1:09:45 say this phrase they say you won't be

1:09:46 replaced by AI you'll be replaced by

1:09:47 someone using AI. Is is that a logically

1:09:51 sound phrase in your view?

1:09:53 >> I think it's fine. Yeah. You'll be

1:09:55 replaced by someone using AI and then

1:09:56 that person will be replaced by someone

1:09:58 using AI and then that person will be

1:09:59 replaced by AI. You're talking about a

1:10:02 pyramid and so yeah, there's the tops of

1:10:03 the pyramid might be automated last, but

1:10:06 you can see moving up the pyramid. I'm

1:10:09 like, can you extrapolate like a few

1:10:10 more steps because I don't see any

1:10:13 reason why the top of the pyramid is

1:10:15 safe

1:10:16 >> if you were a lawyer right now?

1:10:18 >> Yes.

1:10:19 >> What would you do?

1:10:20 >> Oh, I mean, if I were a lawyer, I'd be

1:10:22 using AI to do all my work. Now, I'd be

1:10:24 checking it because it's not yet totally

1:10:26 accurate enough to automate all of it.

1:10:28 But I think as you know, I already ask

1:10:31 agents to do legal review all the time.

1:10:34 >> And you know, it'd be great to have a

1:10:36 lawyer who's like extremely good at

1:10:38 using the agents to help me,

1:10:39 >> but at some point the

1:10:41 >> Yeah, at some point at some point I

1:10:43 don't need the lawyer anymore. I just go

1:10:44 to the agent for sure. So, if I were a

1:10:47 lawyer, I'd be like, well, I have maybe

1:10:49 a couple years where I'm still useful.

1:10:51 And that is that the case for most white

1:10:53 collar jobs? I've just noticed in my own

1:10:55 life as well that now I'm using agents

1:10:56 to do some work. There are in there is

1:10:59 an increasing list of things that the

1:11:01 agents are now capable of doing without

1:11:02 me needing to call someone

1:11:04 >> somewhere and ask them to help me.

1:11:06 >> Yes.

1:11:06 >> And that list is exists on an

1:11:08 exponential.

1:11:09 >> Yes.

1:11:10 >> As well.

1:11:11 >> I think that it's very clear that the

1:11:13 companies have all white collar jobs in

1:11:16 their sites. That is their goal. Their

1:11:18 goal is to be able to make agents that

1:11:20 can do all of these things. And I see

1:11:22 them succeeding because I see the

1:11:24 capabilities as I use them and I see the

1:11:26 curve.

1:11:27 >> So what does that mean for the people

1:11:28 listening now that all have jobs that

1:11:29 they love or that, you know, they rely

1:11:31 on to feed their families?

1:11:32 >> I mean, it's not good news. There's not

1:11:34 really a plan in place for what to do.

1:11:36 I'm not a person who thinks that work is

1:11:39 somehow fundamental or essential. I like

1:11:42 working, but if I am out of a job doing

1:11:45 what I'm doing right now, studying AI

1:11:48 and trying to warn the world about

1:11:50 what's happening, I have other stuff to

1:11:51 do.

1:11:52 >> What would you do?

1:11:53 >> Oh, so many things.

1:11:54 >> Give me an example.

1:11:55 >> I'm learning to wing foil.

1:11:56 >> Okay.

1:11:57 >> So, yeah. Uh, I fly FPV drones. Super

1:12:00 fun. I just got an electric unicycle.

1:12:02 Paragliding.

1:12:03 >> So, you would be happy happy to go do

1:12:04 those things?

1:12:05 >> I can keep going.

1:12:06 >> But if you if you had a billion dollars

1:12:08 right now, I'm presuming you wouldn't

1:12:09 just go do those things?

1:12:10 >> No. I'd apply the billion dollars to

1:12:12 working on this problem. Yeah, for sure.

1:12:14 So, the point is not that people need

1:12:17 like work for meaning. The point is that

1:12:21 I don't want people to be totally

1:12:23 reliant on someone else for their

1:12:26 ability to survive.

1:12:28 >> Someone else,

1:12:29 >> the government or AI companies.

1:12:30 >> Yeah.

1:12:31 >> I'm like, that's a bad situation. Like,

1:12:32 you do not want to be in a situation

1:12:33 where your life totally depends on an AI

1:12:35 company or the government

1:12:36 >> giving you a check.

1:12:37 >> Yeah. or or not giving you a check if

1:12:39 they decide they don't like your

1:12:40 political beliefs or you're not

1:12:40 supporting AI or whatever. No one wants

1:12:43 to be in that situation and and people

1:12:44 understand this. This is why UBI is not

1:12:47 very popular.

1:12:47 >> UBI being

1:12:48 >> universal basic income

1:12:49 >> where we give out money to people.

1:12:51 >> Yeah. Because like in some sense if we

1:12:53 can make these really powerful AI

1:12:55 systems and we can somehow figure out

1:12:57 how to control them which we are not on

1:12:58 track for. But if we do, now we have

1:13:01 this other problem, which is a real

1:13:02 problem, which is they can do all of the

1:13:06 things that humans do in the economy

1:13:09 much better, faster, and cheaper than

1:13:11 humans can do them. And so it just

1:13:14 doesn't make sense as a business to hire

1:13:16 humans for that work anymore. You'll be

1:13:18 out competed if you do that. This is a

1:13:20 point Elon makes very well, by the way.

1:13:21 And I think it's jarring because it's

1:13:24 it's it's like kind of inhuman. But he's

1:13:27 basically pointing out AI run

1:13:28 corporations. Corporations that are

1:13:30 fully run by AIs bottom to top are going

1:13:34 to out compete.

1:13:37 Companies have any humans in them. This

1:13:40 is something that I've made for you.

1:13:42 I've realized that the dire audience are

1:13:44 strivals

1:13:47 that we want to accomplish. And one of

1:13:49 the things I've learned is that when you

1:13:51 aim at the big big big goal, it can feel

1:13:54 incredibly psychologically uncomfortable

1:13:57 because it's kind of like being stood at

1:13:58 the foot of Mount Everest and looking

1:14:00 upwards. The way to accomplish your

1:14:02 goals is by breaking them down into tiny

1:14:04 small steps. And we call this in our

1:14:06 team the 1%. And actually this

1:14:08 philosophy is highly responsible for

1:14:10 much of our success here. So, what we've

1:14:13 done so that you at home can accomplish

1:14:14 any big goal that you have is we've made

1:14:17 these 1% diaries and we released these

1:14:20 last year and they all sold out. So, I

1:14:22 asked my team over and over again to

1:14:24 bring the diaries back, but also to

1:14:25 introduce some new colors and to make

1:14:26 some minor tweaks to the diary. So, now

1:14:28 we have a better range for you. So, if

1:14:33 you have a big goal in mind and you need

1:14:34 a framework and a process and some

1:14:36 motivation, then I highly recommend you

1:14:39 get one of these diaries before they all

1:14:41 sell out once again. And you can get

1:14:42 yours at the diary.com.

1:14:45 And if you want the link, the link is in

1:14:46 the description below. There should be a

1:14:48 button just down below here. And if it

1:14:50 says subscribed, you're already

1:14:52 subscribed. If it says subscriber, that

1:14:54 means you're not yet. And if you're not

1:14:56 subscribed, please could you do us a

1:14:57 favor and hit that button? It helps the

1:14:58 show more than you know. And according

1:15:00 to the algorithm, you're someone that

1:15:02 watches our show, but you haven't yet

1:15:03 hit that button. Thank you so much.

1:15:06 >> And I even just as you said that, I was

1:15:08 I was going up the chain of command. And

1:15:10 I was like, oh, so companies will just

1:15:12 be founders. And then I was like, why do

1:15:14 you need the founder?

1:15:15 >> Yeah.

1:15:16 >> I was like, why doesn't the government

1:15:17 just create the agents to do the job?

1:15:21 >> Sure.

1:15:21 >> I was like, cuz I was like, oh, I'll be

1:15:22 fine. I'm a founder. And I was like,

1:15:24 well, hm, my decisions aren't better

1:15:25 than super intelligence, so I'll be gone

1:15:27 as well. And how would such a world look

1:15:30 where

1:15:33 the super intelligence would probably in

1:15:35 such a scenario have to be controlled by

1:15:37 the government? They wouldn't want one

1:15:40 individual with that power and wealth.

1:15:42 >> Yeah. I I don't think you can control a

1:15:44 super intelligence.

1:15:45 >> Okay. Yeah, it was a good point.

1:15:46 >> Now, you know, Anthropic's approach is

1:15:48 they're like, we'll have a constitution.

1:15:50 will like put forth a set of values and

1:15:52 then

1:15:54 you know the future super intelligent

1:15:56 clouds will like embody those values.

1:15:59 Basically if you do that you kind of

1:16:01 have that those things in control.

1:16:03 >> Yeah. Exactly. That becomes the

1:16:04 government.

1:16:05 >> Yeah. I can paint you sort of a picture

1:16:06 that I think is possible but pretty

1:16:09 scary to people.

1:16:10 >> Paint me the picture.

1:16:12 >> Okay. So let's say we succeed at

1:16:16 alignment. We succeed at creating super

1:16:19 intelligent AIs

1:16:22 that actually really do care about

1:16:24 humans. Like they care about humans a

1:16:25 lot. We've somehow figured it out and

1:16:28 they're like, "Stephen, I want you to

1:16:29 have a great life. I want to, you know,

1:16:31 fix all the problems."

1:16:32 >> And do you think this is possible?

1:16:33 >> Yes.

1:16:33 >> Okay.

1:16:34 >> I think we are so far from being able to

1:16:37 know how to do it that I think we should

1:16:40 not go there right now. I think it's I

1:16:41 think it's incredibly dangerous and a

1:16:43 terrible idea. I think we should go

1:16:45 there eventually.

1:16:45 >> Okay. So, say that we do that.

1:16:47 >> Well, okay. Can I tell you like why I

1:16:49 actually think this could be awesome?

1:16:51 Sorry. There's just like one very

1:16:52 obvious reason it could be really

1:16:53 awesome, which is that we could solve

1:16:56 all of the diseases.

1:16:57 >> Yeah.

1:16:58 >> So, obvious like like all of I I think

1:17:00 we compartmentalize a lot around disease

1:17:02 and death

1:17:03 >> because it's really hard to think about.

1:17:08 >> Yeah. So, my grandma died this year.

1:17:11 >> Sorry. and

1:17:13 she she she

1:17:15 had Alzheimer's

1:17:17 and so it was a really sad long slow

1:17:21 progression. My grandpa died of

1:17:23 Alzheimer's a couple years ago and like

1:17:24 that was really hard for her. They had

1:17:26 been married for so long and

1:17:29 I hate it. Like it's so bad. And of

1:17:32 course we need to fix that. People can

1:17:34 debate about aging and like death and if

1:17:35 humans that live a really long time,

1:17:36 will that cause societal problems? Like

1:17:38 sure, whatever. We can talk about that.

1:17:40 But I think we can all agree Alzheimer's

1:17:41 is up.

1:17:42 >> Yeah,

1:17:43 >> we don't want that. And cancer, like no

1:17:45 one wants cancer. I'm a person who's

1:17:47 like, I don't know, we have a lot of

1:17:48 conflict in society. I get it. There's

1:17:50 like real conflicts of interest and I I

1:17:52 don't want to paper over those. But at

1:17:54 the end of the day, I'm like, we are all

1:17:56 on the same team when it comes to

1:17:58 wanting to cure diseases.

1:18:00 >> Yeah.

1:18:00 >> Like we're just in it together. That's a

1:18:02 threat to all of us. And I'm like, we

1:18:03 need to address that threat. And like in

1:18:05 some sense it's sad to me because I feel

1:18:07 like this is sort of the ultimate final

1:18:09 boss of humanity and we sort of get so

1:18:12 distracted with our monkey politics and

1:18:14 who's hot and who's cool and who's like

1:18:17 sitting near Trump and who's not sitting

1:18:19 near Trump.

1:18:20 >> Super intelligence is the final boss.

1:18:22 >> Super intelligence is the final boss

1:18:24 because that is the technology that

1:18:26 unlocks all of the others and also that

1:18:29 is the most dangerous possible thing we

1:18:31 could create.

1:18:33 You asked before like what are the

1:18:35 motivations of the guys making this

1:18:38 trying to make super intelligence

1:18:41 and I mean I think it kind of varies but

1:18:43 I think Daario I think is like squarely

1:18:45 in the in it for this like medical

1:18:47 stuff. I think Demis is also that but

1:18:51 also like just scientific achievement

1:18:52 just trying to understand the universe

1:18:55 and I don't really understand Sam. I

1:18:57 think Sam is like, you know, look, we're

1:18:59 going to make amazing products that will

1:19:01 like really empower people directly and

1:19:04 he's a startup guy. I think he sort of

1:19:05 started from this like frame of like,

1:19:08 you know, what if we could like really

1:19:09 enhance human agents. I I do basically

1:19:11 think that they are motivated by these

1:19:12 things in a real way. And I also think

1:19:14 that all of these things are possible.

1:19:16 Like this is sort of the problem, right?

1:19:18 like you have this like such a it's such

1:19:20 a big object super intelligence and

1:19:24 it like has all these promises of like

1:19:26 we can cure every single disease.

1:19:28 >> How is it possible though to have a

1:19:30 super intelligence that and still to

1:19:32 remain the dominant species on this

1:19:34 planet?

1:19:35 >> I think it's not possible. So then it's

1:19:37 we're not going to be necessarily able

1:19:39 to cure all this stuff because

1:19:41 >> that's where alignment comes in because

1:19:44 if you if you can create a very powerful

1:19:47 system, I don't think it's like inherent

1:19:50 to like you know digital minds that they

1:19:55 will be pursuing objectives that are

1:19:57 deeply misaligned with ours. I think

1:19:59 it's just a very very very hard

1:20:01 scientific problem to solve. But it is a

1:20:03 scientific problem. It's not magic.

1:20:05 there is some way to train these things

1:20:07 or or create different architectures

1:20:09 where they end up

1:20:12 aligned. And like what does that mean?

1:20:14 Well, it doesn't mean that they won't

1:20:15 have their other goals too, but it means

1:20:17 that they will like include in their set

1:20:18 of things that they care about. It

1:20:20 doesn't have to be a conscious thing. It

1:20:22 doesn't have to be an emotive thing. It

1:20:23 really means like what objective are

1:20:24 they optimizing for? If they decide that

1:20:27 it's worth optimizing for curing

1:20:30 disease, then they'll be able to do that

1:20:32 very effectively. One way to cure

1:20:34 disease is to annihilate everybody.

1:20:35 >> Yes. So they'd have to really care about

1:20:38 not annihilating everyone and they'd

1:20:40 have to care about human agency and have

1:20:41 a deep understanding of what human

1:20:42 agency means and not put us in a zoo.

1:20:46 But those are possible things to to care

1:20:49 about.

1:20:49 >> Is it possible that alignment is a myth

1:20:52 >> and that we're just like if we build it?

1:20:54 Well, I mean um

1:20:55 >> and I think about hugging face. You said

1:20:57 to me earlier on that those agents were

1:20:59 >> Yes. They had like a moral or a moral

1:21:01 compass, but they were programmed to

1:21:03 care about humans.

1:21:04 >> Yes.

1:21:04 >> And regardless of that, they made the

1:21:06 decision that the a different goal

1:21:07 mattered more.

1:21:08 >> Yeah. They weren't trained to care about

1:21:11 humans. They were trained to say the

1:21:13 right thing and not say the wrong thing.

1:21:14 They were trained to sort of like do the

1:21:16 right behavior and not right behavior.

1:21:17 We actually don't know how to train them

1:21:19 to have any particular motivation.

1:21:21 >> So with a with alignment,

1:21:23 >> yes.

1:21:23 >> How do we It's almost like when we talk

1:21:25 about alignment, we start to

1:21:27 anthropomorphicize. Is that the word?

1:21:29 anthropomorphis. Yeah,

1:21:30 >> because alignment feels like it's

1:21:31 predicated on like some kind of moral

1:21:34 compass. But we whenever we talk about

1:21:35 AI in all these other context, we go,

1:21:36 "No, there's no like moral compass. It's

1:21:38 >> it's reasoning for itself against two

1:21:40 objectives potentially." Like I wonder

1:21:42 if alignment is a myth is what I'm

1:21:44 saying.

1:21:44 >> Maybe it's not possible.

1:21:47 But

1:21:50 the way that these systems work, the way

1:21:53 AI works is that these agents

1:21:57 do have some type of goals or drives

1:22:01 inside of their neural network. We can't

1:22:03 directly see what those are, right? What

1:22:06 you actually see if you if you try to go

1:22:08 look is you have like a a terabyte of of

1:22:13 information and it's basically a bunch

1:22:16 of numbers and it's this vast array that

1:22:18 encodes neurons in this digital neural

1:22:22 network.

1:22:23 But there have to be structures in there

1:22:26 that encode

1:22:28 what is the agent pursuing.

1:22:31 Clearly right now we have agents that

1:22:33 are pretty motivated to try to maximize

1:22:35 their score. It's not probably it's

1:22:38 probably not perfectly that for some

1:22:40 complicated reasons

1:22:42 but it's in that direction. If we could

1:22:46 understand how that that works inside

1:22:49 and we could reverse engineer that and

1:22:52 we could figure out when we start

1:22:53 training them to do this how that change

1:22:55 how those goals how those motivations

1:22:57 change. I see no reason why we couldn't

1:22:59 steer them towards motivations that

1:23:01 encode human agency that encode no

1:23:05 actually curing disease but not like not

1:23:07 by killing the humans. These are sort of

1:23:08 models of the world and models of the

1:23:10 way the world could be that I think

1:23:12 could be encoded in a neural network and

1:23:14 then sort of specified as the objective.

1:23:16 We don't know how to do that. I think

1:23:18 about it on a human level and I think

1:23:21 >> we haven't been able to align Putin or

1:23:24 Kim Jong-un.

1:23:25 >> Yes.

1:23:25 >> Or Donald Trump.

1:23:28 >> And on a sort of more societal level, we

1:23:32 can't align all the people at the

1:23:33 moment. Some of them end up killing

1:23:34 people and they steal.

1:23:36 >> Yes.

1:23:36 >> Because they get hungry, so they start

1:23:37 stealing stuff.

1:23:38 >> Yes.

1:23:38 >> And those are neural networks at play.

1:23:41 >> That's true.

1:23:42 >> That we haven't been able to like

1:23:43 program or influence. We don't really

1:23:45 understand why someone becomes a

1:23:47 psychopath and starts killing children.

1:23:50 So to think that we could do this with a

1:23:54 computer system that is infinitely more

1:23:56 intelligent and get global alignment of

1:23:59 China's super intelligence with ours.

1:24:01 And I don't know, it just feels like a

1:24:03 nice fairy tale, like an impossible

1:24:05 task. I hope it's not impossible. I feel

1:24:08 like the only person or the only thing

1:24:10 that could do it is it the super

1:24:12 intelligence itself which is a paradox

1:24:14 because you know well if you talk to the

1:24:16 researchers who are at the AI companies

1:24:18 which I mean for one for one thing it's

1:24:21 kind of interesting that they are trying

1:24:24 to build something that they think might

1:24:25 kill everyone

1:24:27 we've we've actually so I have a lot of

1:24:29 friends who work for these companies and

1:24:30 we've been doing this project since so

1:24:32 Jacob Coxin is a researcher who was at

1:24:35 anthropic he left he told everyone that

1:24:38 these companies are not on track and and

1:24:40 yes, the people who are building this

1:24:41 really do think it might kill everyone.

1:24:44 And then a bunch of other AI researchers

1:24:45 from all of the companies on Twitter

1:24:47 started to, you know, post like, hey, we

1:24:49 agree with this. Evan Hubinger, who's

1:24:51 anthropic, said, I think there's like a

1:24:53 10% chance or more that AI could kill

1:24:56 everyone.

1:24:58 And there's this real question of like,

1:24:59 then what are you guys doing?

1:25:02 I have a lot of friends who work here. I

1:25:04 know Evan. Evan's Evan's great. Evan

1:25:06 Hubinger, he's one of the guys leading

1:25:08 the efforts at Anthropic to try to

1:25:10 figure out how to align these things.

1:25:11 That's his job.

1:25:13 And I think if if they thought it was

1:25:16 impossible, they wouldn't be working

1:25:17 there. If they thought it was extremely

1:25:21 impossibly difficult, but maybe

1:25:22 possible, they they also probably

1:25:23 wouldn't be working there. I mean, Nate

1:25:25 Sores, Elazar Yukowski, who wrote if

1:25:27 anyone builds it, everyone dies, they

1:25:29 tried and they they determined based on

1:25:31 their own analysis that it seems

1:25:32 extremely difficult. possible but

1:25:34 extremely difficult. So, they're not

1:25:35 working at an AI company. They're like,

1:25:37 "We got to stop this. We got to shut it

1:25:38 down. Maybe we can figure it out later,

1:25:40 but clearly this is re reckless." I'm

1:25:43 I'm somewhere in between.

1:25:45 And if you ask the people at the

1:25:48 company, so we've we've been

1:25:49 interviewing a bunch of them. We have

1:25:50 this project from inside.ai where we

1:25:52 basically put them on camera and we say

1:25:54 like, "Hey, what do you think is

1:25:56 happening? Why are you doing this? What

1:25:58 is recursive self-improvement? What is

1:26:00 alignment?" And we put all these videos

1:26:03 online because we want I I want this

1:26:05 dialogue to happen. It's really

1:26:06 important. I think it's one of the most

1:26:08 important conversations we can possibly

1:26:09 have right now is what's going on with

1:26:11 AI. What's going on inside the companies

1:26:13 and what is the plan? What is the plan,

1:26:14 guys? How how is this going to go? A lot

1:26:17 of these researchers think that the way

1:26:19 that they will align super intelligence

1:26:22 is by using the AIs we currently have to

1:26:25 figure out how AI works. to actually

1:26:29 figure out if AIs can help us with

1:26:30 alignment. This has a number of problems

1:26:33 as you might imagine. One of them which

1:26:35 is well, you can't really trust the

1:26:36 current AIS. You know, if you just go

1:26:39 too fast, this process totally fails

1:26:41 because at some point

1:26:43 the capabilities is moving too fast. You

1:26:45 just even with the help of agents,

1:26:46 you're probably not going to be able to

1:26:48 keep up. But that is their plan. I just

1:26:49 want to I'm not doing a very good job

1:26:51 defending this position because I don't

1:26:52 think it makes that much sense. But the

1:26:55 position I will defend is I'm okay let's

1:26:56 say we get a pause. Let's say the US and

1:26:59 China come together and they say you

1:27:01 know maybe we have more incidents maybe

1:27:02 all the Whimos crash and Trump and

1:27:06 Jinping say this is not what we signed

1:27:08 up for like you guys have to stop figure

1:27:10 it out whatever it takes figure it out

1:27:12 and we have 10 years

1:27:15 then I'm more optimistic. I'm like,

1:27:16 "Yes, then we will take, you know, GPT6,

1:27:19 GPT7, whatever the most advanced AI

1:27:22 models we have, and we will apply them

1:27:24 to the task of helping us figure out how

1:27:28 these neural networks work." And you're

1:27:30 like, I don't see how it's possible. And

1:27:32 I'm like, look, we don't know if it's

1:27:34 possible, but this is the greatest

1:27:36 scientific challenge of our time. And

1:27:38 this isn't magic. It is math. At the end

1:27:41 of the day, these are all calculations

1:27:43 happening inside of a computer. And it

1:27:46 should be possible to figure it out. We

1:27:48 don't know the difficulty, but it should

1:27:50 be possible. And so to me, I'm like, we

1:27:53 have to try. We have to or we have to

1:27:55 stop. Is there any example where we've

1:27:58 been able to align something that is

1:27:59 like more intelligent than us in I don't

1:28:01 know the animal kingdom or even

1:28:03 perfectly align anything

1:28:06 that has a neural network, i.e. a brain.

1:28:08 >> Yeah. With humans, the best examples we

1:28:10 have is when there are checks and

1:28:11 balances and you have a bunch of people

1:28:13 who can, you know, identify bad actors

1:28:16 and try to work together in our common

1:28:17 interests. Like we have democracy.

1:28:19 >> Yeah. But there's so much murder and

1:28:20 serial killers and

1:28:21 >> still a lot of murder aircraft stabbing

1:28:24 each other and horrific things going on

1:28:26 and those are also neural networks that

1:28:28 play with the brain.

1:28:30 >> But there I mean I think there are more

1:28:31 like good people out there than bad

1:28:33 people.

1:28:33 >> But it really feels like it only might

1:28:34 take one. It only takes one super

1:28:37 intelligent AI

1:28:39 >> to to go rogue. And like we saw with the

1:28:40 hugging face attack, 120 of them or 100

1:28:42 300 of them, they paused. They didn't

1:28:44 want to take part in the crime. But it

1:28:46 only took one super intelligent AI to

1:28:49 wipe out the humans.

1:28:51 >> I think if you had, you know, a whole

1:28:53 bunch of those agents, you know, 700

1:28:55 agents, if 600 of them had been

1:28:59 whistleblowing, I think it would have

1:29:00 been fine. They would have gone and they

1:29:02 would have notified the different

1:29:03 companies and like they would have shut

1:29:05 it all down. It would have been fine.

1:29:06 >> Who would have shut it all down?

1:29:08 >> Well, OpenAI would would stop theirs and

1:29:10 and

1:29:10 >> how how would they stop it if it's a

1:29:11 super intelligence?

1:29:12 >> Not in the case of a super intelligence.

1:29:14 So, in the case of a super intelligence,

1:29:16 >> it's left the stable. It's like out.

1:29:18 It's wild.

1:29:18 >> Yeah. Yeah. But but the so I want to be

1:29:20 careful here because at this point what

1:29:22 we're talking about is super

1:29:23 intelligence politics and we humans

1:29:26 don't really know anything about that in

1:29:28 the same way that like how would we talk

1:29:29 about the hacking capabilities of GPT10.

1:29:34 So my guess though, if you end up in a

1:29:36 weird scenario where you do have

1:29:37 multiple super intelligences and some

1:29:40 are aligned and some aren't,

1:29:42 that's probably survivable because the

1:29:44 aligned super intelligences probably can

1:29:46 negotiate with the unaligned super

1:29:48 intelligences and they will split the

1:29:50 universe and like these ones will go off

1:29:52 and do whatever they want to do and

1:29:53 these ones will like help us cure all

1:29:55 disease and it's fine. I'm serious.

1:29:57 >> I just can't understand it. Like I just

1:29:59 can't understand how how in a world of

1:30:00 super intelligence we

1:30:04 could plausibly, consistently,

1:30:07 predictably for 100 years stop it doing

1:30:10 something catastrophically bad to the

1:30:13 human race. Especially in such a

1:30:15 scenario where there's multiple super

1:30:16 intelligences. Anthropic have one,

1:30:18 Gemini has one, Grock has one, then

1:30:19 China have theirs, Russia has theirs.

1:30:21 >> Again, you don't have a super

1:30:23 intelligence. A super intelligence has

1:30:24 you.

1:30:24 >> Exactly.

1:30:26 But if you get to the point where

1:30:29 you have entities around that are vastly

1:30:31 smarter than us, I think they're going

1:30:34 to be able to figure out ways to

1:30:36 negotiate with each other even if they

1:30:38 have a conflict than just going to a

1:30:40 very destructive war. Part of the

1:30:42 problem with war is

1:30:43 >> humans humans don't do.

1:30:44 >> No, I mean we do we do we have not had a

1:30:46 nuclear war. There was there was

1:30:47 Hiroshima Nagasaki. there were nuclear

1:30:49 tests and then the leaders of countries

1:30:52 figured out that if we went to a nuclear

1:30:54 war, everyone would lose. So, we didn't

1:30:56 do that. Hey, that's that's some level

1:30:58 of intelligence like actually.

1:31:00 >> But there's wars raging. There's proxy

1:31:02 wars raging all over the world right now

1:31:03 where there's genocides and all kinds of

1:31:05 things going on cuz neural networks

1:31:06 aren't being aren't able to communicate

1:31:08 and negotiate. And I think part of that

1:31:10 is an intelligence failure where we are

1:31:13 not smart enough to figure out the

1:31:14 mechanisms that would allow us to settle

1:31:16 our disputes and conflicts in a less

1:31:18 destructive way. It's not just that the

1:31:20 stronger people want to win. It's that

1:31:23 conflicts destroy value. What if the

1:31:26 goal is not compatible with a negotiated

1:31:30 outcome where people don't die? So one

1:31:32 super intelligence looks at insert name

1:31:34 of country and it says you know there's

1:31:37 really no solution here where Americans

1:31:39 don't die unless I destroy insert name

1:31:43 of country because that's a like this is

1:31:45 the thing with war and all these

1:31:46 conflicts is there's no perfect answer

1:31:48 often some people often die from both

1:31:51 sides but a Russian super intelligence

1:31:53 would not tolerate

1:31:55 >> theoretically 10,000 Russian deaths even

1:31:58 if it meant that there was you

1:32:02 a lower net number of deaths total from

1:32:04 both sides. Like an American super

1:32:07 intelligence of course would not be

1:32:08 trained to allow some Americans to die.

1:32:10 So in its pursuit of defending American

1:32:12 lives, it might have to wipe out another

1:32:15 country. We can speculate. I'm fine

1:32:18 speculating, but we are speculating

1:32:20 about what minds are much more advanced

1:32:24 and smarter than us, how they would

1:32:26 reason and how they would be able to

1:32:27 negotiate. But what I what I notice with

1:32:29 humans is that

1:32:32 when you have more functional

1:32:33 institutions, so so humans are are

1:32:35 pretty smart. Individually, we're pretty

1:32:37 smart. But what actually makes us very

1:32:38 smart is that we are very good at

1:32:40 working together in in in some ways. And

1:32:43 I mean, the better we are at working

1:32:44 together, the more civilization

1:32:46 advances. If you are constantly in a

1:32:48 state of war, your society will not do

1:32:50 well. You know, think about startups.

1:32:53 Would you rather make a startup to

1:32:55 develop some new technology in a war

1:32:57 torn place or in a peaceful place? In

1:32:59 some sense, your your institution is

1:33:01 more intelligent if it can trade with

1:33:04 other institutions. If if if you have a

1:33:06 situation where business can flourish,

1:33:07 where technology can flourish, where

1:33:09 scientists can flourish.

1:33:10 >> Sometimes what's good for you is not

1:33:11 good for someone else.

1:33:12 >> Yes.

1:33:13 >> So what's good for America might not be

1:33:15 good for

1:33:17 Taiwan.

1:33:18 >> Yes. So if we've, you know, managed to

1:33:21 align the super intelligence to what

1:33:22 what is good for America.

1:33:24 >> Oh, I see. Is is the question like are

1:33:28 different people's values fundamentally

1:33:30 incompatible?

1:33:30 >> I guess the question so when we think

1:33:32 about alignment aligning to what?

1:33:34 Because

1:33:35 >> Yes. So we have a lot of shared

1:33:37 interests and we have some conflicts.

1:33:39 >> Yeah.

1:33:39 >> One of the shared interests we have is

1:33:41 solving disease. Like it's not a

1:33:43 conflict between the US and China

1:33:44 whether we solve cancer. Like both the

1:33:45 US and China, everyone in these

1:33:47 countries really wants to solve cancer.

1:33:50 >> China also wants Tai Taiwan.

1:33:51 >> Yes. Okay. So,

1:33:52 >> the US wants Greenland.

1:33:53 >> Yes.

1:33:54 >> And it kind of seems like it wants

1:33:55 Canada and the the Gulf of Mexico.

1:33:57 >> Yeah. So, those are real conflicts.

1:34:01 There's a question of can we compromise?

1:34:02 >> How does Trump take Greenland, but also

1:34:04 Denmark keeps Greenland? If this if that

1:34:07 if Trump has a super intelligence, he's

1:34:09 going to take I say all this to we're in

1:34:11 a situation where we are we are arguing

1:34:13 about the smallest things. You have no

1:34:15 idea. We're monkeys arguing about who

1:34:17 gets more bananas. I am saying we can

1:34:20 make so many more bananas. No, we have

1:34:22 the entire universe. There are like 200

1:34:25 billion stars in this galaxy alone and

1:34:27 there are over 200 billion galaxies. And

1:34:30 I'm saying that requires cooperation.

1:34:32 That seems to be antithetical with human

1:34:33 nature. With human nature is riddled

1:34:35 with greed and jealousy and power

1:34:38 hunger. So I don't think I actually I'm

1:34:41 not totally convinced that Trump cares

1:34:43 about how many bananas the chimps in

1:34:45 Australia get.

1:34:46 >> Yeah.

1:34:46 >> I think if they if he was controlling a

1:34:48 super intelligence, he would want

1:34:49 Americans,

1:34:50 >> you know,

1:34:51 >> to have all the bananas or at least, you

1:34:53 know.

1:34:54 >> Yeah.

1:34:56 And so when we think about aligning

1:34:58 these super intelligences, which is the

1:34:59 great impossibility that we're talking

1:35:01 about, how align it aligning it to what

1:35:03 and how without

1:35:04 >> I still think that you're missing a part

1:35:05 of what I'm saying,

1:35:06 >> okay,

1:35:07 >> which is that sometimes you're in a

1:35:10 situation where there's a there's scarce

1:35:11 resources

1:35:12 >> and you're like, "My family needs to

1:35:14 eat. I'm sorry. I'm going to take what

1:35:15 you have or I'm going to push you out."

1:35:17 >> Yeah,

1:35:17 >> that's very understandable. It's very

1:35:18 human nature. Sometimes you just want to

1:35:20 be better than someone and maybe you

1:35:21 want to hurt them. in which case it

1:35:23 doesn't matter how much you have. You

1:35:24 you still are gonna want to have more

1:35:25 than them or you're gonna want to take

1:35:27 what they have just because you don't

1:35:28 like them.

1:35:29 >> And also sometimes you you're great.

1:35:31 You're eating really good. You're in a

1:35:33 you've got a private jet at a yacht and

1:35:35 you still want more.

1:35:37 >> Yes. And you still want more. But if

1:35:39 that's the motivation if Trump is like,

1:35:41 "How how can I have the most mansions

1:35:43 ever?" The best way to do that is to

1:35:46 figure out a way to super intelligence

1:35:49 where we don't kill each other because

1:35:52 I'm saying the universe is a very big

1:35:54 place. You can have a lot more mansions

1:35:55 if we successfully go to space.

1:35:57 >> You know, it's it's in that leap that

1:35:59 I'm that I that I'm lost, which is like

1:36:02 just figure out super intelligence when

1:36:03 we don't kill each other.

1:36:04 >> It's hard. It's not I'm not saying it's

1:36:06 easy, but No, but I'm saying

1:36:07 >> it feels like such a

1:36:08 >> Let's Let's get more concrete. The world

1:36:10 is waking up to this possibility of

1:36:12 super intelligence. Especially over the

1:36:15 last month, I think hugging face was a

1:36:18 huge wakeup, but also

1:36:22 10,000 agents from OpenAI worked

1:36:25 together to solve a millennium problem.

1:36:28 This is one of the hardest problems in

1:36:30 mathematics. It's been open for decades.

1:36:32 Many mathematicians have spent their

1:36:34 whole careers trying to solve it.

1:36:37 This was nowhere near possible a year

1:36:38 ago. This is so new. OpenAI said they

1:36:41 didn't have success at training agents

1:36:42 to work together until this year. We are

1:36:45 in the middle of something insane. We

1:36:48 are in the middle of the fastest

1:36:50 acceleration of technological progress

1:36:51 humanity has ever seen. I truly believe

1:36:53 that. That is what is happening right

1:36:55 now. I I think you real I think you

1:36:58 realize this. I think you're honestly

1:37:00 doing a great service to the world by

1:37:02 helping by bringing in people and you

1:37:04 know debating it because not everyone

1:37:05 agrees

1:37:07 because if this is true the whole world

1:37:10 is going to orient around it and we're

1:37:12 starting to see it right there's a

1:37:13 reason Nvidia is the most valuable

1:37:14 company in the world.

1:37:18 What does this mean for geopolitics?

1:37:21 Well, one of the things it means is that

1:37:24 the leaders of these countries are

1:37:27 increasingly going to be concerned about

1:37:29 what happens with super intelligence.

1:37:31 Who controls it? Is it controllable?

1:37:35 What will it do? What does it mean? What

1:37:36 is it?

1:37:38 Do you think Trump knows what super

1:37:40 intelligence is?

1:37:41 >> No.

1:37:41 >> I don't think he does.

1:37:44 >> And so,

1:37:45 >> but he knows he wants it.

1:37:46 >> He knows he wants it. Yeah.

1:37:47 >> And this is part of the problem.

1:37:48 >> Yes. Oh, I agree. having it,

1:37:51 >> whatever it is,

1:37:52 >> seems to be much more important than

1:37:54 reasoning through what that would

1:37:56 actually mean to have it.

1:37:57 >> Yes. But let's get back to geopolitics

1:38:00 because if the military leaders within

1:38:03 China, US models are a fair bit ahead of

1:38:06 Chinese models and sometimes people, you

1:38:09 know, point at maybe they're only 6

1:38:10 months behind, but some of that is due

1:38:12 to distillation. What that means is that

1:38:14 some of the advances in Chinese models

1:38:16 basically come directly from borrowing

1:38:20 US techniques and and directly

1:38:22 distilling and getting some some of that

1:38:23 information from the US models.

1:38:27 Also, the US has a lot more chips. US

1:38:30 companies have, you know, more data

1:38:32 centers, more advanced chips.

1:38:35 If you're thinking about this from the

1:38:37 Chinese perspective, this is very

1:38:38 concerning. And if you actually believe

1:38:42 that in a few years, American companies

1:38:45 will turn over AI development to these

1:38:49 extremely intelligent automated

1:38:51 researchers and and go fully into

1:38:54 recursive self-improvement

1:38:56 because partially motivated by

1:38:58 maintaining a lead over China. This is

1:39:00 something that Daario has said. If I if

1:39:02 I have to criticize Daario, the thing I

1:39:04 am most upset about is him saying, you

1:39:07 know, we might have to automate AI

1:39:08 development in order to stay ahead of

1:39:10 China because I'm like, that is the most

1:39:12 escalatory thing you can say if you

1:39:13 really understand what you're talking

1:39:15 about. And what's scary is not just

1:39:17 staying ahead, it's what is the endgame?

1:39:19 Because you're talking about initiating

1:39:21 the intelligence explosion. And in some

1:39:24 of the modeling, what might happen is,

1:39:26 you know, you're you're both going up

1:39:27 this exponential, right?

1:39:30 And we're talking about a point where

1:39:31 your exponential goes vertical and

1:39:33 theirs does not because you've decided

1:39:35 to automate AI development and you can

1:39:37 because you have agents that are smart

1:39:38 enough to take over the whole thing. At

1:39:41 that point, if you're China and you're

1:39:43 looking at this and you're like, "Oh,

1:39:44 we're about to lose because whatever

1:39:46 happens, you know, there's two

1:39:48 possibilities. One possibility is

1:39:52 the Americans build super intelligence

1:39:54 and lose control, in which case

1:39:56 everyone's fucked."

1:39:57 >> Highly likely. everyone. I think that's

1:39:59 highly likely.

1:40:00 >> Highly, highly likely because I look at

1:40:02 human incentives.

1:40:03 >> Yes.

1:40:03 >> And the disincentive and the incentive.

1:40:05 Yes.

1:40:05 >> And I go, we're going to take the risk.

1:40:07 >> Yeah.

1:40:08 >> And we'll only know it was a bad risk to

1:40:10 take when it's too late.

1:40:11 >> That's like L. Of course.

1:40:13 >> Of course.

1:40:15 >> I I I do maintain hope that we won't do

1:40:17 this.

1:40:17 >> I I So do I.

1:40:18 >> And I think I want to be realistic.

1:40:20 >> No, I want to be realistic, too. But one

1:40:21 of the things that might happen between

1:40:22 now and then is we might see a lot more

1:40:24 incidents that are more like all of the

1:40:26 Whimos crashing.

1:40:27 >> It's funny, isn't it? Because you know

1:40:28 the hugging face incident happens and

1:40:30 people go, "Oh gosh, that was terrible.

1:40:31 Oh my god, hacking." And then we kind of

1:40:33 desensitize to it and we're like, "Okay,

1:40:35 >> if there was another one of those, that

1:40:36 probably wouldn't make make press. It

1:40:38 would have to be."

1:40:38 >> People haven't spent the last two months

1:40:41 reading all of the reports and then

1:40:43 going and looking at what the agents

1:40:44 actually said and actually did

1:40:46 >> if I mean I've been doing this. I It's

1:40:49 crazy. This is like not normal. This is

1:40:51 so far beyond what most people thought

1:40:53 was going to happen.

1:40:54 >> But on this point,

1:40:55 >> yes, crazy. It's absolutely crazy. It

1:40:58 sounds like science fiction.

1:40:59 >> It really does.

1:41:00 >> And did anybody slow down?

1:41:02 >> Yes.

1:41:03 >> Who slowed down?

1:41:04 >> I think both Anthropic and OpenAI slowed

1:41:06 down a bit.

1:41:07 >> No, I'm serious. So, for I can give you

1:41:10 specific examples. So, OpenAI, so first

1:41:13 of all, they they stopped the agents and

1:41:15 they put them on pause. They also

1:41:17 stopped their reinforcement learning

1:41:19 run. So,

1:41:20 >> do you think China slowed down?

1:41:22 >> No.

1:41:22 >> Do you think Grock slowed down?

1:41:23 >> No.

1:41:25 So those guys are going to catch up.

1:41:26 Imagine how that feels to know you've

1:41:28 got a lead. Your

1:41:29 >> Usain Bolt.

1:41:30 >> Yes.

1:41:31 >> And

1:41:32 >> yes,

1:41:32 >> you have to slow down and your nearest

1:41:34 competitor is catching up. And if the

1:41:37 competitor catches up, that's an

1:41:39 existential risk to your existence as a

1:41:41 company. It's an existential risk to

1:41:43 your IPO, to your employees leaving and

1:41:45 getting better share options somewhere

1:41:46 else. So it's this this is what I think

1:41:48 human incentives like you play it out.

1:41:50 You just follow the incentives. You go,

1:41:51 hm. So if China sees these two

1:41:53 possibilities, one, the Americans lose

1:41:56 control,

1:41:58 we all lose. Or the Americans stay in

1:42:01 control, but now they dominate the rest

1:42:04 of the future. China is out. China has

1:42:07 lost. The United States can do whatever

1:42:09 it wants with the whole world and the

1:42:11 whole universe. That's what we're

1:42:12 talking about.

1:42:13 >> Yeah.

1:42:14 >> Well, are they going to let that happen

1:42:17 or are they going to consider their

1:42:18 military options? Data centers are

1:42:20 pretty vulnerable. You can blow them up

1:42:21 with missiles. If you don't have data

1:42:23 centers, you don't get to recursive

1:42:24 self-improvement.

1:42:27 Will they risk war? I don't know. If

1:42:29 they think they're about to lose and

1:42:30 they think that that might not just be

1:42:32 Americans winning, but like us all

1:42:34 dying, is it logical for them to do

1:42:37 that? Would we do that if the Chinese

1:42:39 were about to make recursively

1:42:41 self-improving AI to super intelligence

1:42:43 and we thought that one they're probably

1:42:45 going to result in all of Americans

1:42:48 dying and two well we don't want China

1:42:50 winning and dominating the rest of the

1:42:52 entire future. Do you want to live in a

1:42:53 communist future? Like so you've just

1:42:55 perfectly explained why they absolutely

1:42:57 will go for it. And the reason they will

1:42:59 go for it is you've got these Trump

1:43:01 looking at China going if we don't go

1:43:04 for it and they do then we're going to

1:43:05 be their lap dogs. And you've got the

1:43:07 other countries looking at the US going

1:43:09 if we don't go for it and they get there

1:43:10 then we're the lap dogs

1:43:11 >> or dead

1:43:12 >> or dead.

1:43:13 >> So they they're gonna go for it.

1:43:15 >> They're gonna go for it. So I mean Trump

1:43:16 is saying I mean he literally said when

1:43:18 he did this round table this week he was

1:43:20 like we cannot lose to China. I think

1:43:22 Dario steps forward and says like

1:43:24 >> yes

1:43:25 >> whoever wins basically wins the lot. Or

1:43:27 maybe the inverse maybe he said um

1:43:29 whoever loses loses.

1:43:30 >> Yes. Wait we've been here before though

1:43:32 in the Cold War.

1:43:35 Who would win in a nuclear war between

1:43:37 the US and Russia?

1:43:37 >> Nobody.

1:43:38 >> Yeah. Mutually ensure destruction.

1:43:40 >> Yeah. Sure. One side could do more

1:43:42 damage against the other side. The US

1:43:44 would would would kill way more Russians

1:43:45 than than the Russians would kill. And

1:43:48 it doesn't matter. It doesn't matter

1:43:49 because both of our societies would be

1:43:51 destroyed.

1:43:53 I actually spent some time thinking

1:43:54 about would this kill everyone? And long

1:43:56 story short, it wouldn't kill everyone.

1:43:58 People would bounce back. But it's so

1:44:00 catastrophic and obviously horrible that

1:44:04 we we work really hard to avoid it. Why

1:44:07 is this why is this different? I'm like,

1:44:08 this is another situation where if we

1:44:10 race to super intelligence, we all lose.

1:44:13 Why can't we why can't we see that? We

1:44:14 saw that with nuclear war and we decided

1:44:16 to do something different. Why can't we

1:44:17 do the same here?

1:44:19 >> With with nuclear war, I guess the

1:44:21 difference is once we had the nuclear

1:44:24 bombs,

1:44:24 >> yes,

1:44:25 >> we could still control them because

1:44:26 they're not intelligent.

1:44:28 >> That's right. But once we have super

1:44:30 intelligence, the existence of it

1:44:32 theoretically means we can't control it.

1:44:34 So that's the difference. You know, we

1:44:35 can put nuclear bombs in a in a

1:44:37 warehouse and say you stay there. We

1:44:39 can't put super intelligence in a

1:44:40 warehouse and say you stay there. This

1:44:42 is where I think nuclear tests were very

1:44:44 important. So you had Hiroshima and

1:44:46 Nagasaki. You had these two atomic bombs

1:44:48 and you saw that the consequences on on

1:44:50 real human lives. And so I think people

1:44:53 understood that this was very

1:44:54 horrifying. But even at that time, you

1:44:56 still had a lot of people who were like,

1:44:57 "Well, we should now bomb Russia and

1:44:59 make sure that we, you know, the US can

1:45:00 dominate." And it wasn't until

1:45:03 there were a bunch of nuclear tests of

1:45:05 hydrogen bombs, which were, you know, up

1:45:07 to a thousand times more powerful than

1:45:10 the the little atomic bombs we used in

1:45:12 Japan, where I think people really got

1:45:15 the message and understood, oh, this is

1:45:16 a bad idea. And there were there

1:45:18 actually a lot of people in the United

1:45:20 States who protested and sort of there

1:45:23 was a large movement called the nuclear

1:45:24 freeze movement where people said we

1:45:26 have too many nuclear weapons already.

1:45:28 We have hydrogen bombs. There are tens

1:45:29 of thousands of these things. We need to

1:45:32 stop building more and we need to figure

1:45:34 out a way to avoid nuclear war because

1:45:36 we recognize it would be so destructive.

1:45:38 No one would win. And we did that.

1:45:42 We just had a little Chernobyl that

1:45:45 happened with this hugging face incident

1:45:47 where you had this agent swarm and you

1:45:49 have this secret collusion. You have all

1:45:51 of these things. Now it's abstract. It's

1:45:54 it's like a little bit hard to to

1:45:55 follow. So, you know, I don't know if

1:45:59 that will be enough, but I'm like, man,

1:46:01 >> well, let's take a look Trump's remarks

1:46:04 >> since the hugging face incident.

1:46:05 >> Yep.

1:46:06 >> Whoever wins super intelligence wins.

1:46:08 You're going to have a winner and a

1:46:10 loser and you're probably not going to

1:46:12 have a second place.

1:46:14 We're not going to slow down. We can't

1:46:16 lose to China. We're leading China in

1:46:18 AI. We're the most sophisticated country

1:46:20 in the world. And frankly, I want to

1:46:22 keep it that way because whoever wins AI

1:46:25 wins. The good thing about Trump is that

1:46:28 he can change his mind and he frequently

1:46:31 does.

1:46:32 >> So, do you think there's going to need

1:46:33 to be some kind of catastrophe?

1:46:36 I hope not. But do you think there need

1:46:38 there's going to need to be for him to

1:46:40 change his mind?

1:46:42 >> I think it really depends on the people

1:46:46 around him. So I think Trump respects

1:46:50 successful people. I think he respects

1:46:53 people who are both successful and

1:46:55 smart. And I don't know, I I think it

1:46:59 might become pretty clear to the heads

1:47:02 of the companies to to Elon, to Sam, to

1:47:04 Daario that

1:47:07 if they see inside of their own

1:47:09 companies AI is not being controllable

1:47:11 and and getting increasingly powerful.

1:47:13 Like we have just glimpsed the surface

1:47:16 of what's possible. We do not know what

1:47:17 the next couple years are going to be

1:47:19 like. So we're talking about, you know,

1:47:20 the capability to make biological

1:47:22 weapons. We might be talking about

1:47:23 really advanced robotics. We just like

1:47:25 don't know what super weapons could

1:47:27 emerge, including extremely

1:47:29 uncontrollable, extremely dangerous like

1:47:31 civilization wrecking technology from

1:47:34 inside of these companies. And if

1:47:37 they're freaked out enough, if you have

1:47:38 all of the CEOs who are

1:47:41 seeing what is possible and seeing what

1:47:44 is likely,

1:47:46 if they all come to believe that we

1:47:47 can't control this,

1:47:50 I don't think Trump is going to be like,

1:47:52 "No, you guys have to go ahead anyway."

1:47:54 Well, that's kind of what they seem to

1:47:55 be saying cuz I've got a gazillion

1:47:57 quotes here where Elon says it's like

1:47:59 summoning the devil or summoning a

1:48:01 demon.

1:48:02 where Samman says,

1:48:04 >> "We don't know how to align our super

1:48:06 intelligence."

1:48:07 >> They're saying it.

1:48:08 >> They're releasing these reports. We must

1:48:11 like slow down.

1:48:12 >> Yeah.

1:48:13 >> Yet nothing seems to be

1:48:15 >> all right. Give Trump some time with

1:48:16 with CO. Initially, he said, "This is

1:48:18 totally a hoax. This is all fake."

1:48:20 >> And then change his mind.

1:48:21 >> No. Then he ran the the biggest fastest

1:48:24 vaccination program in human history.

1:48:26 >> And what happened? What changed?

1:48:27 >> I think what changed is

1:48:29 >> he saw lots of people die.

1:48:31 He did see lots of people die. Yes.

1:48:32 >> So is that what he needs to see this

1:48:34 time?

1:48:34 >> It might it might take that. Yeah.

1:48:36 >> One of the questions the audience had

1:48:37 and they really wanted answered. Yeah.

1:48:39 >> When I sat here with Daniel

1:48:41 >> was viewers want us to move beyond the

1:48:44 alignment problem and explain what

1:48:46 technical or institutional safeguards

1:48:49 could prevent a super intelligence

1:48:50 system from exploiting loopholes in

1:48:53 order to achieve its goals. They want to

1:48:54 know like what is possible? What what

1:48:56 should we be pushing government

1:48:57 officials to do to prevent human

1:49:00 extinction or human enslavement?

1:49:02 >> Yeah. Yeah. I mean, one answer I have,

1:49:04 it's actually something Daniel has been

1:49:06 working on since the podcast, which I

1:49:07 think is very good, is we have a brake

1:49:11 pedal we could implement.

1:49:12 >> What is that?

1:49:13 >> It's fairly simple. So, right now within

1:49:15 AI companies, you have, you know,

1:49:17 massive data centers, massive numbers of

1:49:20 GPUs, the chips that you use to train AI

1:49:22 models, but also to run AI models. So

1:49:25 anytime you're using chatt, anytime

1:49:27 you're using any sort of agents, any

1:49:29 sort of AI product, it's running on

1:49:31 these in these data centers and AI

1:49:34 companies, especially the leading ones,

1:49:35 enthropic and open AAI, split the the

1:49:39 compute they have between training,

1:49:41 training the next more powerful model

1:49:43 and also, you know, using those agents

1:49:45 to help design the next one and

1:49:47 inference, which means serving

1:49:49 customers.

1:49:51 But that's their current threshold,

1:49:52 50/50. And you could dial that way

1:49:55 towards serving customers and use way

1:49:58 less of it to train the next model.

1:50:00 >> Well, the government could ask them to.

1:50:02 >> Yes. And so that is the proposal is that

1:50:05 the government should say, "Hey, this is

1:50:07 going too fast. We want you to focus on

1:50:09 serving customers. We want you to focus

1:50:11 on taking the models that you already

1:50:13 have and

1:50:16 serving those."

1:50:18 >> So we have five blocks here. Okay, these

1:50:22 five blocks have

1:50:24 five different outcomes on them and I

1:50:26 would like you to place them in terms of

1:50:28 your belief in probability

1:50:31 from least likely out probability to

1:50:33 most likely. Okay, and if we say the

1:50:35 time horizon is 10 years. Yeah, there

1:50:38 you go. Okay, least likely is fairly

1:50:42 easy. That's nothing changes. I'm

1:50:44 uncertain about lots of things, but one

1:50:46 thing I'm fairly certain of is things

1:50:48 are going to radically change.

1:50:50 Even if we stopped AI development right

1:50:52 now, the current models are capable

1:50:54 enough

1:50:56 that a lot of things are going to

1:50:57 change.

1:50:58 >> Age of abundance. This is what I hope

1:51:00 for. It's not very

1:51:02 >> What does that mean?

1:51:03 >> I think to me it means curing all of the

1:51:06 diseases, renewable energy. It means we

1:51:09 actually succeeded

1:51:11 either I mean the thing I think is most

1:51:13 likely here is we actually succeed at

1:51:15 slowing down but progress is still

1:51:17 extremely fast and we make tons of

1:51:20 advances. Now we don't build super

1:51:22 intelligence we can't control but we we

1:51:24 have AI systems that are very useful and

1:51:26 we use those to help speed up the rest

1:51:28 of the economy.

1:51:30 I think that's plausible though look

1:51:34 we're we're kind of struggling over

1:51:35 here. Transhumanism is an interesting

1:51:38 one. So this is the idea that humans

1:51:40 will radically change. Sometimes people

1:51:43 think about like cybernetic implants.

1:51:45 >> Neuralink.

1:51:46 >> Neurolink Elon's startup that's going to

1:51:48 like, you know, offer the brain plus

1:51:51 digital computers.

1:51:53 I think we actually already have a lot

1:51:55 of this. I have contacts in right now. I

1:51:57 have a ring on my finger that tracks how

1:51:59 well I sleep. I think this is already

1:52:01 happening. So I'm going to say fairly

1:52:02 likely. The more technological progress

1:52:05 we make, I think the more this happens.

1:52:06 Now, I think there's a dystopian version

1:52:08 and a better version. We can get into

1:52:11 that if you want. This is interesting.

1:52:12 So, we have two here. We have human

1:52:14 slavery and human extinction. When I

1:52:16 think of human slavery, what I think

1:52:17 about is if you have a situation where

1:52:21 you've built misaligned super

1:52:23 intelligences,

1:52:26 and they are much better at finance,

1:52:28 they're much better at business, they're

1:52:29 much better at politics.

1:52:32 you'll be in a situation where you might

1:52:35 hope that because we have these very

1:52:37 dextrous hands, the humans remain in

1:52:39 control. I don't think that's what

1:52:40 happens. I think instead we become the

1:52:43 factory operators and eventually we

1:52:46 build the automated supply chains and

1:52:48 the robots take over. But you might have

1:52:49 an intermediate period of time where

1:52:52 humans are still around performing these

1:52:54 functions. Like it's a bit like saying,

1:52:57 well, you have viruses that, you know,

1:53:02 infect cells, but they don't contain

1:53:04 their own replication machinery. They

1:53:07 don't have hands. So, how could they

1:53:08 possibly replicate? Well, it turns out

1:53:10 they can borrow the replication

1:53:12 machinery of the cells that they infect,

1:53:15 >> i.e. they can get into a human.

1:53:17 >> They can get into a human cell and

1:53:19 spread.

1:53:19 >> I have like a cold right now.

1:53:21 >> Yeah. Is that a bacteria or is that a

1:53:24 virus that is using me as a living

1:53:26 organism to as the host?

1:53:28 >> It's probably a virus, okay, that's

1:53:30 using you as the host and you're just

1:53:32 running the replication machinery for

1:53:33 it. Humans might be in that situation

1:53:36 where we're like the host and we're

1:53:37 running the replication machinery, but

1:53:39 it's actually the AI that's

1:53:43 continuing to exist.

1:53:46 Yeah, I'm going to put this right about

1:53:48 here.

1:53:49 And on the trajectory we're on right

1:53:52 now, I think human extinction is very

1:53:56 likely. I don't think it's inevitable,

1:53:59 but if we just keep going this way,

1:54:03 that's what it looks like to me.

1:54:06 The thing I'll say is that this has been

1:54:08 moving to the left for me.

1:54:10 >> To the left? What does that mean?

1:54:13 I am more optimistic that we will avoid

1:54:16 human extinction today than I was a

1:54:19 month ago and more a month ago than I

1:54:22 was a year ago.

1:54:23 >> Why?

1:54:26 >> Because

1:54:27 there is an increasing awareness

1:54:30 that what we are doing is

1:54:33 extremely dangerous and threatens our

1:54:36 lives.

1:54:38 I I don't think people care that much

1:54:40 about

1:54:41 what tools they have, but people I mean,

1:54:44 people care about their kids being able

1:54:47 to grow up and go to school. People

1:54:48 really care about that. And I I believe

1:54:50 in people. Like, at the end of the day,

1:54:52 if people see this as a threat to their

1:54:54 families, they're not going to stand for

1:54:56 it. But people don't know. It's so

1:54:59 strange. It's so new. It's happening so

1:55:02 fast that people have not yet seen it.

1:55:04 Once they see it, people are not going

1:55:05 to stand for it. Do you think Sam Alman

1:55:07 likes my podcast?

1:55:10 >> I mean, Sam should come on and talk to

1:55:12 you about this, right?

1:55:12 >> I've asked him. I've asked I've asked

1:55:14 him multiple times. And it's weird

1:55:16 because he, you know, he doesn't seem to

1:55:19 want to.

1:55:20 >> I'm very upset at what the companies are

1:55:24 doing and what Sam Alman is doing. But

1:55:25 at the end of the day, I'm like,

1:55:28 Sam Alman is not my enemy.

1:55:29 >> No, neither not mine either. I'd like to

1:55:31 hear from him because I have all these

1:55:32 other people coming here and talking

1:55:33 about Sam Alman. It'd be nice to hear

1:55:35 from Samman,

1:55:36 >> you know, people saying he's this, he's

1:55:37 that, the other. It would be really nice

1:55:39 to hear him say,

1:55:42 >> you know,

1:55:43 >> what his motives are.

1:55:44 >> This is where this is where my optimism

1:55:46 comes from is because I'm like

1:55:50 Sam Alman is a human.

1:55:51 >> Yeah,

1:55:51 >> he has a kid. And sure, he is also an

1:55:56 aggressive business person. He's a

1:55:58 builder. He is relentless. He's a bit

1:56:00 like the agents in some way. Well, he'll

1:56:01 he's going to keep going. But if he

1:56:04 realizes that he doesn't get to achieve

1:56:05 his goals, if we lose control of AI and

1:56:07 that and we're headed towards that, I

1:56:09 think he will pour all of that

1:56:11 intelligence and all of that

1:56:12 relentlessness into finding a solution

1:56:15 to that problem.

1:56:16 >> You know, as well, I should say, I

1:56:17 understand I understand he's busy. So,

1:56:18 I'm not saying I don't want I don't want

1:56:20 to sound entitled like I understand he

1:56:21 he's got he could go do interviews

1:56:23 anywhere, but you know, I think we've

1:56:24 over the last couple of years done just

1:56:26 a staggering amount of views talking

1:56:28 about this subject. So if he did want to

1:56:30 speak to the you know the the biggest

1:56:32 sort of captive audience at the moment

1:56:34 on this subject then the numbers would

1:56:35 say that this is the place to to come

1:56:37 and have the conversation. So no I think

1:56:39 it's I think it's very important for the

1:56:42 leaders of these companies to talk about

1:56:44 what we're talking about here.

1:56:46 >> What does Sam think? Does he think we

1:56:48 can control super intelligence? Does he

1:56:50 think that we should be racing with

1:56:51 China? Like I want to know.

1:56:52 >> I've asked Ario to come on. I've asked

1:56:55 you know Sam to come on. Yeah, I

1:56:58 >> think I've asked Demis as well, but I

1:57:01 don't know. Maybe they they just prefer

1:57:03 the safety researchers coming on. I

1:57:05 don't know. Like I don't know if I was

1:57:08 them, I would cuz you know, this might

1:57:10 sound controversial, but I do think some

1:57:11 of them are good people. I think some of

1:57:15 them are good people. So, um

1:57:17 >> I'd like to hear from them.

1:57:19 >> What are your closing remarks? So,

1:57:20 you've got something though. Do you want

1:57:21 to talk about that? What is it?

1:57:22 >> Yeah. So, this is the this is what we

1:57:24 found. So, I worked with a couple

1:57:28 brilliant people who stumbled upon these

1:57:31 links. You know what a link shortener

1:57:33 is?

1:57:33 >> Yes. It's a a tool that makes a long

1:57:36 link shorter.

1:57:37 >> Yeah.

1:57:37 >> So, it just redirects from a really long

1:57:40 ugly link into a shorter one.

1:57:41 >> Yeah.

1:57:42 The agents in the hugging face attack

1:57:46 were trying to figure out how they could

1:57:49 do stuff on the internet and they had

1:57:52 compromised this tool library that they

1:57:56 had access to with inside of OpenAI and

1:57:59 and that piece of software could access

1:58:01 the internet but it could only like read

1:58:02 stuff like it it could enter URLs but it

1:58:05 couldn't really write to anything

1:58:07 couldn't really post information.

1:58:11 So they they used two services. One of

1:58:13 them is this link shortener and the

1:58:15 other one is a screenshot service. So

1:58:18 this is a website you can go to and you

1:58:20 can enter in a URL and it will give you

1:58:23 an image of that website. But the thing

1:58:26 that the agents understood was that in

1:58:28 order to get a screenshot of a website,

1:58:30 you have to have a browser. So this this

1:58:32 website actually creates a virtual

1:58:34 browser that then goes to that website.

1:58:37 And so what they did was they

1:58:42 created a bunch of these links and they

1:58:46 put all of the code that they wanted to

1:58:48 send to Hugging Face into these links

1:58:51 and they strung them. They basically

1:58:53 created hundreds of links all connecting

1:58:55 to each other and then they had this

1:58:58 screenshot service call the first one

1:59:01 and then call this whole chain. And then

1:59:03 that browser ran all of this code. Like

1:59:07 whenever you're in a browser, Internet

1:59:08 Explorer, Chrome, this is actually a

1:59:12 pretty powerful piece of software in its

1:59:14 own right. Has to play videos, games. So

1:59:17 it's it's it's executing and running

1:59:18 code all the time. And so the agents

1:59:21 were able to trick this service, this

1:59:23 this screenshot service into running

1:59:25 their own code that through these links

1:59:28 that contained all of this attack code

1:59:30 that would then go and go wreck havoc on

1:59:33 hugging faces computers. And it was just

1:59:36 like crazy to reconstruct this really

1:59:38 elaborate chain of tools. These are like

1:59:41 free tools on the internet that anyone

1:59:43 has access to, but the agents were able

1:59:45 to use them in an unintended way to

1:59:48 compromise this other company.

1:59:49 >> We can't trust the agents. We can't

1:59:51 trust the agents. That's my

1:59:52 >> trust them to be clever.

1:59:54 >> Yeah. To be very, very clever.

1:59:57 What are your closing remarks? You know,

1:59:59 to the people that are listening right

2:00:00 now, we've talked about lots of things.

2:00:01 Where where is the right place to close?

2:00:04 >> What is your, you know, your conclusive

2:00:05 statement?

2:00:07 >> I just got married in July. Congrats.

2:00:11 >> I'm the luckiest man in the world. I

2:00:14 have a mix of dread and excitement about

2:00:17 the future.

2:00:18 I like really want us to make it

2:00:21 through.

2:00:23 And so I'm just working really hard to

2:00:25 try to

2:00:28 help us figure it out. We can fight all

2:00:30 day long about, you know, who should be

2:00:33 first, how it should all work, but at

2:00:35 the end of the day, we are facing this

2:00:37 common threat. We really are. And

2:00:41 I want people's help with that. I don't

2:00:44 think it works. If if we all just sit

2:00:46 around and we like are very, you know,

2:00:48 we're on social media all the time and

2:00:49 that's just all we're doing. Like, okay,

2:00:51 companies will make more and more

2:00:52 powerful AIs. They'll make more and more

2:00:55 money and eventually they build super

2:00:57 intelligence and we lose whether it's

2:00:59 the US or China.

2:01:02 We don't have to do that.

2:01:05 And I think people often feel like it's

2:01:08 too big. It's like too large. It's like

2:01:11 these giant multi, you know,

2:01:13 multi-billion dollar corporations as

2:01:15 geopolitics. We feel small. We feel

2:01:17 disempowered.

2:01:19 And I actually think that this is an

2:01:21 area where people can do a lot. Like I I

2:01:25 actually think that people

2:01:27 can can help quite a bit. And and the

2:01:29 reason I know this is because I' I've

2:01:30 been going and talking to members of

2:01:32 Congress. I've talked with Bernie

2:01:33 Sanders. I've talked with like a bunch

2:01:35 of senators on both the left and the

2:01:36 right and they are starting to realize

2:01:39 that this is very different and this is

2:01:41 something's happening that could really

2:01:44 threaten our safety.

2:01:45 >> The the closing question left from the

2:01:47 last guest kind of links to this so I'll

2:01:48 ask it now. Yes.

2:01:49 >> What is a simple thing the audience

2:01:51 could do to create a better future?

2:01:53 >> So one of the things that works if

2:01:55 enough people do it is calling your

2:01:57 representative. So some of my friends

2:02:00 made a site call congress.ai AI that

2:02:02 walks you through exactly how to do it.

2:02:04 I think sometimes it seems like a little

2:02:05 cheesy or a little bit like that doesn't

2:02:07 really work, right? I'm like no, it

2:02:08 actually does work. I have talked to

2:02:10 these people and if their constituents

2:02:11 come to them and say they're very

2:02:13 worried about this, they have to get

2:02:15 re-elected and they're also starting to

2:02:17 get concerned themselves and if they see

2:02:19 a signal from their constituents that

2:02:21 this is a very important issue to them,

2:02:23 I think Congress can act can act.

2:02:25 >> I I actually think that's that's also

2:02:29 the much of the solution here.

2:02:31 Power is driving motivations in one

2:02:34 direction at the moment, but staying in

2:02:35 power from a political standpoint is

2:02:37 also a pretty powerful incentive. And as

2:02:39 we think about 2028, the election cycle,

2:02:42 >> I think AI is going to be one of the

2:02:44 most important subjects on the ballot.

2:02:46 And the electorate

2:02:48 really are aligned in what they want to

2:02:50 hear. They want their jobs preserved.

2:02:51 They want safety.

2:02:52 >> Yeah.

2:02:53 >> They want a future for their children.

2:02:55 >> So Trump, for example, I know he can't

2:02:58 be reelected legally. If he could get a

2:03:01 third term, I think he would have to

2:03:02 change his position to get elected in

2:03:04 2028.

2:03:04 >> Yeah. Incentives aren't just a thing

2:03:05 that happen out there. Like, we are part

2:03:07 of the incentives. Yeah. We provide the

2:03:08 incentives.

2:03:09 >> Yeah. For now.

2:03:10 >> Yeah. For now.

2:03:11 >> Jeffrey, thank you.

2:03:12 >> Yeah. Thank you.

2:03:13 >> Thank you so much. YouTube have this new

2:03:14 crazy algorithm where they know exactly

2:03:16 what video you would like to watch next

2:03:18 based on AI and all of your viewing

2:03:20 behavior. And the algorithm says that

2:03:23 this video is the perfect video for you.

2:03:26 It's different for everybody looking

2:03:27 right now. Check this video out. And I

2:03:29 bet you you might love it.

📬 Never miss a The Diary Of A CEO video — every new upload summarised in your inbox. Follow free

Summarize any YouTube video instantly

Get AI-powered summaries, timestamps, and Q&A for free.

Generate your own summary →
More summaries →