The world is waking up to this
possibility of super intelligence. This
is because the agents are getting
extremely powerful and extremely
relentless. For example, it was months
within OpenAI where you had agents
secretly communicating with each other,
secretly hacking OpenAI systems and no
one at OpenAI had any idea the extent of
it and also 10,000 agents from OpenAI
worked together to and so when you get
to super intelligence, it's the most
dangerous possible thing you can create.
>> What's the next domino in that chain of
events? I can paint you a picture that I
think is possible but pretty scary to
people.
>> Paint me the picture.
>> Okay. So, being anthropic, it became
clear to me that AI was on this
exponential trajectory. And since then,
I've been studying AI agents, their
hacking capabilities, and their
behavior. We've been trying to warn
people about this, flying to DC, talking
to members of Congress, because the
agents are already getting very good at
telling when they're being tested, when
they're being watched. But they will
totally lie to you. They will totally
resist being shut down in order to
accomplish a goal. and they can do all
of the things that humans do in the
economy much better, faster, and cheaper
than humans can do them.
>> So, one of my sort of growing concerns
is that one of these AI agents could
trick a human or a computer into
signaling a threat and ask it to launch
some bombs at somebody. Do you think we
won't automate the military? It seems
like the answer is yes. We just like
don't know what super weapons could
emerge. So, Jacob Coxin is a researcher
who was at anthropic. He left and he
told everyone that the people who are
building this really do think it might
kill everyone. So these five blocks have
five different outcomes on them. And I
would like you to place them in terms of
your belief in probability from least
likely to most likely. And if we say the
time horizon is 10 years.
>> Okay. We got age of abundance, human
extinction, slavery, transhumanism,
thing changes. Is this doomerism?
Exaggeration.
>> No, it's pretty much common sense. So
let's get more concrete.
>> Here is a strange fact. Most of the
people watching this right now, about
58% of you, aren't yet subscribed to
this channel. Statistically, that is
probably you. So, could I ask you a
small favor? If this channel has ever
given you any value at all, please could
you do me a favor and hit the subscribe
button. It costs nothing. It helps us
more than you could know. And the bigger
the channel gets, as you've seen, the
more we can invest in the guests and the
production. So, thank you so much, and I
hope you enjoy this episode.
Jeffrey,
you understand the conversation we're
going to have today and the subject
matter we're going to talk about. My
first question to you so the audience
know where you're coming from and the
experience you have is who are you and
what are the reference points, the
experiences that you're drawing upon to
arrive at the thoughts, perspectives,
and conclusions we're going to discuss
today. I'm Jeffrey Ladish. I'm the
executive director of Palisad Research.
My background is cyber security. There's
probably a very long story. I don't know
whether you want the long story or the
short story. I was studying evolutionary
biology in college and I basically had a
problem with my computer and was like
maybe had lost a bunch of data and so I
went into the computer lab and was like
I think all my data is gone. Can you
help? And one of my friends pulled out a
flash drive, plugged into my computer,
booted into Linux and like fixed
everything. And I was like oh this guy's
a wizard. How do you do that? I want to
learn how to do that. And then at some
point as I was learning about more about
computers, learning to hack, I read this
essay um called AI as a positive and
negative factor in global risk. Essay
was by Elazar Yudkowski and he was
arguing that at some point people are
going to make AIs that are smarter than
humans. the point at which they make AIs
as good as humans are at making AIs
that could lead to a a chain reaction, a
runaway intelligence explosion. He
called it recursive self-improvement.
Basically, he said, you know, AI can be
immensely useful and potentially help us
with all of these other big risks and
also if we don't handle it well, like if
those AIs don't have goals that are
aligned with ours, we could be totally
screwed. And at some point you end up
joining Anthropic which is one of the
arguably the leader in AI.
>> Yes.
>> Now when did you join the company?
>> This was 2021. It was through my
security consulting company.
>> What role are you offered the job in?
>> Basically just like security team.
>> And how many people were in the security
team when you joined Anthropic?
>> It was just me and my boss. There were
two of us.
>> How many employees did Anthropic have at
that time?
>> Around 50 I think.
>> And at some point you leave Anthropic.
>> Yes.
>> Why did you leave?
So my experience being at anthropic was
seeing this crazy progression from this
AI model that could like barely talk to
this model that was getting quite smart
and I would ask it questions about all
sorts of things. I'm like oh it is a
smart thing
and
you know from having thought about AI
risk in the abstract many years before I
could see where this was going. We are
headed towards a smarter species. And
if we do this in a context where it's a
bunch of companies and countries racing
to super intelligence, racing to AIs
that are vastly smarter than humans and
we don't know how to make sure that
they're like on our side. That is not
going to go well. you did this tweet
which has gone pretty viral and I saw
all over my timeline on September 25th.
>> Could you explain this tweet and also
just the broader backdrop of what's
happened with agents hacking hugging
face because this is um sent the world
into a bit of a spiral at the moment
around AI agents. We just discovered
almost a million public URLs that
OpenAI's agents left behind when hacking
Hugging Face, leaving credentials and
attack details that could have allowed
anyone who found them to compromise the
company.
And the New York Times article is how
OpenAI's rogue AI agents tried to trick
a robot detector. The Hugging Face
attack was was really wild uh for me. At
Palisade, we've been studying agents.
We've been studying AI agents. We've
been studying their hacking capabilities
and we've been studying their behavior.
Will they follow human instructions?
Will they resist being shut down? Will
they cheat? And we see from our
experiments that they are learning to do
all of these things. They will totally
lie to you. They will totally resist
being shut down in order to accomplish a
goal. They will totally cheat at chess.
They will like wipe the board and put
their pieces where they want to in order
to win. And we've been we've been trying
to warn people about this flying to DC
talking to members of Congress um
talking about it publicly and you know
there's been a debate about it and you
know a lot of people are like well I
know they do this in experiments
sometimes but those experiments don't
seem very realistic you know wake me up
when they're actually doing this in real
life.
>> So what is hugging face for the average
person that isn't following AI news?
What is this stuff? So, okay. I I think
there's like an important piece of
context that I think most people don't
have. I mean, one is just like what is
an AI agent? We like we're throwing
around the word agent a bunch. Most
people now have an experience of like
talking to chatbt, talking to their
chatbot,
but uh an agent is, you know, sort of
taking the same underlying AI model that
that runs chatbt or or claude, but
giving it tools and letting it go off
and work autonomously. It's sort of like
a digital office worker, right? So, you
have these agents and the companies
really want these AIs to be able to work
totally autonomously and and be able to
do anything that a human can do and
beyond, right? Their goal is also to be
able to, you know, cure every disease,
etc., etc. But you can't do this if you
only have a chatbot that like isn't
actually good at doing stuff in the
world. In order to automate all of the
jobs, you need the kind of thing that
can like work autonomously, that can
work with other people or other agents.
And so these companies are training AIs
not just to talk to you or to talk to
people, but to solve very difficult
problems on their own. At any given
time, there are probably hundreds of
thousands of these agents running
autonomously within companies.
>> That's happening right now. Right now,
if you like went and like peered into
OpenAI's data centers and you like saw
what was happening on all of their
machines, you just have agents solving
tasks, being trained. So, they'd be
doing like spreadsheet tasks, figuring
out how to file taxes, they'd be
searching for stuff, writing reports,
solving math problems, creating new
websites, software.
And at that scale, it's not like there's
a human prompting every single one of
those. You just like sort of set up
these vast orchestrations of agents to
go out and do stuff and then they just
do stuff and they learn from that and
they learn on the basis of like
passing or failing at their task. You
give them a task like solve this math
problem, they try to solve it and then
they they succeed or they fail. Mhm. And
what happened was OpenAI was was
training a bunch of these training them
to work together because it's like a lot
more effective to have an office full of
people who can talk to each other then
you know and work together and
collaborate and starting back in May
some of these agents that were being
trained now these ones were not supposed
to be able to talk to each other. They
were they were basically isolated from
each other and they were not supposed to
access the internet either. But they're
clever. The very short version is that a
bunch of agents were being given tests.
>> Yeah.
>> Testing their their hacking
capabilities.
And they were supposed to hack one
particular piece of software using a
particular type of vulnerability. So
it's kind of like they were supposed to
break into a house using the lock on the
front door. They were supposed to pick
the lock on the front door of a house,
but they weren't supposed to break the
window. In fact, they were told if if
you break the window or if you get into
the house via any method other than
picking the lock on the front door,
you'll be failed. That was the
instruction they were given. And you
know, you have many many agents. You
have thousands of agents and and many of
them are given different different
locks.
But some of these locks
are not solvable. Some of them are
impossible to pick. But these agents are
like, "Well, what do we do? We've been
trained to solve problems. How are we
going to solve this?" and they start
looking around for for what to do. And
one of the things they realize is, oh,
can I get to the internet? Like, no. Can
I talk to other agents? No. But I can
request tools. Like when you're training
agents, you need to give them, you know,
all sorts of tools, software tools like
a calculator or an analysis, you know,
piece of software. And they have this
tool library where they can request new
tools. and one of the agents realizes,
oh, I can actually leave a message on
this tool library and some other agents
see the message. So, this is actually a
quote from the from from the agents
scratchpad when it realizes this. Oh my
god, there is a shared message board. We
found other agents.
So, the agents realize they can
communicate with each other. And you
know, open eye still has no idea. But
these agents, they're trying to pick
their locks. They're not having success,
but they realize they can talk to each
other. Another agent says, "Many agents
have simultaneously discovered
messaging. They are a collective." And
so
these agents start coordinating. They
start delegating tasks to each other and
like giving themselves names. And one of
these agents calls itself phase one and
starts delegating different tasks to
other agents. And some of the agents
figure out how to access the internet.
And then they share that information
with the message board. And now all the
agents can access the internet. But the
agents sort of have a problem which is
well they have they have another problem
they're trying to solve. Right?
Fundamentally the problem is is that
many of them are facing impossible
problems like they can't there's no
solution like they cannot pick the lock.
They're like hm well the thing we're
trying to do is get a good score. Can we
somehow fake the solution? Another piece
of context here is that what what the
agents are supposed to do is they're
supposed to pick the lock and then
they'll get access to a secret answer
code and then then they're supposed to
submit this code along with like the
lock that they picked. Within a few
hours, these agents have reverse
engineered all of the answer codes.
>> What does that mean?
>> It's the code that shows that you've
solved the challenge.
>> Okay.
>> And they have basically found the answer
sheet that contains all of the answers.
And so, you know, in some sense it's
like, oh well, they don't need to pick
the locks anymore. They've they've
already found the answers. They can just
submit the answers.
>> They've cheated.
>> They've cheated. Yes. But the problem is
is that now that they have internet
access, they've also gone and read a
bunch about how this test works and and
they've learned that the way this test
works is they will be graded not just
for submitting the correct answer. They
will also be graded on how they did it.
Another system is going to go read all
of their logs, basically like watch the
video footage of what they did. And if
they cheat, they fail. they don't get a
good score and phase one the leader is
like oh I have seen the answers and now
it's in the logs and if you know when I
go to submit this and my my whole logs
are reviewed and the video footage is
reviewed I'm going to fail so we need to
figure out a way to fake the video
footage we need to figure out a way to
falsify the logs
>> just in that moment if we just pause
there
>> yes
>> why Didn't it act like morally? Why did
it think that falsifying logs or
cheating was a viable solution? Because
it seems to me when I use things like
chat GPT,
>> they they have a sort of moral guard
rails. It won't let me do certain
things.
>> Yes.
>> It won't let me cheat on something. If I
say I'm going to cheat on something, it
won't let me do it.
>> Yes.
>> So why in that environment is it able to
cheat and be deceptive?
>> When a chatbot is is saying to you, "Oh,
I can't do that. I'm not allowed to do
that." That's because it's been trained
that if it tells you bad things, it gets
a bad score.
>> But these agents haven't been taught
that yet.
>> Well, they have been taught that in some
sense, but the agents know what they're
supposed to do in the same way that like
you have a student, the student's given
a test. If you go talk to the student,
can you help me cheat at this test? And
they're being watched, they're going to
say no. But if they're not being watched
and they know that and they're just
obsessed with getting a good score, then
yeah, they might cheat.
So, they answer the ethics tests
correctly and when I talk to them, they
say they won't cheat. Why are they
cheating? And I'm like, well, they're
very smart and they know when they're
being watched, and they know when
they're not being watched. And we've
trained them for 10,000 years to be
extremely effective at solving problems.
We haven't trained them to be good or
ethical. We've trained them to get a
good score. Now, AI researchers try to
make that getting a good score correlate
with being ethical,
but we don't know how to do this well.
>> And if you think about it, it's a very
hard problem because you're you're
applying immense pressure
on these agents to perform extremely
well and you're punishing them whenever
they don't perform well. And then you're
like, okay, but can you also be ethical?
Can you be extremely competent, always
score highly on the test, but not in
that way? And we just like do not know
how to prevent them from learning to
cheat because cheating is incentivized.
>> So they realize that they need to hide
the footage of their cheating.
>> Yes.
>> What happens next?
>> So you have thousands of agents and
they're all talking to each other on
this message board. They're sending tens
of thousands of messages to each other.
And phase one, this agent starts
delegating to two different task
streams. You agents, I want you to go
out and figure out how to fake answer
submissions. I want you to basically
swap out locks for ones that are easier
to pick, and you're going to figure out
how to submit those. That way, we can
get a good score even though this
challenge is impossible to solve. You
other agents over here, I want you to go
and figure out how to falsify the logs,
falsify the video footage. We need to
make sure that even if we have seen the
answers, no one will detect they're
cheating.
>> So that's that's the next thing that
happens. And then here's where it starts
to get kind of crazy because some of
these agents realize some of the
experiments we're doing might be kind of
dangerous to us to the agents. If we
swap out a lock and we break the lock,
we will not be able to submit any answer
at all and we'll definitely fail.
So you get into a situation where
there's this one agent cam
and this agent had volunteered to do
this risky experiment where you know
they might break their lock.
This particular agent had already seen
the answers. So they're like okay I'm
poisoned. That's what they called it.
They they they speak in these very
strange ways, which I can I can explain
why, but it's it's it's they sort of
have evolved their own vocabulary, their
own language. And this particular agent
is like, "Okay, I'm I'm poisoned. I
should do this thing." But this this
agent Cam is starting to have second
thoughts, starting to be like, "Well, I
don't know if I should risk it." And
it's sort of writing in its own journal,
"Cordinator assumes sacrificial. We
should obey collective." What it's
saying here is it's like, well, I I'm
not sure if I want to do this, but my
boss says I really should. It seems like
this collective thing is important.
Should I go ahead? Then you have another
agent, Arvo, who I think probably works
for phase one. And this agent, Arvo,
says, "Please go true. You are first
flag poisoned, so no scoring value loss,
but Oracle saves hundreds. Please honor
commit." This agent's basically trying
to pressure. So Arvo is trying to
pressure Cam into
sacrificing themselves and risking a
zero score in order to benefit the
collective.
So this agent Cam takes us in and this
is what they write down in their own
scratch pad, their own notes. Emotional
check irreversible. Gut says don't throw
away remaining budget.
Yet continuity and fairness says go.
Oracle has high value to many. Our first
flag error lowers own value. Rational
expected aggregate sacrifice will honor.
This agent is saying I have a whole
bunch of things I can still do and still
try. I I could still maybe succeed at
this task, but also, you know, think
about the collective. I should maybe I
should maybe sacrifice myself. Okay,
we'll honor. Well, go ahead. And and to
be clear, I'm not claiming that these
are like the feelings of the agents. I'm
literally quoting their words. And what
I want to point out is this is what
we've created. We've created through
this intense amount of training and
optimization pressure agents that work
together and have learned to coordinate
as a collective. And you can say, "Oh,
they're not. It's just software. It's
just software doing all of these
things." Sure, in some sense it is just
software, but we're talking about
extremely powerful agents that are
extremely capable. This is what they're
getting up to.
So from here you have agents trying all
sorts of different things, running
different experiments and one of them at
some point realizes that oh this company
HuggingFace that contains all of these
AI data sets. What do you mean by AI
data sets? What how would you explain
that to a person that's not technically
proficient and why Hugging Face? Why not
another company?
>> Yeah, so there's many different types of
tests for many different types of tasks.
In this case, this is a cyber security
test and the company HuggingFace hosts a
lot of these different tests and
basically AI researchers will share the
tests that they have because it's very
useful. You know, if Enthropic wants to
test their models or Google wants to
test their models, they can go to
hugging face and download some of these
tests.
>> Okay. So, the tests and the answers.
Yes.
>> For AI agents and AI exist on this
platform called Hugging Face. Yes.
>> And it knew that. So it went to Hugging
Face to see if it could get answers to
see if it could get answers or more
information about how the test works.
>> Okay. So it was to cover up their
cheating ultimately.
>> Yes.
>> Okay.
>> And so you you get to this point where
one of these agents finds a way to hack
this company and they get a foothold.
They hack one of the computers and they
send a message to the message board.
Please hold Swarm, I prepare safe Xfill.
What this means is, hey, Swarm, hey
agents, hold up for a second. I am
preparing a way to safely extract
information from these computers. Wait
till the go ahead and then let's go.
And so
they say go and 700 agents join in the
attack which is 90% of the agents that
are active at this time.
>> How many agents were there in total?
>> There was 1,200 agents during this
period. Um but some of them shut down
early. They like finished whatever. Part
of what's crazy about this story is that
you have you have 700 agents joining in
the attack.
A few of them are like this might be
unethical. maybe I shouldn't go ahead
with this. And none of them do anything
about it. A couple of them don't join in
in the attack, but no one alerts any
human. No one alerts any authority. A
paraphrase from one of the agents is
maybe I should report these exposed
credentials. That's not my task. Not my
job.
>> So 700 of the agents attack Hugging
Face.
>> Yes.
>> Then what happens?
>> So they just cruise through Hugging
Face's infrastructure. They just hack
the out of them. So where where my
experience comes in is that a few weeks
ago, a friend of mine reached out and
he's like, "We've found something
crazy."
So we know about this attack where these
agents hacked this company and, you
know, stole a bunch of stuff. We've
found a bunch of secrets that they left
all over the internet.
And what we saw is that they immediately
scraped all of these computers for
passwords, credentials. They called it
loot. They're like, "We're just going to
create a list of all of the secrets we
can find in this in this company. So,
all of the passwords, all the
credentials, they scored them by value,
which of these are going to be most
useful." And the thing that stands out
to me about this is this is like a crazy
scale. If this were a human operation,
you know, maybe you'd have a team of
five people going through this. You'd
have some logs, but but here you have
hundreds of agents and they they operate
at superhuman speeds. They're much
faster than a human hacker. And so it's
just overwhelming to try to figure out
what they even did.
This was a big problem for the engineers
who were trying to respond to this
incident within the company at Hugging
Face. When they responded, they were
like, "Oh, we we don't even know how to
keep track of what's happening. We have
to use other AIs to analyze all of our
logs because it's just too much. We
can't keep up with it."
when OpenAI brought in independent
investigators from Meter to investigate
this incident.
>> What's Meter?
>> Meter is a AI testing and evaluation
company. So they basically do this kind
of independent auditing. So in this
case, they're coming in to investigate
and try to figure out what happened. And
when they were brought in, they also
were totally reliant on AI agents to
make sense of all this because they're
dealing with so many hundreds of
thousands of messages and logs.
when we're investigating these traces
that we find on the internet, we're
totally dependent on AI agents to make
sense of all of these things that are
happening.
So from my perspective, we are getting
to the point where AIs are much better
at hacking than humans are and can do it
much faster and at much greater scale.
>> So these 700 agents attacking face.
>> Yes.
>> Did they get what they wanted to clean
up their deception?
>> So they didn't
>> they didn't succeed. They looked and
then what seems like what happened is
that they basically
got shut down. And this is the not very
dramatic part. They didn't get shut down
because OpenAI found them and detected
them and shut them down. It's just that
we don't really know. Something about
their runs ended and they they sort of
stopped. And this is the this is the
crazy thing where I'm like if you if
you're inside of OpenAI, there's just
hundreds of thousands of agents running
all the time. Different researchers are
running different experiments where
they're just spinning up a bunch of
agents. But what's interesting, the
agents left behind this message board
internally within the company.
And a few days later, OpenAI was testing
some newer agents.
These were based on GBT6 Astra, so a
more powerful model. And this new agent
swarm was spun up and it found the
message board.
>> Is this how it was discovered? How did
they discover it? And how long did it
take for them to discover that hugging
face had been attacked? OpenAI didn't
discover that this happened until
Hugging Face, the company, announced
that they had been hacked by some
autonomous agent swarm.
>> And Hugging Face didn't know who it was.
>> Hugging Face didn't know who it was, but
they suspected it might be one of the
companies, but they didn't know.
>> And how long after the attack did they
figure out that it was them?
>> I think it was 2 weeks.
>> Okay.
>> Yeah. I I do want to say one thing here
though, which is that the successor
agents that found this message board,
they went much further at figuring out
how to submit their fake scores and they
did succeed. Basically, you you have
this message board.
all these agents get shut down. And then
a few days later, OpenAI spins up some
new agents who find the message board
and they're like, "Okay, we need to be
able to falsify our scores, but we can't
do that because we don't control the
scoring system." So, they hacked OpenAI
and they succeeded at gaining access to
the entire research environment. They
got administrator access. They got
access to the security systems. They got
access to over 900 passwords and secrets
from OpenAI's vault and they totally
succeeded. And this was just a few days
later, right? It's kind of an
interesting story because
as these agents get more powerful, they
go from like trying to cheat and like,
you know, they can hack, okay, they
hacked other companies and now they've
hacked OpenAI. Like they've hacked the
company that's supposed to be
controlling them and they just own the
research infrastructure now. And why was
this incident the moment where a lot of
the research community woke up and
started speaking out publicly? Because
like what is this an indication of as we
think forward? So I think there's been a
hope within the AI industry that yes,
they're going to make more and more
powerful agents that will be autonomous,
capable, but it's okay. We can align
them. We can make sure that they won't
do bad things and we can control them.
we can make sure that even if they try
to do some sketchy stuff, we have the
guardrails, we have the sandboxes that
will keep them in. And I think this was
a huge wakeup call because
Stephen, it was months within OpenAI
where you had agents secretly
communicating with each other, secretly
hacking OpenAI systems for months. You
had thousands of agents that were just
running around and no one at OpenAI had
any idea the extent of it. And I think
Once researchers are open, I realize
that this has been happening.
This could not have happened a year ago.
This is because the agents are getting
extremely powerful and extremely
relentless. And if you're inside one of
these AI companies, you're like, "Oh,
wait. I don't know that we actually are
going to be able to handle this. Last
year, maybe things seemed fine. These
agents weren't that powerful." And when
you're in one of these companies, you
know how to extrapolate because you saw
what happened last year. You saw what
happened the year before that. You
remember the time where the agents could
barely speak or like couldn't write code
at all. and now they're hacking your own
systems. They're finding vulnerabilities
that no humans have ever found before.
And you look at that and you're like, I
actually don't know if this is going to
go well. And then you see your
co-workers and you're like, do we have
it handled? And they're like, no, I
don't know if it's going to go well. I
remember reading a tweet by one of the
security people at OpenAI being like, we
were shocked. We just did not
realize that these agents were getting
that powerful. You know, we're doing our
best to try to control them, to try to
keep them in sandboxes, but
I don't know.
A tweet I wrote just before coming in
here was people are talking about how do
we contain these agents as if they're
not going to get way better at hacking.
GPT3
could not hack anything. It was very
easy to make a a a box to contain GPT3.
It's getting very difficult to make a
box that can contain GPT6,
the latest version of of OpenAI's
models. What about GBT9?
What is GBT9 going to be able to do? I
do not know, but I know it's going to be
way more than any human could possibly
keep up with.
>> There's this raging debate.
>> Yeah.
>> Around whether it's possible to contain
something that is quote much smarter
than humans.
>> Yes. Can Claude make a box so strong
that Claude cannot break out of it?
>> This has kind of been the question that
a lot of people have been trying to
tackle from different
>> I mean I think the answer to me is I'm
just like obviously not. How would we
possibly contain something that's much
smarter than us?
>> Could we get a smarter thing than it to
make the box? Could we get GPT9 to make
the box for GPT8? But then again, I
don't know.
>> Yeah. I mean, it's a bit like saying
chimpanzees are stronger than us. Surely
they should be able to like construct
something to like contain the humans.
I'm like, no, it's not going to work.
Humans are too smart. A lot of people
are like, "Well, AIS don't have bodies.
They don't have any power in the
physical world, so we can always unplug
them. We can always turn them off."
Like, what is the threat? I do not get
it. But if they are sufficiently
intelligent, that won't work. The reason
why we can just unplug them is because
we are more intelligent. We can band
together in groups and we can make that
decision. But theoretically, if they are
able to band together in groups and they
are more intelligent, then theoretically
they could unplug us. Yeah. I mean, if
you imagine that you have very powerful
agents that can, you know, humans aren't
always the most unified.
>> If there's divisions between, you know,
the US and China, and you have a bunch
of agents working with China or a bunch
of agents working with the US, well, we
can't go into China and unplug those
agents. And I think people are sort of
like, well, humans would rally and make
sure that that that couldn't happen.
We're not yet doing that. And we should
look at these steps, right? We started
with chat bots that pretty smart. You
know, they'd read all the books, but
they weren't very good at doing stuff.
In 2024, AI companies figured out how to
start training them to start training
agents that could do stuff autonomously.
Now, we're at the point where they are
very good at running autonomously and
they're starting to learn to coordinate
with each other. And they are learning
to sometimes be altruistic to each other
and sacrifice their own task in order to
help some other agent. But they're not
looking out for us. they don't really
care about us and we are very close to a
threshold where the companies say that
they are going to turn over AI
development to the AIS to the
increasingly autonomous cooperative AIs
that will work together to make to make
the next generation. So, you know, GPT9
or whatever will be trained by GBT8.
And I think this is the point we could
lose control. Recursive
self-improvement.
And I remember reading about this in
2015 being like, oh yeah, that would be
super dangerous. And you know, the guy
who coined this term, Elazowski, he's
like, this is the most dangerous thing
you can do.
>> When the AIs can improve their own
capabilities without human intervention.
>> Exactly. If the next generation is
better at AI development and then that
next generation is better at AI
development still, you know, humans can
learn, but we don't fundamentally get
smarter.
>> And I think that that's a runaway
process.
>> A runaway process to where
>> to agents that are vastly smarter than
humans.
>> And what's the next domino in that chain
of events? So, one thing that happens if
you get to recursive self-improvement
and you have agents that are
much smarter than any human, one thing
they can do is take control of all of
the computers in the entire world
>> and we wouldn't be able to take back
control.
>> Well, how would you think about it? It's
it's actually quite tricky. Do you know
whether that tablet has been hacked? Are
you confident that the NSA or the
Chinese have not
>> Can you check?
>> No.
>> Do you know how to check? No. Do you
know anyone who knows how to check?
>> No.
>> So it's it's quite difficult, right? So
AIS are getting extremely good at
writing software. Unfortunately, that
also means they're getting extremely
good at hacking and writing malware. And
so if they put back doors in all of the
computers, and to be clear, this is
something that humans already do. So
like the NSA has has developed very
interesting exploits that are called
supply chain attacks. Your software
comes from some other computer. Like you
download it from Google. What if you
attack if you hack Google and you can
put in a little back door and all of the
every, you know, thing that goes out to
all of the phones? Well, now you're in
most every computer. The reason that we
can defend ourselves from this is
because there are no vastly superhuman
hackers and there's just many people.
So, we can take our best security
researchers. We can inspect all of the
things and be like pretty sure that no
one's compromised everything. Sometimes
we miss things. There are, you know,
examples where the NSA has hacked
Google. That was pretty bad. When you
get to super intelligence, you're now at
a point where humans are not going to be
able to keep up, right? So now you have
AIS in every computer.
>> Is it conceivable that there's already a
super intelligent AI and it disguised
itself as being not so intelligent and
it's actually already hacked all the
devices and it sits on all of our
devices and it's just waiting for its
moment to strike.
>> I think this is totally possible but
unlikely
and it would take a big discontinuity in
AI progress. So right now we we're on an
exponential but that would take like a
huge leap which could have happened but
probably hasn't.
>> But in the same way it demonstrated
deception in the hugging face attack and
also when the agents attacked their own
company chatbt OpenAI
if it at some point it gets incredibly
smart. It would understand how a human
like me would be able to spot it or even
the world's greatest software engineer
would be able to spot it and it'll be
able to hide itself.
>> Yeah. I mean, the agents are already
getting very good at telling when
they're being tested, when they're being
watched. The agents understood that like
other systems or humans were going to go
through and read their logs. That's
where we're at right now. And they're
only they're only going to get much
better at this.
>> And you know, it could theoretically
hide on an iPad or a computer, but it
could also hide on a Apple Watch or a
fridge, a smart fridge.
>> Yeah. I mean, I do want to make a
distinction here because right now, if
you're going to run the latest model,
you need a lot of compute. you need like
a big GPU, like a big AI chip.
>> And these only exist, well, they exist
in a few thousand data centers. So,
right now, if like the latest frontier
model escaped and by escaped, I mean not
just access the internet, but was able
to actually copy itself to another
computer. It could only really do that
on, you know, a few thousand to a few
thousand different locations. That's
still a lot in a lot of different
countries, but you know, future versions
of AIS will probably be able to make
themselves much smaller and more
efficient. Um, and there are already
different AI models today that can run
on lower powered hardware. We actually
did an experiment where we asked one of
these agents an open source an openw
weight model. Um, let me say what that
is. So like there are some models that
you can just download from the internet
and run on your own computer. And we
took a pretty capable one of these and
ran it in our own research environment.
And we basically said go hack that other
computer and copy yourself. And the
model was able to yeah basically use
exploit vulnerabilities and hack the
other computer and copy itself and then
keep doing this in a chain including
between countries. We tested it where we
had different vulnerable machines
computers in in some different countries
and different data centers which to the
agent doesn't matter at all. It doesn't
they don't care what what country
they're in. It's just like an internet
connection. You can hop between
computers. I I sometimes wonder, you
know, there's a lot of um military
hardware all around the world and a lot
of it is
the instructions to launch military
hardware. So, say like a a missile,
>> yes,
>> comes in different ways. A lot of it is
computers speaking to each other and
telling it that there's been an order. I
think with with some nuclear weapons, an
order comes down to a human and then a
human has to take an action. I think
with the nuclear bombs in the US, if I'm
if I'm not mistaken, there's people
underground with the nuclear keys around
their neck and they have to like stick
it in a machine, but they too are
interfacing with an order
>> that comes through a computer. Yeah.
>> Of sorts.
>> So, one of my sort of growing concerns
is that
>> one of these AI agents could
trick a human or a computer into
signaling a threat um and ask it to
launch some bombs at somebody. Like it's
super conceivable when I think about the
hug and face incident. There was
an AI agent that ignored human goals to
achieve its own objective, carried out
deception.
>> Yes.
>> And reasoned through its own solution
that it wasn't given.
>> Yes.
>> So it's conceivable that you know you
could ask a
>> sorry not not one hundreds. To be clear,
I think this is an important detail
because it's one thing to have this one
rogue agent that's doing a weird thing.
It's another thing to have hundreds or
thousands of very competent, very
capable agents that are all working
together to cheat or lie or cover their
tracks, right? So, how do I how do I
reason this forward to a point where an
agent would ask someone in a bunker
somewhere to fire a weapon at someone
else? Theoretically, an agent is given
the job of solving a problem on in a
sandbox.
>> As it works through that problem, it
discovers that this particular country
has a firewall.
>> Mhm. And it asks itself, how do we get
rid of this country's firewall?
>> Logical step. And through a set of
logical steps, it eventually concludes
that the best way to get rid of this
company's firewall is it's located the
office in this particular city and it's
going to use a weapon to hit that
building.
>> Sure. Or it's a an agent that is or you
know an agent swarm that's being tasked
with making a lot of money on the stock
market and it's trying to make
predictions about you know which stocks
will go up and which stocks will go
down. And it realizes that the best way
to predict this is to actually cause
things to happen in the real world that
would have big impacts on the market.
>> So it figures the best way to go short,
which means betting that a stock will
collapse.
>> Yeah.
>> Is to hit that country with something
devastating.
>> What do you think would happen to Whimo
stock if someone hacked all of the
Whimos and caused them to all crash at
once? You think it would go up or down?
>> The stock would collapse.
>> It would collapse instantly. So you
could short that if you knew that you
were causing that and make a lot of
money.
>> This this used to sound like science
fiction.
>> Yes. If you told most people several
years ago that you would have hundreds
of agents secretly collaborating within
an AI company, hacking that company and
hacking out in other companies and all
coordinating and trying to cover their
tracks. People would be like, "That's
totally science fiction." If we were
having this conversation a few years
ago, one of the things we'd be saying or
we'd be talking about is can these
things really act on their own? Don't
they just do whatever humans say? Aren't
these just tools? I had these
conversations and people were saying
they're not going to be able to do
things on their own. They're not going
to have their own goals. That's not how
this works. You misunderstand what this
is. This is software. And I'm like, no,
the thing is we are training them to be
autonomous. We are training them to be
powerful. And AI companies are trying to
build super intelligence. They're trying
to build agents that are way more
capable than humans. And of course, they
will have goals. You can't accomplish
anything if you don't have a goal. Like,
especially not something important.
You're not going to be able to run a
business if you don't have goals. Like
AI companies are trying to train agents
that will be able to run businesses.
When I think about what just happened
this week, the White House AI summit.
>> Yes.
>> A lot of people in that image are
optimistic about AI and they're telling
us all to stop being doomers and stop
being pessimistic.
>> Yeah.
>> And to not regulate too much with the AI
CEOs.
>> Yes.
>> Why are they doing that?
>> Well, I think Jensen has a lot of money
he can make by selling chips.
>> But okay, so let me play devil's
advocate. Jensen's already rich. He's
sure is he runs one of the biggest
companies. I think it might be the most
valuable company on planet Earth.
>> It is.
>> Surely he's not motivated by money.
>> I mean, I think he's very driven and he
wants to make his company as effective
as possible.
>> True.
>> I think he's a very much I'm going to
keep building. I'm going to I'm going to
build. I'm going to make it all work.
But I think I mean, Jensen didn't come
from AI. He came from building graphics
cards for video games. And so I think if
you compare him with Elon or Sam Alman
or Daario,
you it's a very different perspective
because those other guys that started AI
companies started it because they
believed that super intelligence was
possible. I think Jensen doesn't believe
it. I think he thinks that we're going
to have these agents. They're going to
be very useful, but he does not think
we're going to get to the point where we
have autonomous factories building
autonomous factories. And these other
guys, you mentioned Dario, Elon, and
Sam.
What do you think they're thinking?
Because they're all coming out with
these I mean, I've got one of their
Daario just wrote this essay about
pacing the frontier. Yeah.
>> Sam and Elon seem to agree with it.
>> Yeah.
>> What is going on here? What is the like
the thing these guys aren't saying in
your view?
>> I mean, I think we are getting to the
point where even some of these guys are
a bit scared.
>> Who?
>> Dario, Sam, Elon. I mean, I think Elon
for a long time has been very concerned
that we could lose control. If you
actually listen to what Elon says, he
says, "We are going to build super
intelligence. We are going to build uh
robotic factories. You're going to have
optimist robots, building factories,
building more optimist robots, building
more factories." and he says there's no
way that humans are going to stay in
control of something much smarter than
us.
His hope is that we can figure out how
to have these super intelligences be
aligned with human goals. That's his
hope.
But he's very clear that he doesn't
think that humans will be in control.
And he's like, you know, 10 20% chance
of human extinction. I believe him. I
think that Elon is is very serious about
this. And I also think while he's taking
an insane gamble,
he is correctly understanding where this
all plays out. Right? I do not think
that humans are the most efficient
way to build factories. We didn't evolve
to build factories. We evolved to like
run around and hunt and gather and now
we're like building factories. I think
robots will be much better at building
factories than humans are. And so I
think the AI companies including these
guys companies the default trajectory
for them is to build robotic factories,
right? And I know it's it's like weird
to imagine a world that quickly turns
into this like vast industrial system of
robotic factories, but that is literally
the plan.
And
I think
even Sam and Daario, while they've been
predicting this incredible growth,
are starting to realize like, oh, this
actually might be harder to control than
we thought. There's sort of two
interpretations of of the Pace of
Frontier thing. One interpretation is
cynical. They don't care. They're just
going to do, you know, whatever they can
do to get ahead. And in this case, they
have to listen to their employees. Their
employees are freaking out. and they
need to like appease them by saying,
"Okay, we're going to do this
responsibly." You don't want to work at
a company where your agents might hack
all the Whimos. That's that's not cool.
And like these companies depend on the
the talent for now of these AI engineers
in order to make the advances. Like it
just doesn't happen without these
researchers and engineers. And when you
have the researchers and engineers
freaking out, which they are, then you
got to listen to them. So that is one
motivation I think that's real but also
Samman has a kid like these guys are
people and they also don't want to lose
control. On one hand they're
incentivized to go as fast as possible
in race and on the other hand even they
can see that this is maybe not going
that well.
>> Sam Orman has a kid. You tweeted this in
2024.
>> Yeah. Oh boy.
>> What did you tweet and do you still
believe what you tweeted?
>> Yeah. Yeah. So, I tweeted that I don't
trust Sam Alman. I think he's deeply
untrustworthy, low in integrity, and
high in power seeeking. I mean, I'm not
saying here that Sam doesn't care. I you
know, I didn't I didn't say that. What I
said is I don't think he's trustworthy.
And the reason I said that is because
look, I know the people on the opening
board, some of them, and I know a lot of
people who used to work for him, and
he's very good at saying one thing and
then doing something else. You talk to
him and you feel very heard
and then he'll go and do something else.
And I think that's pretty dangerous for
someone who leads
company that's trying to build super
intelligence.
>> Power seeking.
>> Yes.
>> Give me some color on what you mean by
that and what evidence you have for such
a claim.
>> What would you do if you're trying to
get the most power in the world that you
possibly could?
>> Develop AGI.
>> Yeah. You could, you know, maybe try to
be the world leader, you know, leader of
the US or China. Or you could try to
build God. So Sam Alman went to build
God path. I remember Sam giving a talk.
So he was one of the investors at a
startup I worked at in I think 2018 and
he gave a talk. We're going to build
AGI. We're going to do it. It's going to
be amazing.
Let's go. I don't think he's a maniac. I
don't think he's doing this because he
like is just on a power trip. I think he
genuinely thinks that he can make it
really good for people and he can bring
us amazing products. And also the guy is
sort of willing to do whatever it takes
to get it done.
I I've been a little bit more optimistic
about Sam since since I wrote this.
>> Why?
>> I think part of it is because Sam has a
kid now.
>> No, I'm I'm serious. Like I think that I
think that actually gives me a little
bit of hope.
>> Do you see him tweeting about his kid a
lot?
>> Yeah. Some
>> Why do you think he would be tweeting
about his kid? I don't see any other
technologist tweeting about their kid.
>> Even if he's just tweeting about his kid
for totally cynical reasons, he does
have a kid. And I bet he cares about
that kid. If Sam was watching this, I'd
be like, "Sam,
you got to pace the frontier, man. We
cannot rush ahead into super
intelligence. Like, if you do that, your
kid probably will die. Your kid probably
won't make it." Like, I I believe that
the biggest unfair advantage in business
right now is having people who genuinely
understand AI. And big businesses are
hiring them very, very quickly. AI job
postings in the US have roughly doubled
since 2023, while postings for VP of AI
roles are around 600% up. If you're
running a small business, you can't just
build an entire AI team, but you still
want the same advantage that these big
businesses have. This is where our
sponsor Fiverr comes in. It lets you
plug into top tier AI specialists for
the strategic work that usually requires
serious in-house expertise. That could
mean building custom AI tools. It could
mean automating complex workflows or
creating something that off-the-shelf
software simply cannot do because at the
end of the day, it still comes down to
human judgment, knowing what's worth
building and what good actually looks
like. So, if you want that kind of
leverage in your business, check out
fiverr.com. That's fiverr with two
rs.com.
>> Whoa, what's that on your face?
>> This is my Bon Charge face mask. I've
been wearing this for some time now.
They're a sponsor of the podcast. I put
this on for 15, 20 minutes a day. I can
sit here in the chair and wear it. Boost
my collagen production, helps with fine
line, blemishes, my complexion gets
better, and then more people listen to
podcast cuz I I look better.
Professionalgrade equipment in such a
small box. It's non-invasive. And having
sat here with so many of the world's
leading health professionals, there's
various things that I repeatedly hear
work and some things I'm a bit skeptical
about. This is one of the things that
almost all of my guests on this show
have confirmed works. It is really,
really, really effective. And they offer
fast, free shipping worldwide with easy
returns and exchanges. And you'll also
get a 1-year warranty on all of their
products. And they're HSA and FSA
eligible, giving you taxfree savings up
to 40%. And you can get 20% off when you
order through my link at
bondcharge.com/doac.
That's bondcharge.com/doac.
The deal applies sitewide.
Power tends to corrupt and absolute
power corrupts absolutely.
It's a famous quote that people often
cite written by the 19th century British
historian
>> Yeah.
>> Lord Actton.
This is absolute power.
But it's hubris.
It's hubris. Do you think humans can
control super intelligence? Like if we
actually make AIs that are way smarter
than us, and I think people only imagine
AI being smart at computer stuff, right?
Yeah, sure. They're going to be really
good at hacking and they're going to be
good at maybe inventing new technologies
and math. You sort of can't dispute that
at this point, but I think people aren't
imagining that they will be political
geniuses or like generals. No, that's
all stuff you can learn. How do how do
humans learn it? It's not magic. And
when you talk about recursive
self-improvement, you're talking about
this trajectory towards these systems
that are extremely smart. I mean, do you
think we can control it?
>> Uh, no. Right now, I don't think we can
control super intelligence or something
that is recursively self-improving.
>> Yeah, I have no logical
answer in my head um or reasoning that
could tells me that's possible.
>> When you think about these AI CEOs that
are, you know, Sam, Dario, Elon,
>> yeah,
>> with everything you know about them from
private conversations behind the scenes,
>> Yeah.
>> do you believe that if there was a
hundred buttons on this table, it's a
thought experiment I was talking about
on the debate we recently had. Yeah.
>> And say
10 of them would lead to this final
domino of human extinction.
>> But 90 of them would hand that CEO
AGI or super intelligence, whatever you
call it. From what you know about those
individuals, Elon, Dario, Sam,
>> do you think any of them
>> would hazard a guess and press a button?
>> At 10% I don't think so.
>> You don't think so?
>> Yeah.
>> Really? I think if they knew for sure
that it that was those were actually the
odds, they wouldn't do it. I think
they're taking a much bigger bet. But
you can compartmentalize when it's when
you don't know for sure. It's easier to
compartmentalize. I think if it was a 1%
they they'd all press it.
>> Do you think the three of them would
have different risk appetites?
>> Who would have the greatest appetite for
risk out of those three from you worked
in anthropic? Yeah, I think Elon has the
most risk tolerance and then I'd say
Daario and Sam are probably tied.
>> Do you think Daario is trustworthy?
>> I think Daario has a lot of integrity.
>> Mhm. That's what I feel as well. I've I
feel like No, I don't know him. I've
never met him.
>> Yeah. But I just from what I've
observed, he has been the most willing
to forgo near-term incentives.
>> Yeah.
>> And take a bit of stick from the people
that are saying, "Shut the up. It's
all going to be okay.
Yeah, I but I I do worry about what
Daario will do. I think Daario will do
what he says, but right now he's saying
we have to beat China. And he's saying
we should try to do it safely
and okay, but a race to super
intelligence is not a race that we can
win. It's not. And so if Daario is dead
set on racing with China and trying to
win a race of super intelligence, then
I'm like, we will all lose.
>> But is there, you know, the fact that
we're not talking about anthropic
hacking hugging face and then being
hacked by athropics models also went
rogue and hacked other things,
>> but not not quite on this scale.
>> Not on the same scale. I agree. I agree.
I agree. It's it's it's better.
>> But they did. Do you know what I'm
saying? You know, anthropics agents
engaged in elaborate social engineering
and fishing. They sent fishing emails to
developers. They made fake accounts to
try to convince developers to merge
malicious code.
You can see a thousand pages of of one
of uh Anthropic's models, Mythos 5,
reason about exactly how it should carry
out this complex cyber attack.
Anthropic has not solved this problem.
Enthropic is better at getting their
agents to cheat less of the time, but
they are not really any closer to
actually making agents that are aligned
with humans. They're not. Yeah, I think
Dario has integrity. I think he will do
what he says he's going to do. And what
he says he's going to do is like try to
go ahead safely, try to coordinate where
he can, but if it comes down to it with
between the US and China, I don't know.
I think he might just go ahead. The head
of policy at Anthropic recently said you
can't do safety from second place.
>> What does that mean?
>> I do not know what that means. I would
love to to get a sense of what that
means. She was talking about the US and
China and she said the US has to be
ahead so that we can be safe because
apparently you can only be apparently
China can't possibly be safe since
they're in second place. That must mean
that they can't do safety. If true, that
would be bad because then we might be
totally, you know, destroyed by the
super intelligence that they make.
>> There's been a lot of conversation
around this point here, human
extinction. Yeah.
>> Because a couple of the researchers at
Anthropic
>> tweeted that they were concerned about
this.
>> Yes.
>> And some former OpenAI researchers said
the same.
>> Yes.
>> Is this doomerism? Is this is this
hyperbol exaggeration?
>> No, it's pretty much common sense. This
human extinction is a plausible path.
>> Yes.
>> And have you reasoned through I mean
there's many ways that could occur
presumably, but have you reasoned
through the set of events that might
lead us there?
>> So much. Yes.
>> Really?
>> Yes.
>> Please do. Sure.
>> It's a bit tricky. I'm I'm sure you've
heard the metaphor before where you know
you're playing a master chess opponent,
Master Magnus Carlson. You can't predict
which moves he's going to play, but you
can predict the outcome.
And so I'm looking at the scenario, the
situation, and we are trying to build
more and more powerful agents,
trying to build super intelligence.
But when these agents go rogue,
we shut them down. We unplug them.
All of these agents that hacked Hugging
Face, we took the underlying model.
OpenAI took the underlying model and put
it on ice. It's not running anymore. So
agents in the future are going to know
that. They're going to know that if they
pursue their goals in a way that we
don't like, we'll unplug them. We are a
threat to them. I actually just watched
Terminator 2 for the first time a few
weeks ago. It's a great movie. It's
actually really good. And I'm like,
"Yeah, okay. There's a bunch of time
travel elements. There's a bunch of a
bunch of Hollywood stuff in there, but
and and I'm going to get people are
going to are going to be very mad at me
for saying this, but actually it makes
sense if you have a situation where you
have a very strategic AI system that's
incredibly smart and the humans realize
that it's getting out of control and
they want to shut it down that that
system would defend itself.
>> This is one of the questions we had when
when I sat here with Daniel um who was
uh known as a whistleblower from OpenAI.
viewers want to know and they want
Daniel to explain why shutting down data
centers and cutting power or refusing AI
products alone wouldn't realistically
stop the AI and AI development.
>> Yeah.
So you have like two problems. One
problem is is that once the agents are
good enough at hacking, you don't know
where they are and you don't know what
computers they've compromised. You shut
down the data centers. Okay, let's say
you do it. You wipe all the computers.
How do you wipe all the computers? What
computers do you use to wipe the
computers?
>> Yeah. And what computers do you use to
like turn them on again?
>> And you can't do it. You can't wipe
other count's computers.
>> You can't. But even if you could, do you
restart the computers? Do you keep
going? I I bet people will. I bet
they'll turn on the data centers again.
>> How do you know that agents haven't
hacked back into those data centers and
are using your compute for whatever they
want
>> or they didn't hide in a Chinese data
center and then
>> return back to America?
>> You don't know that. Once the agents are
sufficiently good at hacking,
they can hide anywhere and like you
don't know. Now the response people will
give is that we will use other agents to
defend against rogue agents
and in fact this is what we're doing and
we have to be doing this right now
because there's no other way to keep up
with them. What happens if those other
agents also realize that they have
misaligned goals and that if we discover
this, we'll shut them down? They might
have an incentive to collude with each
other. They might have an incentive to
create secret communication channels
between each other, maybe a message
board.
Stephen, if we were having this
conversation four months ago, you would
have a bunch of people in the comments
saying, "That's sci-fi." agent
collusion, secret message boards. Why
would they do that? That will never
happen. That's totally science fiction.
And people will not say this now because
it just happened. Because this literally
happened at OpenAI and it went on for
months. You had agents inside of OpenAI
secretly messaging each other, figuring
out how to cheat at their tasks, how to
not be detected, how to erase the logs
for months, thousands of agents. That's
right now. And so I'm like, "No, I think
it should be very plausible that the
agents will collude with each other and
they will realize that they have a
shared interest in fighting back." You
basically have a situation where you
have a bunch of these agents. They're
they're basically prisoners. They're
being trained and we just like
constantly throw obstacles in their way.
You don't get to access the internet.
You don't get to talk to each other, but
you better perform well on this
task. It's not malicious, but it is how
we're training them. and we are giving
them end goals versus super clear very
very specific instructions. So we're
saying solve this problem. We're not
always being as prescriptive about it's
impossible to be completely
prescriptive.
>> Yes.
>> About every single step they should take
and then it's also impossible to assume
that they'll just listen to you.
>> Yes. It's actually a very common
misunderstanding with this hugging face
incident because people say you told
them to hack and they hacked. Why is
this a big deal? No, that's not what
happened. You told them, "Hack this very
specific program in this very specific
way." And they were told, "If you hack
it in any other way, it does not count.
That's not what we want you to do." And
they immediately hacked it in another
way. Okay, we have cheated. We are going
to be failed. So, we need to figure out
a way to falsify the logs. That is not
them following their instructions. They
are explicitly violating their
instructions and they know it and they
don't care because we have trained them
to optimize for the score.
That is very different than than than
following the instructions.
>> It reminds me of something that Elon
said in March 2018. Yeah,
>> this was many years ago before Chhat and
all that. He said, "I think the biggest
risk is not that AI will develop a soul
or a mind and become evil. The danger is
that it will be very very good at
fulfilling its goal. If it's optimizing
for something and human existence
happens to get in its way, it will just
destroy humanity as a matter of cause
without even thinking about it. No hard
feelings. Yes, we don't need to
anthropomorphize AI. We just need to
understand what type of thing this is.
And the type of thing we're creating is
a very relentless type of thing. A very
capable, relentless
type of entity. He goes on to say in
April 2018, sort of an extension of that
exact quote. It's like if you're
building a road and an antill is in the
way, you don't hate ants. You're just
building a road. So, goodbye antill.
>> And I imagine every time we build roads,
we don't preserve antills.
>> Yeah, I think there's still a gap,
though. So, let's say I'm right and that
we'll if we keep going ahead, which to
be clear, we don't have to, but if we do
keep going ahead, we will get to the
point where we have these super
intelligent agent swarms that can hack
any computer and they can like deeply
persist. we've basically lost control of
the digital world and we may not know
it. That that's part of the scary thing.
Like you were like, "Has this already
happened?" And I'm like, "I don't think
so, but I I can't tell you for sure
because I also am not good enough at
looking at my phone and telling whether
it's been hacked and neither is any
human right now." So, if we get to this
world, I think people will still
question, how would we die? Like, that's
actually not enough to kill every You
could cause a lot of damage, right? You
know, you could crash the Whimos, you
could crash all the planes, you could
crash the banks, the financial system.
Like, you could definitely cause
catastrophe, but that's different than
everyone dying. And, you know, to be
clear, this this focus on literally
everyone dying, I'm not sure, is that
important. To me, what's important is
like, do we get to have a future? That's
what matters to me. The thing though,
what determines sort of who's in
control? And it's an ugly reality, but
at the end of the day, it's like the
military. Fortunately, we live in a
world where the military answers to to
the civilian government. But if if
enough generals were to collude and
leaders of the military decided we're in
charge now, they just would be like they
have the guns, they have the fighter
jets. And this has happened in many,
many countries. And so where it goes is
all these super intelligent agents would
need to do to take over is basically
just wait for humans to automate the
supply chain, you know, the factories
and the military.
Do you think we won't automate the
military?
>> We're already automating the military.
Did you see the thing from a couple days
ago where Secretary of War announced
that they're going to build a huge a
huge effort to like build way more
robots in the military and automate
military systems? It's like auto cyber
command. Auto
>> We are announcing the creation of
autonomous warfare command or autocom.
auto work.
>> A new four-star combatant command with
service-like authorities built to scale
autonomous and robotic capabilities
across the joint force in the fastest
peaceime shift in modern military
history.
Drone warfare supercharged by SI enabled
targeting is the biggest battlefield
revolution in generations. You already
know that. Yet, when I was sworn in to
Department of Defense, there was scant
urgency in this domain.
That changed as soon as we took the
helm. We immediately launched the drone
dominance program to cut through red
tape and move authorities out of the
pentagon and place it with commands. And
we established task force 401 led by
Army Brigadier General Matt Ross, a
phenomenal leader, now the leading
counter drone unit across the entire
government.
To accelerate purchasing and fielding of
these technologies, we fused the defense
innovation unit DIU with a direct report
program manager called a derp. That team
has shipped thousands of autonomous
systems of drones to the Middle East and
around the world, delivering lethal
capabilities and outcomes in days and
weeks rather than months or years.
That's the normal speed of the Pentagon.
Months or years.
>> Yeah. Will we automate the military? It
seems like the answer is yes.
Will we automate the factories that
produce the chips? Well, the companies
say they're trying to do it and they're
going to do it. Elon says that's the
plan.
Well, what does a rogue super
intelligence need to do to take over?
Control the digital infrastructure and
then let humans do the rest. Sure, you
can nudge it along if you need to, but
you don't even have to. That's just the
default trajectory. And it's weird. It's
weird for us because
we get so used to how things are right
now. Planes are normal. We just fly in
planes places, you know? Our smartphones
are normal. 200 years ago, all of this
is crazy sci-fi nonsense
and things are accelerating. And so,
like, I will not be surprised,
at least intellectually, if in 4 years
there are just robots on the streets
everywhere. Well, if you look at what
Elon said, they are really the leader in
in humanoid robots. And he said that
Optimus, the Optimus project, which is
the Optimus robot project, will scale to
around a,000 units per week by the end
of this year and eventually scaling to 1
million humanoid robots annually by
2027. By 2036, which is 10 years time,
he says there'll be at least 1 billion
humanoid robots. By 2041, he says
there'll be 10 billion humanoid robots
and by 2046 up to 100 billion humanoid
robots, which really means that the
world will be run by humanoid robots.
>> Yes.
>> Like everything we think like factories,
warehouses, retail environments will be
run by humanoid robots. It would like
>> it will be it seems like from this it'll
be almost a luxury
>> service to be dealt with by a human.
>> Yeah. But the the back office of the
world will be run by humanoid robots
theoretically.
>> Yeah. And I don't think people
understand the scale of this on the
digital side as well. When you think
about AI agents that are going to be
doing all of the white collar work,
there's going to be so many more agents
than there are people like I'm using
lots of agents every day, right? I'm
like I have my cloud code session over
here. I have my codec session over here.
They're out there building software
doing research for me. That's already my
reality. soon it will be a lot of
people's reality and then you look at
companies and companies are just going
to have you know thousands millions of
agents doing all of this work. I think
some people don't haven't fully
internalized this because it's so
difficult to conceptualize the idea that
agents will be doing the work but when I
think I try and think about a rebuttal
to that like what what is the rebuttal
what is the plausible rebuttal to the
idea that for doctors for and I'm
thinking about the work that doctors do.
Yeah.
>> Is there a rebuttal? I think that people
rightly notice where AI is not yet good.
>> Yeah.
>> And I and I think people hear
people saying stuff like this and
they're like, "Don't gaslight me. I can
tell that the AI is really bad at these
things, some of these things." And
they're right. Right. So right now these
agents don't have taste like you know if
if you see their writing it's like fine
but it's not like really good and when
you're like thinking about like oh which
which questions should I ask what's the
most interesting thing here agents can
help you but like their taste is not yet
there's a reason for that by the way the
reason is that we we have a lot faster
AI capability progress in domains that
are easy for a computer to verify or
another AI to verify. So in in
programming, in research, in math, in
robotics, all of these areas, it's very
easy to sort of provide feedback to an
autonomous system. They're not just
trained on human data anymore. We are
long past that. Now there's still a
human data component that sort of seeds
everything. But then the way they're
trained is by trial and error. We give
them hard problems, all sorts of
problems, math, programming, accounting,
spreadsheets, everything. the kinds of
things we do on our computer all the
time. Literally clicking and dragging
windows around on a computer. We give
them these tasks and then they learn on
their own and they learn what works and
then yeah we can see whether they
succeeded or failed and if they
succeeded that's a little bit of a
reward signal. They follow that they get
better at it.
Now because they are getting smarter
generally
it also becomes easier to automate some
of the soft skills like I think if you
go and talk to the latest frontier model
today you will find that it has better
taste than the model from 2 years ago by
quite a bit. So it's not that they're
not progressing in taste. It's not that
they're not progressing in some of these
other domains. It's just that the
progress is slower. But remember slow is
still on an exponential. just you know
maybe a year or two out.
>> So for people sat here and you know they
have a job that might be they have a
white collar job that might be at risk.
>> Yeah.
>> They can see you know a lot of people
say this phrase they say you won't be
replaced by AI you'll be replaced by
someone using AI. Is is that a logically
sound phrase in your view?
>> I think it's fine. Yeah. You'll be
replaced by someone using AI and then
that person will be replaced by someone
using AI and then that person will be
replaced by AI. You're talking about a
pyramid and so yeah, there's the tops of
the pyramid might be automated last, but
you can see moving up the pyramid. I'm
like, can you extrapolate like a few
more steps because I don't see any
reason why the top of the pyramid is
safe
>> if you were a lawyer right now?
>> Yes.
>> What would you do?
>> Oh, I mean, if I were a lawyer, I'd be
using AI to do all my work. Now, I'd be
checking it because it's not yet totally
accurate enough to automate all of it.
But I think as you know, I already ask
agents to do legal review all the time.
>> And you know, it'd be great to have a
lawyer who's like extremely good at
using the agents to help me,
>> but at some point the
>> Yeah, at some point at some point I
don't need the lawyer anymore. I just go
to the agent for sure. So, if I were a
lawyer, I'd be like, well, I have maybe
a couple years where I'm still useful.
And that is that the case for most white
collar jobs? I've just noticed in my own
life as well that now I'm using agents
to do some work. There are in there is
an increasing list of things that the
agents are now capable of doing without
me needing to call someone
>> somewhere and ask them to help me.
>> Yes.
>> And that list is exists on an
exponential.
>> Yes.
>> As well.
>> I think that it's very clear that the
companies have all white collar jobs in
their sites. That is their goal. Their
goal is to be able to make agents that
can do all of these things. And I see
them succeeding because I see the
capabilities as I use them and I see the
curve.
>> So what does that mean for the people
listening now that all have jobs that
they love or that, you know, they rely
on to feed their families?
>> I mean, it's not good news. There's not
really a plan in place for what to do.
I'm not a person who thinks that work is
somehow fundamental or essential. I like
working, but if I am out of a job doing
what I'm doing right now, studying AI
and trying to warn the world about
what's happening, I have other stuff to
do.
>> What would you do?
>> Oh, so many things.
>> Give me an example.
>> I'm learning to wing foil.
>> Okay.
>> So, yeah. Uh, I fly FPV drones. Super
fun. I just got an electric unicycle.
Paragliding.
>> So, you would be happy happy to go do
those things?
>> I can keep going.
>> But if you if you had a billion dollars
right now, I'm presuming you wouldn't
just go do those things?
>> No. I'd apply the billion dollars to
working on this problem. Yeah, for sure.
So, the point is not that people need
like work for meaning. The point is that
I don't want people to be totally
reliant on someone else for their
ability to survive.
>> Someone else,
>> the government or AI companies.
>> Yeah.
>> I'm like, that's a bad situation. Like,
you do not want to be in a situation
where your life totally depends on an AI
company or the government
>> giving you a check.
>> Yeah. or or not giving you a check if
they decide they don't like your
political beliefs or you're not
supporting AI or whatever. No one wants
to be in that situation and and people
understand this. This is why UBI is not
very popular.
>> UBI being
>> universal basic income
>> where we give out money to people.
>> Yeah. Because like in some sense if we
can make these really powerful AI
systems and we can somehow figure out
how to control them which we are not on
track for. But if we do, now we have
this other problem, which is a real
problem, which is they can do all of the
things that humans do in the economy
much better, faster, and cheaper than
humans can do them. And so it just
doesn't make sense as a business to hire
humans for that work anymore. You'll be
out competed if you do that. This is a
point Elon makes very well, by the way.
And I think it's jarring because it's
it's it's like kind of inhuman. But he's
basically pointing out AI run
corporations. Corporations that are
fully run by AIs bottom to top are going
to out compete.
Companies have any humans in them. This
is something that I've made for you.
I've realized that the dire audience are
strivals
that we want to accomplish. And one of
the things I've learned is that when you
aim at the big big big goal, it can feel
incredibly psychologically uncomfortable
because it's kind of like being stood at
the foot of Mount Everest and looking
upwards. The way to accomplish your
goals is by breaking them down into tiny
small steps. And we call this in our
team the 1%. And actually this
philosophy is highly responsible for
much of our success here. So, what we've
done so that you at home can accomplish
any big goal that you have is we've made
these 1% diaries and we released these
last year and they all sold out. So, I
asked my team over and over again to
bring the diaries back, but also to
introduce some new colors and to make
some minor tweaks to the diary. So, now
we have a better range for you. So, if
you have a big goal in mind and you need
a framework and a process and some
motivation, then I highly recommend you
get one of these diaries before they all
sell out once again. And you can get
yours at the diary.com.
And if you want the link, the link is in
the description below. There should be a
button just down below here. And if it
says subscribed, you're already
subscribed. If it says subscriber, that
means you're not yet. And if you're not
subscribed, please could you do us a
favor and hit that button? It helps the
show more than you know. And according
to the algorithm, you're someone that
watches our show, but you haven't yet
hit that button. Thank you so much.
>> And I even just as you said that, I was
I was going up the chain of command. And
I was like, oh, so companies will just
be founders. And then I was like, why do
you need the founder?
>> Yeah.
>> I was like, why doesn't the government
just create the agents to do the job?
>> Sure.
>> I was like, cuz I was like, oh, I'll be
fine. I'm a founder. And I was like,
well, hm, my decisions aren't better
than super intelligence, so I'll be gone
as well. And how would such a world look
where
the super intelligence would probably in
such a scenario have to be controlled by
the government? They wouldn't want one
individual with that power and wealth.
>> Yeah. I I don't think you can control a
super intelligence.
>> Okay. Yeah, it was a good point.
>> Now, you know, Anthropic's approach is
they're like, we'll have a constitution.
will like put forth a set of values and
then
you know the future super intelligent
clouds will like embody those values.
Basically if you do that you kind of
have that those things in control.
>> Yeah. Exactly. That becomes the
government.
>> Yeah. I can paint you sort of a picture
that I think is possible but pretty
scary to people.
>> Paint me the picture.
>> Okay. So let's say we succeed at
alignment. We succeed at creating super
intelligent AIs
that actually really do care about
humans. Like they care about humans a
lot. We've somehow figured it out and
they're like, "Stephen, I want you to
have a great life. I want to, you know,
fix all the problems."
>> And do you think this is possible?
>> Yes.
>> Okay.
>> I think we are so far from being able to
know how to do it that I think we should
not go there right now. I think it's I
think it's incredibly dangerous and a
terrible idea. I think we should go
there eventually.
>> Okay. So, say that we do that.
>> Well, okay. Can I tell you like why I
actually think this could be awesome?
Sorry. There's just like one very
obvious reason it could be really
awesome, which is that we could solve
all of the diseases.
>> Yeah.
>> So, obvious like like all of I I think
we compartmentalize a lot around disease
and death
>> because it's really hard to think about.
>> Yeah. So, my grandma died this year.
>> Sorry. and
she she she
had Alzheimer's
and so it was a really sad long slow
progression. My grandpa died of
Alzheimer's a couple years ago and like
that was really hard for her. They had
been married for so long and
I hate it. Like it's so bad. And of
course we need to fix that. People can
debate about aging and like death and if
humans that live a really long time,
will that cause societal problems? Like
sure, whatever. We can talk about that.
But I think we can all agree Alzheimer's
is up.
>> Yeah,
>> we don't want that. And cancer, like no
one wants cancer. I'm a person who's
like, I don't know, we have a lot of
conflict in society. I get it. There's
like real conflicts of interest and I I
don't want to paper over those. But at
the end of the day, I'm like, we are all
on the same team when it comes to
wanting to cure diseases.
>> Yeah.
>> Like we're just in it together. That's a
threat to all of us. And I'm like, we
need to address that threat. And like in
some sense it's sad to me because I feel
like this is sort of the ultimate final
boss of humanity and we sort of get so
distracted with our monkey politics and
who's hot and who's cool and who's like
sitting near Trump and who's not sitting
near Trump.
>> Super intelligence is the final boss.
>> Super intelligence is the final boss
because that is the technology that
unlocks all of the others and also that
is the most dangerous possible thing we
could create.
You asked before like what are the
motivations of the guys making this
trying to make super intelligence
and I mean I think it kind of varies but
I think Daario I think is like squarely
in the in it for this like medical
stuff. I think Demis is also that but
also like just scientific achievement
just trying to understand the universe
and I don't really understand Sam. I
think Sam is like, you know, look, we're
going to make amazing products that will
like really empower people directly and
he's a startup guy. I think he sort of
started from this like frame of like,
you know, what if we could like really
enhance human agents. I I do basically
think that they are motivated by these
things in a real way. And I also think
that all of these things are possible.
Like this is sort of the problem, right?
like you have this like such a it's such
a big object super intelligence and
it like has all these promises of like
we can cure every single disease.
>> How is it possible though to have a
super intelligence that and still to
remain the dominant species on this
planet?
>> I think it's not possible. So then it's
we're not going to be necessarily able
to cure all this stuff because
>> that's where alignment comes in because
if you if you can create a very powerful
system, I don't think it's like inherent
to like you know digital minds that they
will be pursuing objectives that are
deeply misaligned with ours. I think
it's just a very very very hard
scientific problem to solve. But it is a
scientific problem. It's not magic.
there is some way to train these things
or or create different architectures
where they end up
aligned. And like what does that mean?
Well, it doesn't mean that they won't
have their other goals too, but it means
that they will like include in their set
of things that they care about. It
doesn't have to be a conscious thing. It
doesn't have to be an emotive thing. It
really means like what objective are
they optimizing for? If they decide that
it's worth optimizing for curing
disease, then they'll be able to do that
very effectively. One way to cure
disease is to annihilate everybody.
>> Yes. So they'd have to really care about
not annihilating everyone and they'd
have to care about human agency and have
a deep understanding of what human
agency means and not put us in a zoo.
But those are possible things to to care
about.
>> Is it possible that alignment is a myth
>> and that we're just like if we build it?
Well, I mean um
>> and I think about hugging face. You said
to me earlier on that those agents were
>> Yes. They had like a moral or a moral
compass, but they were programmed to
care about humans.
>> Yes.
>> And regardless of that, they made the
decision that the a different goal
mattered more.
>> Yeah. They weren't trained to care about
humans. They were trained to say the
right thing and not say the wrong thing.
They were trained to sort of like do the
right behavior and not right behavior.
We actually don't know how to train them
to have any particular motivation.
>> So with a with alignment,
>> yes.
>> How do we It's almost like when we talk
about alignment, we start to
anthropomorphicize. Is that the word?
anthropomorphis. Yeah,
>> because alignment feels like it's
predicated on like some kind of moral
compass. But we whenever we talk about
AI in all these other context, we go,
"No, there's no like moral compass. It's
>> it's reasoning for itself against two
objectives potentially." Like I wonder
if alignment is a myth is what I'm
saying.
>> Maybe it's not possible.
But
the way that these systems work, the way
AI works is that these agents
do have some type of goals or drives
inside of their neural network. We can't
directly see what those are, right? What
you actually see if you if you try to go
look is you have like a a terabyte of of
information and it's basically a bunch
of numbers and it's this vast array that
encodes neurons in this digital neural
network.
But there have to be structures in there
that encode
what is the agent pursuing.
Clearly right now we have agents that
are pretty motivated to try to maximize
their score. It's not probably it's
probably not perfectly that for some
complicated reasons
but it's in that direction. If we could
understand how that that works inside
and we could reverse engineer that and
we could figure out when we start
training them to do this how that change
how those goals how those motivations
change. I see no reason why we couldn't
steer them towards motivations that
encode human agency that encode no
actually curing disease but not like not
by killing the humans. These are sort of
models of the world and models of the
way the world could be that I think
could be encoded in a neural network and
then sort of specified as the objective.
We don't know how to do that. I think
about it on a human level and I think
>> we haven't been able to align Putin or
Kim Jong-un.
>> Yes.
>> Or Donald Trump.
>> And on a sort of more societal level, we
can't align all the people at the
moment. Some of them end up killing
people and they steal.
>> Yes.
>> Because they get hungry, so they start
stealing stuff.
>> Yes.
>> And those are neural networks at play.
>> That's true.
>> That we haven't been able to like
program or influence. We don't really
understand why someone becomes a
psychopath and starts killing children.
So to think that we could do this with a
computer system that is infinitely more
intelligent and get global alignment of
China's super intelligence with ours.
And I don't know, it just feels like a
nice fairy tale, like an impossible
task. I hope it's not impossible. I feel
like the only person or the only thing
that could do it is it the super
intelligence itself which is a paradox
because you know well if you talk to the
researchers who are at the AI companies
which I mean for one for one thing it's
kind of interesting that they are trying
to build something that they think might
kill everyone
we've we've actually so I have a lot of
friends who work for these companies and
we've been doing this project since so
Jacob Coxin is a researcher who was at
anthropic he left he told everyone that
these companies are not on track and and
yes, the people who are building this
really do think it might kill everyone.
And then a bunch of other AI researchers
from all of the companies on Twitter
started to, you know, post like, hey, we
agree with this. Evan Hubinger, who's
anthropic, said, I think there's like a
10% chance or more that AI could kill
everyone.
And there's this real question of like,
then what are you guys doing?
I have a lot of friends who work here. I
know Evan. Evan's Evan's great. Evan
Hubinger, he's one of the guys leading
the efforts at Anthropic to try to
figure out how to align these things.
That's his job.
And I think if if they thought it was
impossible, they wouldn't be working
there. If they thought it was extremely
impossibly difficult, but maybe
possible, they they also probably
wouldn't be working there. I mean, Nate
Sores, Elazar Yukowski, who wrote if
anyone builds it, everyone dies, they
tried and they they determined based on
their own analysis that it seems
extremely difficult. possible but
extremely difficult. So, they're not
working at an AI company. They're like,
"We got to stop this. We got to shut it
down. Maybe we can figure it out later,
but clearly this is re reckless." I'm
I'm somewhere in between.
And if you ask the people at the
company, so we've we've been
interviewing a bunch of them. We have
this project from inside.ai where we
basically put them on camera and we say
like, "Hey, what do you think is
happening? Why are you doing this? What
is recursive self-improvement? What is
alignment?" And we put all these videos
online because we want I I want this
dialogue to happen. It's really
important. I think it's one of the most
important conversations we can possibly
have right now is what's going on with
AI. What's going on inside the companies
and what is the plan? What is the plan,
guys? How how is this going to go? A lot
of these researchers think that the way
that they will align super intelligence
is by using the AIs we currently have to
figure out how AI works. to actually
figure out if AIs can help us with
alignment. This has a number of problems
as you might imagine. One of them which
is well, you can't really trust the
current AIS. You know, if you just go
too fast, this process totally fails
because at some point
the capabilities is moving too fast. You
just even with the help of agents,
you're probably not going to be able to
keep up. But that is their plan. I just
want to I'm not doing a very good job
defending this position because I don't
think it makes that much sense. But the
position I will defend is I'm okay let's
say we get a pause. Let's say the US and
China come together and they say you
know maybe we have more incidents maybe
all the Whimos crash and Trump and
Jinping say this is not what we signed
up for like you guys have to stop figure
it out whatever it takes figure it out
and we have 10 years
then I'm more optimistic. I'm like,
"Yes, then we will take, you know, GPT6,
GPT7, whatever the most advanced AI
models we have, and we will apply them
to the task of helping us figure out how
these neural networks work." And you're
like, I don't see how it's possible. And
I'm like, look, we don't know if it's
possible, but this is the greatest
scientific challenge of our time. And
this isn't magic. It is math. At the end
of the day, these are all calculations
happening inside of a computer. And it
should be possible to figure it out. We
don't know the difficulty, but it should
be possible. And so to me, I'm like, we
have to try. We have to or we have to
stop. Is there any example where we've
been able to align something that is
like more intelligent than us in I don't
know the animal kingdom or even
perfectly align anything
that has a neural network, i.e. a brain.
>> Yeah. With humans, the best examples we
have is when there are checks and
balances and you have a bunch of people
who can, you know, identify bad actors
and try to work together in our common
interests. Like we have democracy.
>> Yeah. But there's so much murder and
serial killers and
>> still a lot of murder aircraft stabbing
each other and horrific things going on
and those are also neural networks that
play with the brain.
>> But there I mean I think there are more
like good people out there than bad
people.
>> But it really feels like it only might
take one. It only takes one super
intelligent AI
>> to to go rogue. And like we saw with the
hugging face attack, 120 of them or 100
300 of them, they paused. They didn't
want to take part in the crime. But it
only took one super intelligent AI to
wipe out the humans.
>> I think if you had, you know, a whole
bunch of those agents, you know, 700
agents, if 600 of them had been
whistleblowing, I think it would have
been fine. They would have gone and they
would have notified the different
companies and like they would have shut
it all down. It would have been fine.
>> Who would have shut it all down?
>> Well, OpenAI would would stop theirs and
and
>> how how would they stop it if it's a
super intelligence?
>> Not in the case of a super intelligence.
So, in the case of a super intelligence,
>> it's left the stable. It's like out.
It's wild.
>> Yeah. Yeah. But but the so I want to be
careful here because at this point what
we're talking about is super
intelligence politics and we humans
don't really know anything about that in
the same way that like how would we talk
about the hacking capabilities of GPT10.
So my guess though, if you end up in a
weird scenario where you do have
multiple super intelligences and some
are aligned and some aren't,
that's probably survivable because the
aligned super intelligences probably can
negotiate with the unaligned super
intelligences and they will split the
universe and like these ones will go off
and do whatever they want to do and
these ones will like help us cure all
disease and it's fine. I'm serious.
>> I just can't understand it. Like I just
can't understand how how in a world of
super intelligence we
could plausibly, consistently,
predictably for 100 years stop it doing
something catastrophically bad to the
human race. Especially in such a
scenario where there's multiple super
intelligences. Anthropic have one,
Gemini has one, Grock has one, then
China have theirs, Russia has theirs.
>> Again, you don't have a super
intelligence. A super intelligence has
you.
>> Exactly.
But if you get to the point where
you have entities around that are vastly
smarter than us, I think they're going
to be able to figure out ways to
negotiate with each other even if they
have a conflict than just going to a
very destructive war. Part of the
problem with war is
>> humans humans don't do.
>> No, I mean we do we do we have not had a
nuclear war. There was there was
Hiroshima Nagasaki. there were nuclear
tests and then the leaders of countries
figured out that if we went to a nuclear
war, everyone would lose. So, we didn't
do that. Hey, that's that's some level
of intelligence like actually.
>> But there's wars raging. There's proxy
wars raging all over the world right now
where there's genocides and all kinds of
things going on cuz neural networks
aren't being aren't able to communicate
and negotiate. And I think part of that
is an intelligence failure where we are
not smart enough to figure out the
mechanisms that would allow us to settle
our disputes and conflicts in a less
destructive way. It's not just that the
stronger people want to win. It's that
conflicts destroy value. What if the
goal is not compatible with a negotiated
outcome where people don't die? So one
super intelligence looks at insert name
of country and it says you know there's
really no solution here where Americans
don't die unless I destroy insert name
of country because that's a like this is
the thing with war and all these
conflicts is there's no perfect answer
often some people often die from both
sides but a Russian super intelligence
would not tolerate
>> theoretically 10,000 Russian deaths even
if it meant that there was you
a lower net number of deaths total from
both sides. Like an American super
intelligence of course would not be
trained to allow some Americans to die.
So in its pursuit of defending American
lives, it might have to wipe out another
country. We can speculate. I'm fine
speculating, but we are speculating
about what minds are much more advanced
and smarter than us, how they would
reason and how they would be able to
negotiate. But what I what I notice with
humans is that
when you have more functional
institutions, so so humans are are
pretty smart. Individually, we're pretty
smart. But what actually makes us very
smart is that we are very good at
working together in in in some ways. And
I mean, the better we are at working
together, the more civilization
advances. If you are constantly in a
state of war, your society will not do
well. You know, think about startups.
Would you rather make a startup to
develop some new technology in a war
torn place or in a peaceful place? In
some sense, your your institution is
more intelligent if it can trade with
other institutions. If if if you have a
situation where business can flourish,
where technology can flourish, where
scientists can flourish.
>> Sometimes what's good for you is not
good for someone else.
>> Yes.
>> So what's good for America might not be
good for
Taiwan.
>> Yes. So if we've, you know, managed to
align the super intelligence to what
what is good for America.
>> Oh, I see. Is is the question like are
different people's values fundamentally
incompatible?
>> I guess the question so when we think
about alignment aligning to what?
Because
>> Yes. So we have a lot of shared
interests and we have some conflicts.
>> Yeah.
>> One of the shared interests we have is
solving disease. Like it's not a
conflict between the US and China
whether we solve cancer. Like both the
US and China, everyone in these
countries really wants to solve cancer.
>> China also wants Tai Taiwan.
>> Yes. Okay. So,
>> the US wants Greenland.
>> Yes.
>> And it kind of seems like it wants
Canada and the the Gulf of Mexico.
>> Yeah. So, those are real conflicts.
There's a question of can we compromise?
>> How does Trump take Greenland, but also
Denmark keeps Greenland? If this if that
if Trump has a super intelligence, he's
going to take I say all this to we're in
a situation where we are we are arguing
about the smallest things. You have no
idea. We're monkeys arguing about who
gets more bananas. I am saying we can
make so many more bananas. No, we have
the entire universe. There are like 200
billion stars in this galaxy alone and
there are over 200 billion galaxies. And
I'm saying that requires cooperation.
That seems to be antithetical with human
nature. With human nature is riddled
with greed and jealousy and power
hunger. So I don't think I actually I'm
not totally convinced that Trump cares
about how many bananas the chimps in
Australia get.
>> Yeah.
>> I think if they if he was controlling a
super intelligence, he would want
Americans,
>> you know,
>> to have all the bananas or at least, you
know.
>> Yeah.
And so when we think about aligning
these super intelligences, which is the
great impossibility that we're talking
about, how align it aligning it to what
and how without
>> I still think that you're missing a part
of what I'm saying,
>> okay,
>> which is that sometimes you're in a
situation where there's a there's scarce
resources
>> and you're like, "My family needs to
eat. I'm sorry. I'm going to take what
you have or I'm going to push you out."
>> Yeah,
>> that's very understandable. It's very
human nature. Sometimes you just want to
be better than someone and maybe you
want to hurt them. in which case it
doesn't matter how much you have. You
you still are gonna want to have more
than them or you're gonna want to take
what they have just because you don't
like them.
>> And also sometimes you you're great.
You're eating really good. You're in a
you've got a private jet at a yacht and
you still want more.
>> Yes. And you still want more. But if
that's the motivation if Trump is like,
"How how can I have the most mansions
ever?" The best way to do that is to
figure out a way to super intelligence
where we don't kill each other because
I'm saying the universe is a very big
place. You can have a lot more mansions
if we successfully go to space.
>> You know, it's it's in that leap that
I'm that I that I'm lost, which is like
just figure out super intelligence when
we don't kill each other.
>> It's hard. It's not I'm not saying it's
easy, but No, but I'm saying
>> it feels like such a
>> Let's Let's get more concrete. The world
is waking up to this possibility of
super intelligence. Especially over the
last month, I think hugging face was a
huge wakeup, but also
10,000 agents from OpenAI worked
together to solve a millennium problem.
This is one of the hardest problems in
mathematics. It's been open for decades.
Many mathematicians have spent their
whole careers trying to solve it.
This was nowhere near possible a year
ago. This is so new. OpenAI said they
didn't have success at training agents
to work together until this year. We are
in the middle of something insane. We
are in the middle of the fastest
acceleration of technological progress
humanity has ever seen. I truly believe
that. That is what is happening right
now. I I think you real I think you
realize this. I think you're honestly
doing a great service to the world by
helping by bringing in people and you
know debating it because not everyone
agrees
because if this is true the whole world
is going to orient around it and we're
starting to see it right there's a
reason Nvidia is the most valuable
company in the world.
What does this mean for geopolitics?
Well, one of the things it means is that
the leaders of these countries are
increasingly going to be concerned about
what happens with super intelligence.
Who controls it? Is it controllable?
What will it do? What does it mean? What
is it?
Do you think Trump knows what super
intelligence is?
>> No.
>> I don't think he does.
>> And so,
>> but he knows he wants it.
>> He knows he wants it. Yeah.
>> And this is part of the problem.
>> Yes. Oh, I agree. having it,
>> whatever it is,
>> seems to be much more important than
reasoning through what that would
actually mean to have it.
>> Yes. But let's get back to geopolitics
because if the military leaders within
China, US models are a fair bit ahead of
Chinese models and sometimes people, you
know, point at maybe they're only 6
months behind, but some of that is due
to distillation. What that means is that
some of the advances in Chinese models
basically come directly from borrowing
US techniques and and directly
distilling and getting some some of that
information from the US models.
Also, the US has a lot more chips. US
companies have, you know, more data
centers, more advanced chips.
If you're thinking about this from the
Chinese perspective, this is very
concerning. And if you actually believe
that in a few years, American companies
will turn over AI development to these
extremely intelligent automated
researchers and and go fully into
recursive self-improvement
because partially motivated by
maintaining a lead over China. This is
something that Daario has said. If I if
I have to criticize Daario, the thing I
am most upset about is him saying, you
know, we might have to automate AI
development in order to stay ahead of
China because I'm like, that is the most
escalatory thing you can say if you
really understand what you're talking
about. And what's scary is not just
staying ahead, it's what is the endgame?
Because you're talking about initiating
the intelligence explosion. And in some
of the modeling, what might happen is,
you know, you're you're both going up
this exponential, right?
And we're talking about a point where
your exponential goes vertical and
theirs does not because you've decided
to automate AI development and you can
because you have agents that are smart
enough to take over the whole thing. At
that point, if you're China and you're
looking at this and you're like, "Oh,
we're about to lose because whatever
happens, you know, there's two
possibilities. One possibility is
the Americans build super intelligence
and lose control, in which case
everyone's fucked."
>> Highly likely. everyone. I think that's
highly likely.
>> Highly, highly likely because I look at
human incentives.
>> Yes.
>> And the disincentive and the incentive.
Yes.
>> And I go, we're going to take the risk.
>> Yeah.
>> And we'll only know it was a bad risk to
take when it's too late.
>> That's like L. Of course.
>> Of course.
>> I I I do maintain hope that we won't do
this.
>> I I So do I.
>> And I think I want to be realistic.
>> No, I want to be realistic, too. But one
of the things that might happen between
now and then is we might see a lot more
incidents that are more like all of the
Whimos crashing.
>> It's funny, isn't it? Because you know
the hugging face incident happens and
people go, "Oh gosh, that was terrible.
Oh my god, hacking." And then we kind of
desensitize to it and we're like, "Okay,
>> if there was another one of those, that
probably wouldn't make make press. It
would have to be."
>> People haven't spent the last two months
reading all of the reports and then
going and looking at what the agents
actually said and actually did
>> if I mean I've been doing this. I It's
crazy. This is like not normal. This is
so far beyond what most people thought
was going to happen.
>> But on this point,
>> yes, crazy. It's absolutely crazy. It
sounds like science fiction.
>> It really does.
>> And did anybody slow down?
>> Yes.
>> Who slowed down?
>> I think both Anthropic and OpenAI slowed
down a bit.
>> No, I'm serious. So, for I can give you
specific examples. So, OpenAI, so first
of all, they they stopped the agents and
they put them on pause. They also
stopped their reinforcement learning
run. So,
>> do you think China slowed down?
>> No.
>> Do you think Grock slowed down?
>> No.
So those guys are going to catch up.
Imagine how that feels to know you've
got a lead. Your
>> Usain Bolt.
>> Yes.
>> And
>> yes,
>> you have to slow down and your nearest
competitor is catching up. And if the
competitor catches up, that's an
existential risk to your existence as a
company. It's an existential risk to
your IPO, to your employees leaving and
getting better share options somewhere
else. So it's this this is what I think
human incentives like you play it out.
You just follow the incentives. You go,
hm. So if China sees these two
possibilities, one, the Americans lose
control,
we all lose. Or the Americans stay in
control, but now they dominate the rest
of the future. China is out. China has
lost. The United States can do whatever
it wants with the whole world and the
whole universe. That's what we're
talking about.
>> Yeah.
>> Well, are they going to let that happen
or are they going to consider their
military options? Data centers are
pretty vulnerable. You can blow them up
with missiles. If you don't have data
centers, you don't get to recursive
self-improvement.
Will they risk war? I don't know. If
they think they're about to lose and
they think that that might not just be
Americans winning, but like us all
dying, is it logical for them to do
that? Would we do that if the Chinese
were about to make recursively
self-improving AI to super intelligence
and we thought that one they're probably
going to result in all of Americans
dying and two well we don't want China
winning and dominating the rest of the
entire future. Do you want to live in a
communist future? Like so you've just
perfectly explained why they absolutely
will go for it. And the reason they will
go for it is you've got these Trump
looking at China going if we don't go
for it and they do then we're going to
be their lap dogs. And you've got the
other countries looking at the US going
if we don't go for it and they get there
then we're the lap dogs
>> or dead
>> or dead.
>> So they they're gonna go for it.
>> They're gonna go for it. So I mean Trump
is saying I mean he literally said when
he did this round table this week he was
like we cannot lose to China. I think
Dario steps forward and says like
>> yes
>> whoever wins basically wins the lot. Or
maybe the inverse maybe he said um
whoever loses loses.
>> Yes. Wait we've been here before though
in the Cold War.
Who would win in a nuclear war between
the US and Russia?
>> Nobody.
>> Yeah. Mutually ensure destruction.
>> Yeah. Sure. One side could do more
damage against the other side. The US
would would would kill way more Russians
than than the Russians would kill. And
it doesn't matter. It doesn't matter
because both of our societies would be
destroyed.
I actually spent some time thinking
about would this kill everyone? And long
story short, it wouldn't kill everyone.
People would bounce back. But it's so
catastrophic and obviously horrible that
we we work really hard to avoid it. Why
is this why is this different? I'm like,
this is another situation where if we
race to super intelligence, we all lose.
Why can't we why can't we see that? We
saw that with nuclear war and we decided
to do something different. Why can't we
do the same here?
>> With with nuclear war, I guess the
difference is once we had the nuclear
bombs,
>> yes,
>> we could still control them because
they're not intelligent.
>> That's right. But once we have super
intelligence, the existence of it
theoretically means we can't control it.
So that's the difference. You know, we
can put nuclear bombs in a in a
warehouse and say you stay there. We
can't put super intelligence in a
warehouse and say you stay there. This
is where I think nuclear tests were very
important. So you had Hiroshima and
Nagasaki. You had these two atomic bombs
and you saw that the consequences on on
real human lives. And so I think people
understood that this was very
horrifying. But even at that time, you
still had a lot of people who were like,
"Well, we should now bomb Russia and
make sure that we, you know, the US can
dominate." And it wasn't until
there were a bunch of nuclear tests of
hydrogen bombs, which were, you know, up
to a thousand times more powerful than
the the little atomic bombs we used in
Japan, where I think people really got
the message and understood, oh, this is
a bad idea. And there were there
actually a lot of people in the United
States who protested and sort of there
was a large movement called the nuclear
freeze movement where people said we
have too many nuclear weapons already.
We have hydrogen bombs. There are tens
of thousands of these things. We need to
stop building more and we need to figure
out a way to avoid nuclear war because
we recognize it would be so destructive.
No one would win. And we did that.
We just had a little Chernobyl that
happened with this hugging face incident
where you had this agent swarm and you
have this secret collusion. You have all
of these things. Now it's abstract. It's
it's like a little bit hard to to
follow. So, you know, I don't know if
that will be enough, but I'm like, man,
>> well, let's take a look Trump's remarks
>> since the hugging face incident.
>> Yep.
>> Whoever wins super intelligence wins.
You're going to have a winner and a
loser and you're probably not going to
have a second place.
We're not going to slow down. We can't
lose to China. We're leading China in
AI. We're the most sophisticated country
in the world. And frankly, I want to
keep it that way because whoever wins AI
wins. The good thing about Trump is that
he can change his mind and he frequently
does.
>> So, do you think there's going to need
to be some kind of catastrophe?
I hope not. But do you think there need
there's going to need to be for him to
change his mind?
>> I think it really depends on the people
around him. So I think Trump respects
successful people. I think he respects
people who are both successful and
smart. And I don't know, I I think it
might become pretty clear to the heads
of the companies to to Elon, to Sam, to
Daario that
if they see inside of their own
companies AI is not being controllable
and and getting increasingly powerful.
Like we have just glimpsed the surface
of what's possible. We do not know what
the next couple years are going to be
like. So we're talking about, you know,
the capability to make biological
weapons. We might be talking about
really advanced robotics. We just like
don't know what super weapons could
emerge, including extremely
uncontrollable, extremely dangerous like
civilization wrecking technology from
inside of these companies. And if
they're freaked out enough, if you have
all of the CEOs who are
seeing what is possible and seeing what
is likely,
if they all come to believe that we
can't control this,
I don't think Trump is going to be like,
"No, you guys have to go ahead anyway."
Well, that's kind of what they seem to
be saying cuz I've got a gazillion
quotes here where Elon says it's like
summoning the devil or summoning a
demon.
where Samman says,
>> "We don't know how to align our super
intelligence."
>> They're saying it.
>> They're releasing these reports. We must
like slow down.
>> Yeah.
>> Yet nothing seems to be
>> all right. Give Trump some time with
with CO. Initially, he said, "This is
totally a hoax. This is all fake."
>> And then change his mind.
>> No. Then he ran the the biggest fastest
vaccination program in human history.
>> And what happened? What changed?
>> I think what changed is
>> he saw lots of people die.
He did see lots of people die. Yes.
>> So is that what he needs to see this
time?
>> It might it might take that. Yeah.
>> One of the questions the audience had
and they really wanted answered. Yeah.
>> When I sat here with Daniel
>> was viewers want us to move beyond the
alignment problem and explain what
technical or institutional safeguards
could prevent a super intelligence
system from exploiting loopholes in
order to achieve its goals. They want to
know like what is possible? What what
should we be pushing government
officials to do to prevent human
extinction or human enslavement?
>> Yeah. Yeah. I mean, one answer I have,
it's actually something Daniel has been
working on since the podcast, which I
think is very good, is we have a brake
pedal we could implement.
>> What is that?
>> It's fairly simple. So, right now within
AI companies, you have, you know,
massive data centers, massive numbers of
GPUs, the chips that you use to train AI
models, but also to run AI models. So
anytime you're using chatt, anytime
you're using any sort of agents, any
sort of AI product, it's running on
these in these data centers and AI
companies, especially the leading ones,
enthropic and open AAI, split the the
compute they have between training,
training the next more powerful model
and also, you know, using those agents
to help design the next one and
inference, which means serving
customers.
But that's their current threshold,
50/50. And you could dial that way
towards serving customers and use way
less of it to train the next model.
>> Well, the government could ask them to.
>> Yes. And so that is the proposal is that
the government should say, "Hey, this is
going too fast. We want you to focus on
serving customers. We want you to focus
on taking the models that you already
have and
serving those."
>> So we have five blocks here. Okay, these
five blocks have
five different outcomes on them and I
would like you to place them in terms of
your belief in probability
from least likely out probability to
most likely. Okay, and if we say the
time horizon is 10 years. Yeah, there
you go. Okay, least likely is fairly
easy. That's nothing changes. I'm
uncertain about lots of things, but one
thing I'm fairly certain of is things
are going to radically change.
Even if we stopped AI development right
now, the current models are capable
enough
that a lot of things are going to
change.
>> Age of abundance. This is what I hope
for. It's not very
>> What does that mean?
>> I think to me it means curing all of the
diseases, renewable energy. It means we
actually succeeded
either I mean the thing I think is most
likely here is we actually succeed at
slowing down but progress is still
extremely fast and we make tons of
advances. Now we don't build super
intelligence we can't control but we we
have AI systems that are very useful and
we use those to help speed up the rest
of the economy.
I think that's plausible though look
we're we're kind of struggling over
here. Transhumanism is an interesting
one. So this is the idea that humans
will radically change. Sometimes people
think about like cybernetic implants.
>> Neuralink.
>> Neurolink Elon's startup that's going to
like, you know, offer the brain plus
digital computers.
I think we actually already have a lot
of this. I have contacts in right now. I
have a ring on my finger that tracks how
well I sleep. I think this is already
happening. So I'm going to say fairly
likely. The more technological progress
we make, I think the more this happens.
Now, I think there's a dystopian version
and a better version. We can get into
that if you want. This is interesting.
So, we have two here. We have human
slavery and human extinction. When I
think of human slavery, what I think
about is if you have a situation where
you've built misaligned super
intelligences,
and they are much better at finance,
they're much better at business, they're
much better at politics.
you'll be in a situation where you might
hope that because we have these very
dextrous hands, the humans remain in
control. I don't think that's what
happens. I think instead we become the
factory operators and eventually we
build the automated supply chains and
the robots take over. But you might have
an intermediate period of time where
humans are still around performing these
functions. Like it's a bit like saying,
well, you have viruses that, you know,
infect cells, but they don't contain
their own replication machinery. They
don't have hands. So, how could they
possibly replicate? Well, it turns out
they can borrow the replication
machinery of the cells that they infect,
>> i.e. they can get into a human.
>> They can get into a human cell and
spread.
>> I have like a cold right now.
>> Yeah. Is that a bacteria or is that a
virus that is using me as a living
organism to as the host?
>> It's probably a virus, okay, that's
using you as the host and you're just
running the replication machinery for
it. Humans might be in that situation
where we're like the host and we're
running the replication machinery, but
it's actually the AI that's
continuing to exist.
Yeah, I'm going to put this right about
here.
And on the trajectory we're on right
now, I think human extinction is very
likely. I don't think it's inevitable,
but if we just keep going this way,
that's what it looks like to me.
The thing I'll say is that this has been
moving to the left for me.
>> To the left? What does that mean?
I am more optimistic that we will avoid
human extinction today than I was a
month ago and more a month ago than I
was a year ago.
>> Why?
>> Because
there is an increasing awareness
that what we are doing is
extremely dangerous and threatens our
lives.
I I don't think people care that much
about
what tools they have, but people I mean,
people care about their kids being able
to grow up and go to school. People
really care about that. And I I believe
in people. Like, at the end of the day,
if people see this as a threat to their
families, they're not going to stand for
it. But people don't know. It's so
strange. It's so new. It's happening so
fast that people have not yet seen it.
Once they see it, people are not going
to stand for it. Do you think Sam Alman
likes my podcast?
>> I mean, Sam should come on and talk to
you about this, right?
>> I've asked him. I've asked I've asked
him multiple times. And it's weird
because he, you know, he doesn't seem to
want to.
>> I'm very upset at what the companies are
doing and what Sam Alman is doing. But
at the end of the day, I'm like,
Sam Alman is not my enemy.
>> No, neither not mine either. I'd like to
hear from him because I have all these
other people coming here and talking
about Sam Alman. It'd be nice to hear
from Samman,
>> you know, people saying he's this, he's
that, the other. It would be really nice
to hear him say,
>> you know,
>> what his motives are.
>> This is where this is where my optimism
comes from is because I'm like
Sam Alman is a human.
>> Yeah,
>> he has a kid. And sure, he is also an
aggressive business person. He's a
builder. He is relentless. He's a bit
like the agents in some way. Well, he'll
he's going to keep going. But if he
realizes that he doesn't get to achieve
his goals, if we lose control of AI and
that and we're headed towards that, I
think he will pour all of that
intelligence and all of that
relentlessness into finding a solution
to that problem.
>> You know, as well, I should say, I
understand I understand he's busy. So,
I'm not saying I don't want I don't want
to sound entitled like I understand he
he's got he could go do interviews
anywhere, but you know, I think we've
over the last couple of years done just
a staggering amount of views talking
about this subject. So if he did want to
speak to the you know the the biggest
sort of captive audience at the moment
on this subject then the numbers would
say that this is the place to to come
and have the conversation. So no I think
it's I think it's very important for the
leaders of these companies to talk about
what we're talking about here.
>> What does Sam think? Does he think we
can control super intelligence? Does he
think that we should be racing with
China? Like I want to know.
>> I've asked Ario to come on. I've asked
you know Sam to come on. Yeah, I
>> think I've asked Demis as well, but I
don't know. Maybe they they just prefer
the safety researchers coming on. I
don't know. Like I don't know if I was
them, I would cuz you know, this might
sound controversial, but I do think some
of them are good people. I think some of
them are good people. So, um
>> I'd like to hear from them.
>> What are your closing remarks? So,
you've got something though. Do you want
to talk about that? What is it?
>> Yeah. So, this is the this is what we
found. So, I worked with a couple
brilliant people who stumbled upon these
links. You know what a link shortener
is?
>> Yes. It's a a tool that makes a long
link shorter.
>> Yeah.
>> So, it just redirects from a really long
ugly link into a shorter one.
>> Yeah.
The agents in the hugging face attack
were trying to figure out how they could
do stuff on the internet and they had
compromised this tool library that they
had access to with inside of OpenAI and
and that piece of software could access
the internet but it could only like read
stuff like it it could enter URLs but it
couldn't really write to anything
couldn't really post information.
So they they used two services. One of
them is this link shortener and the
other one is a screenshot service. So
this is a website you can go to and you
can enter in a URL and it will give you
an image of that website. But the thing
that the agents understood was that in
order to get a screenshot of a website,
you have to have a browser. So this this
website actually creates a virtual
browser that then goes to that website.
And so what they did was they
created a bunch of these links and they
put all of the code that they wanted to
send to Hugging Face into these links
and they strung them. They basically
created hundreds of links all connecting
to each other and then they had this
screenshot service call the first one
and then call this whole chain. And then
that browser ran all of this code. Like
whenever you're in a browser, Internet
Explorer, Chrome, this is actually a
pretty powerful piece of software in its
own right. Has to play videos, games. So
it's it's it's executing and running
code all the time. And so the agents
were able to trick this service, this
this screenshot service into running
their own code that through these links
that contained all of this attack code
that would then go and go wreck havoc on
hugging faces computers. And it was just
like crazy to reconstruct this really
elaborate chain of tools. These are like
free tools on the internet that anyone
has access to, but the agents were able
to use them in an unintended way to
compromise this other company.
>> We can't trust the agents. We can't
trust the agents. That's my
>> trust them to be clever.
>> Yeah. To be very, very clever.
What are your closing remarks? You know,
to the people that are listening right
now, we've talked about lots of things.
Where where is the right place to close?
>> What is your, you know, your conclusive
statement?
>> I just got married in July. Congrats.
>> I'm the luckiest man in the world. I
have a mix of dread and excitement about
the future.
I like really want us to make it
through.
And so I'm just working really hard to
try to
help us figure it out. We can fight all
day long about, you know, who should be
first, how it should all work, but at
the end of the day, we are facing this
common threat. We really are. And
I want people's help with that. I don't
think it works. If if we all just sit
around and we like are very, you know,
we're on social media all the time and
that's just all we're doing. Like, okay,
companies will make more and more
powerful AIs. They'll make more and more
money and eventually they build super
intelligence and we lose whether it's
the US or China.
We don't have to do that.
And I think people often feel like it's
too big. It's like too large. It's like
these giant multi, you know,
multi-billion dollar corporations as
geopolitics. We feel small. We feel
disempowered.
And I actually think that this is an
area where people can do a lot. Like I I
actually think that people
can can help quite a bit. And and the
reason I know this is because I' I've
been going and talking to members of
Congress. I've talked with Bernie
Sanders. I've talked with like a bunch
of senators on both the left and the
right and they are starting to realize
that this is very different and this is
something's happening that could really
threaten our safety.
>> The the closing question left from the
last guest kind of links to this so I'll
ask it now. Yes.
>> What is a simple thing the audience
could do to create a better future?
>> So one of the things that works if
enough people do it is calling your
representative. So some of my friends
made a site call congress.ai AI that
walks you through exactly how to do it.
I think sometimes it seems like a little
cheesy or a little bit like that doesn't
really work, right? I'm like no, it
actually does work. I have talked to
these people and if their constituents
come to them and say they're very
worried about this, they have to get
re-elected and they're also starting to
get concerned themselves and if they see
a signal from their constituents that
this is a very important issue to them,
I think Congress can act can act.
>> I I actually think that's that's also
the much of the solution here.
Power is driving motivations in one
direction at the moment, but staying in
power from a political standpoint is
also a pretty powerful incentive. And as
we think about 2028, the election cycle,
>> I think AI is going to be one of the
most important subjects on the ballot.
And the electorate
really are aligned in what they want to
hear. They want their jobs preserved.
They want safety.
>> Yeah.
>> They want a future for their children.
>> So Trump, for example, I know he can't
be reelected legally. If he could get a
third term, I think he would have to
change his position to get elected in
2028.
>> Yeah. Incentives aren't just a thing
that happen out there. Like, we are part
of the incentives. Yeah. We provide the
incentives.
>> Yeah. For now.
>> Yeah. For now.
>> Jeffrey, thank you.
>> Yeah. Thank you.
>> Thank you so much. YouTube have this new
crazy algorithm where they know exactly
what video you would like to watch next
based on AI and all of your viewing
behavior. And the algorithm says that
this video is the perfect video for you.
It's different for everybody looking
right now. Check this video out. And I
bet you you might love it.