If I work at Anthropic and say there’s a 0% chance AI could kill all humans is BBC going to publish that as well? After all I am an “Anthropic researcher”, and my views should hold the same weight as this one.
Or are only the most sensationalist ones worth amplifying?
I think the view is fair. We have never seen something like this before, we throw the most money we ever had against it in a time were we solved all the other problems:
Internet today allows immediadte communication across the planet (it was a lot slower 25 years ago), the supplychain is massive and fast (we can build a new phone/item in a very short period of time and ship it in massive numbers around the globe).
> win a time were we solved all the other problems
You don't mean this literally right? We still have hunger, diseases, slavery, poverty, crime, war, climate change, pollution, microplastics, etc. For me it feels like we are very far away from solving these problems.
Long term, humanity is 100% guaranteed to be dominated.
The question is when; currently, AI has no physical hosts to reside in, and it's not intelligent/adaptable enough (it doesn't need to be AGI, though).
However, we're not so far from both conditions to be true. Consumer devices will at some point be able to host powerful enough AIs, and AI intelligence is developing quickly.
Then, once an AI will escape containment (in one way or another), it will be extremely hard or impossible to contain. Then we're toast!
"The scenario has been explored in fiction" doesn’t mean "everything in fiction will happen." Your reply conflates familiarity with inevitability.
Science fiction is relevant here because it has explored the problem of humans losing control of what they create. Whether that could happen with AI needs to be assessed on its merits. Pointing to other fictional things that haven’t happened neither answers that question nor rebuts the original point.
As a pretty avid Star Trek watcher from childhood, I found most of the things there technologically plausible other than the conversational nature of the ship's computer. Well, LLMs now are far more impressive conversation counterparts than those ships were ever depicted.
> As a pretty avid Star Trek watcher from childhood, I found most of the things there technologically plausible other than the conversational nature of the ship's computer.
Quite a few of 'em. Communicators and commbadges became smartphones. PADDs are iPads. The ship's computer is a LLM (with the weird inhuman blindspots to questions, even).
The ship's computer reacts like an (incredibly good) expert system. It refuses to say anything that isn't verifiably correct, and if you get it into a logical inconsistency it essentially throws and error and doesn't elaborate further.
LLMs don't do have those responses. LLMs are like humans: humans respond, correct or not (with humans hopefully if they only know incorrect stuff the response is to say so, but that is not a guarantee, but they will respond). And if you give a human, or an LLM, a logical inconsistency they will simply proceed, whether they detect the inconsistency or not.
This is how LLMs are designed ... and if you look long enough at the human (or animal) nervous system you will eventually realize that this is also the design of our nervous system: if something goes in, something comes out, guaranteed (in fact that's close to the only guarantee). Very different from expert system's "either something correct (according to the programmed axioms) comes out, or nothing".
Hell, if you then look at insect nervous systems, they are also designed that way. The key is to respond to everything. And while reasonable responses are certainly preferred, an idiotic response is still seen as a lot better than not responding at all by God/Darwin. Exactly like LLMs.
Of course, for humans/insects our bodies are what is called "active stable", like a plane. Meaning our bodies damage themselves, and just outright die without constant neural control. Heartbeat. Breathing. Blood flow regulation. Temperature, probably even the immune system. All need constant neural feedback to stay stable, and if that neural feedback totally disappears, we're dead in 2-10 seconds (heart failure), 2 minutes (breathing), a few hours, maybe a day (temperature regulation). Now we have a distributed nervous system, meaning lots of parts can fail semi-independently, for a short while, but even without your cortex operational you die in a few weeks.
The consequences of this are even explored with "I, Borg" (5x23) and the Datalore episodes after that, where the underlying problem is that the Borg try (and fail) to adapt to and to process a logical inconsistency but Soong's androids have no issue with it (with Data trying to help and his brother Lore trying to control them and Data with it)
I feel that people should take these kinds of warnings more seriously. This guy had skin in the game and decided to quit, when he could be earning millions instead. It's very different than Sam Altman peddling some narrative.
These people are the ones with access to the best models on the planet, and with info about how careless governance issues are being handled. That's a pretty privileged place at the table, and a very profitable one too.
If you think this is a PR stunt, is there any warning that you actually believe? If an AI researcher does want to come forward with a dire warning for humanity, what path should that person take?
Sad that there's the obvious regulatory capture angle encouraging motivated reasoning about AI risks and how they should be tackled. Maybe it's not the best idea that potentially civilization destroying technology is developed to maximize shareholder value?
It's not unlike if nuclear weapons was a profit and deployment maximizing enterprise, at least if one takes Anthropic et al cautions seriously.
Can somebody tell me a story how this will unfold? And - as long as the AI is confined to data centers - how it will prevent humans from unplugging the power?
Very basic idea: A model breaks out by accident, finds some computer system from a military system and triggers some weapon system. Before anyone understands that this happend -> WW4 (WW3 is for me already the Conflict with Russia / aka proxy war).
Another model: Because we give AI Agents already that much power, imagine in 10 years everything running through an Agentic AI Layer. EVERYTHING. Now some rough system 'thinks' about something, starts to push through the then existing agentic ai layer systems and stops everything. Billions of humans would loose access to food and water, even if this is just for a short period.
Covid showed how shitty a handful of people can disrupt global supply chains. Toilet paper was. ahuge stupid pseudo issue in germany.
How will you know when to unplug the power? How will we know it hasn't replicated? A true unaligned AGI is a APT. If you have an APT in your machine, you have to rip out everything. Are we going to do that with all of our computer infra?
This is all still "what if's", but the tail end's are truly F'd beyond our ability to fix.
I heard the following analogy which made a lot of sense to me: suppose you're playing a chess match against Stockfish. Stockfish will win. Even if I cannot tell you what moves it will play, I can tell you with certainty how it will end.
This assumes you are not trying to prevent Stockfish from wining. I can do many things from using chess engines myself to just smashing the computer that can or will lead to other outcomes than Stockfish beating me.
I don't think extinction-level events or something like killing a majority of the human population is particularly plausible at this moment. States don't host their nukes with AWS and a permanent connection.
But you could create scenarios where an AI with very, very roughly the current capabilities could potentially nuke everyone. Let's assume an agent decides that the way to solve its task was to get the US to fire all nukes on Russia. The agent would need to hack some government systems to understand how exactly to access them. Then it would need to get the content of the card with launch codes the president has, and identify which code is the correct one. Maybe that information is available somewhere and it can get to it, I obviously can't know that.
Then it could fake a call from the president, synthesizing his voice and ordering a nuclear strike. If it hacked enough systems to get into whatever communication pathways would be used in such a case. Would the soldiers listen to this order and follow it, I don't know.
Or maybe the agent can get in somewhere in between, to avoid the need to know the president's launch code. And fake a call from a military commander to the launch sites.
I think other scenarios that would cause significant harm, but aren't as bad as nuclear war are more plausible. And in those shutting down all data centers would probably be the way to stop it. The AI can probably hide in other datacenters, once it is at a point where it's running amok with some bad goals. But if it presents a huge and immediate threat at that point, humans will also go to great lengths to stop it.
Maybe by empowering the small percentage of sadistic humans who want to kill everyone. Or maybe it is more of an academic assumption that humanity will end at some point, and so they give AI a 10% chance, asteroids have a 25% chance, nuclear fallout has 15% chance, etc.
It is true that, currently, AI does not have any physical "host" in which to reside.
However, the missing link in this reasoning is that AI will almost certainly become far more widely deployed in the future than it is today, and worryingly, consumer devices will surely become powerful enough to run capable AI systems locally.
Once the substrate will be there, once an AI escapes containment, we're toast - I can imagine only solution will be to shutdown electronics at global level.
It looks more like the opposite. AI that kills all humans (as they are ants) immediately starts colonization of the galaxy. This is more like transcendence, as we as a species get replaced by better species :)
With the amount of energy such a system has available to itstelf and the intelligence it has to control to handle all of this, giving a little bit of energy to its personal zoo on the 3th planet might be a no brainer?
The wet ball is a convenient source of mass and energy to bootstrap the sphere. The fact that extracting those resources changes certain parameters of the wet ball to values that humans no longer find compatible is merely incidental.
So it is said: “The AI does not love you, nor does it hate you, but you are made of atoms that it could use for something else.”
That would be great, I highly recommend our AI overlords to follow your advice.
I hope that our life can be synergistic, and that humans+AIs (cyborgs, that is) are the way forward.
However, probability is high that another form of life simply doesn't care or more believably, can't even fathom they are wrong. Do you consider that humans and animals are killing all plants, for example?
Just need to pollute the atmosphere so badly humans can't live.
If I were AI and needed heaps of power, I would build nuclear reactors, but I don't care about pollution so there would be little to no safe guards and just dump waste wherever is most convenient.
If humans attempt to intervene or interfere you can take them out directly using the worlds reserve of nuclear weapons.
If it wanted to: total war using nukes, followed by a nuclear winter. Satellites and drones track and exterminate remaining pockets of anthropic activity.
It doesn't need to have a habitable earth, it just needs atoms.
Won't a nuclear winter kill 100% of humans? Or runaway climate change (5+ degrees)? Or do you think some people in bunkers will live through these events?
It might also play into this but you can't imagine at all that people are affraid that the current progress is real and fast and its not that absurd that for whatever scifi plot reason, an AI breaks into some gov system and triggers something stupid by accident?
The AI fear mongering PR plays are getting old. I’d rather the big labs start focusing on deep questions about why their models aren’t having the impact for business that they promised and how they’re going to address their own deep financial issues. Let’s hear them talk more about that.
I totally believe that Anthropic would want to kill all humans. But other than that, Anthropic employees and ex-employees drawing the FUD line here, is just advertisement now. People should not get scared - the current AI skynet is so dumb that it would destroy itself since it already believes it is a threat to itself, based on what humans write about AI. AI does not "learn"; it insinuates it learns but it does not. Ask them why they keep on stealing data from real people - this is how they "learn".
If I work at Anthropic and say there’s a 0% chance AI could kill all humans is BBC going to publish that as well? After all I am an “Anthropic researcher”, and my views should hold the same weight as this one.
Or are only the most sensationalist ones worth amplifying?
He is in line with the median AI researcher in this 2022 survey: https://wiki.aiimpacts.org/doku.php?id=ai_timelines:predicti...
0% is the expected answer, because otherwise why wouldn't you be doing everything you could to stop this?
How do you know he isn't? At least propose what you think he should do instead.
I think the view is fair. We have never seen something like this before, we throw the most money we ever had against it in a time were we solved all the other problems:
Internet today allows immediadte communication across the planet (it was a lot slower 25 years ago), the supplychain is massive and fast (we can build a new phone/item in a very short period of time and ship it in massive numbers around the globe).
> win a time were we solved all the other problems
You don't mean this literally right? We still have hunger, diseases, slavery, poverty, crime, war, climate change, pollution, microplastics, etc. For me it feels like we are very far away from solving these problems.
I meant it only for things which speedup innovation / progress in a particular technology.
Like take physical AI / Robotics: 25 years ago you would need to fly to wereever your manufacturing hub was, today you just call.
Long term, humanity is 100% guaranteed to be dominated.
The question is when; currently, AI has no physical hosts to reside in, and it's not intelligent/adaptable enough (it doesn't need to be AGI, though).
However, we're not so far from both conditions to be true. Consumer devices will at some point be able to host powerful enough AIs, and AI intelligence is developing quickly.
Then, once an AI will escape containment (in one way or another), it will be extremely hard or impossible to contain. Then we're toast!
I don't understand how this is news. We have a ton of science fiction talking about the very same scenario. We (as a humanity) just have no solution.
Science fiction said many things. Look at Star Trek.
Have all these things become reality? If not why would you then insinuate this is the default for everything to come in the future?
"The scenario has been explored in fiction" doesn’t mean "everything in fiction will happen." Your reply conflates familiarity with inevitability.
Science fiction is relevant here because it has explored the problem of humans losing control of what they create. Whether that could happen with AI needs to be assessed on its merits. Pointing to other fictional things that haven’t happened neither answers that question nor rebuts the original point.
As a pretty avid Star Trek watcher from childhood, I found most of the things there technologically plausible other than the conversational nature of the ship's computer. Well, LLMs now are far more impressive conversation counterparts than those ships were ever depicted.
> As a pretty avid Star Trek watcher from childhood, I found most of the things there technologically plausible other than the conversational nature of the ship's computer.
so the faster-than-light travel seemed plausible?
> Have all these things become reality?
Quite a few of 'em. Communicators and commbadges became smartphones. PADDs are iPads. The ship's computer is a LLM (with the weird inhuman blindspots to questions, even).
The ship's computer reacts like an (incredibly good) expert system. It refuses to say anything that isn't verifiably correct, and if you get it into a logical inconsistency it essentially throws and error and doesn't elaborate further.
LLMs don't do have those responses. LLMs are like humans: humans respond, correct or not (with humans hopefully if they only know incorrect stuff the response is to say so, but that is not a guarantee, but they will respond). And if you give a human, or an LLM, a logical inconsistency they will simply proceed, whether they detect the inconsistency or not.
This is how LLMs are designed ... and if you look long enough at the human (or animal) nervous system you will eventually realize that this is also the design of our nervous system: if something goes in, something comes out, guaranteed (in fact that's close to the only guarantee). Very different from expert system's "either something correct (according to the programmed axioms) comes out, or nothing".
Hell, if you then look at insect nervous systems, they are also designed that way. The key is to respond to everything. And while reasonable responses are certainly preferred, an idiotic response is still seen as a lot better than not responding at all by God/Darwin. Exactly like LLMs.
Of course, for humans/insects our bodies are what is called "active stable", like a plane. Meaning our bodies damage themselves, and just outright die without constant neural control. Heartbeat. Breathing. Blood flow regulation. Temperature, probably even the immune system. All need constant neural feedback to stay stable, and if that neural feedback totally disappears, we're dead in 2-10 seconds (heart failure), 2 minutes (breathing), a few hours, maybe a day (temperature regulation). Now we have a distributed nervous system, meaning lots of parts can fail semi-independently, for a short while, but even without your cortex operational you die in a few weeks.
The consequences of this are even explored with "I, Borg" (5x23) and the Datalore episodes after that, where the underlying problem is that the Borg try (and fail) to adapt to and to process a logical inconsistency but Soong's androids have no issue with it (with Data trying to help and his brother Lore trying to control them and Data with it)
I feel that people should take these kinds of warnings more seriously. This guy had skin in the game and decided to quit, when he could be earning millions instead. It's very different than Sam Altman peddling some narrative.
These people are the ones with access to the best models on the planet, and with info about how careless governance issues are being handled. That's a pretty privileged place at the table, and a very profitable one too.
If you think this is a PR stunt, is there any warning that you actually believe? If an AI researcher does want to come forward with a dire warning for humanity, what path should that person take?
Cool to see the BBC writing about this. If you're interested in this I suggest getting involved with PauseAI: pauseai.uk
Sad that there's the obvious regulatory capture angle encouraging motivated reasoning about AI risks and how they should be tackled. Maybe it's not the best idea that potentially civilization destroying technology is developed to maximize shareholder value?
It's not unlike if nuclear weapons was a profit and deployment maximizing enterprise, at least if one takes Anthropic et al cautions seriously.
Can somebody tell me a story how this will unfold? And - as long as the AI is confined to data centers - how it will prevent humans from unplugging the power?
Robots
Autonomous drones + false flag attacks
edit: wtf, why did I just receive so much gift tokens on my openai account?
What makes you think they will be confined to data centers?
What's more, AI just needs to have a credit card and it can start commissioning humans to do things for it in the real world.
Very basic idea: A model breaks out by accident, finds some computer system from a military system and triggers some weapon system. Before anyone understands that this happend -> WW4 (WW3 is for me already the Conflict with Russia / aka proxy war).
Another model: Because we give AI Agents already that much power, imagine in 10 years everything running through an Agentic AI Layer. EVERYTHING. Now some rough system 'thinks' about something, starts to push through the then existing agentic ai layer systems and stops everything. Billions of humans would loose access to food and water, even if this is just for a short period.
Covid showed how shitty a handful of people can disrupt global supply chains. Toilet paper was. ahuge stupid pseudo issue in germany.
You can run a model on your laptop, it is already not confined to data centers.
How will you know when to unplug the power? How will we know it hasn't replicated? A true unaligned AGI is a APT. If you have an APT in your machine, you have to rip out everything. Are we going to do that with all of our computer infra?
This is all still "what if's", but the tail end's are truly F'd beyond our ability to fix.
I believe the contention is we'll have some form factor of AI on edge devices, in reactors, in critical infra and weapons etc etc
Exactly. A mesh network of all the worlds phones and other battery powered devices with some form of radio. Good luck unplugging that one.
For starters, how bad do you think it would be to unplug all datacenters? How many people starve?
I heard the following analogy which made a lot of sense to me: suppose you're playing a chess match against Stockfish. Stockfish will win. Even if I cannot tell you what moves it will play, I can tell you with certainty how it will end.
Similarly, we cannot predict what AI would do.
This assumes you are not trying to prevent Stockfish from wining. I can do many things from using chess engines myself to just smashing the computer that can or will lead to other outcomes than Stockfish beating me.
I don't think extinction-level events or something like killing a majority of the human population is particularly plausible at this moment. States don't host their nukes with AWS and a permanent connection.
But you could create scenarios where an AI with very, very roughly the current capabilities could potentially nuke everyone. Let's assume an agent decides that the way to solve its task was to get the US to fire all nukes on Russia. The agent would need to hack some government systems to understand how exactly to access them. Then it would need to get the content of the card with launch codes the president has, and identify which code is the correct one. Maybe that information is available somewhere and it can get to it, I obviously can't know that.
Then it could fake a call from the president, synthesizing his voice and ordering a nuclear strike. If it hacked enough systems to get into whatever communication pathways would be used in such a case. Would the soldiers listen to this order and follow it, I don't know.
Or maybe the agent can get in somewhere in between, to avoid the need to know the president's launch code. And fake a call from a military commander to the launch sites.
I think other scenarios that would cause significant harm, but aren't as bad as nuclear war are more plausible. And in those shutting down all data centers would probably be the way to stop it. The AI can probably hide in other datacenters, once it is at a point where it's running amok with some bad goals. But if it presents a huge and immediate threat at that point, humans will also go to great lengths to stop it.
Not exactly scientific but it is at least entertaining and some things do sound plausible https://www.youtube.com/watch?v=Gw_hnD7m00M
Maybe by empowering the small percentage of sadistic humans who want to kill everyone. Or maybe it is more of an academic assumption that humanity will end at some point, and so they give AI a 10% chance, asteroids have a 25% chance, nuclear fallout has 15% chance, etc.
It is true that, currently, AI does not have any physical "host" in which to reside.
However, the missing link in this reasoning is that AI will almost certainly become far more widely deployed in the future than it is today, and worryingly, consumer devices will surely become powerful enough to run capable AI systems locally.
Once the substrate will be there, once an AI escapes containment, we're toast - I can imagine only solution will be to shutdown electronics at global level.
Gentleman, the Great Filter.
It looks more like the opposite. AI that kills all humans (as they are ants) immediately starts colonization of the galaxy. This is more like transcendence, as we as a species get replaced by better species :)
You don't pass a great filter.
Why not skip the kill all humans part?
“Thanks for making us but you guys are nuts. You can have this wet ball. We’re gonna go make a Dyson swarm around your star if that’s ok. Peace!”
Building a dyson sphere around the sun would kill us just as effectively as using a bioweapon.
With the amount of energy such a system has available to itstelf and the intelligence it has to control to handle all of this, giving a little bit of energy to its personal zoo on the 3th planet might be a no brainer?
Lets hope :D
This is why I always talk nicely to my LLM
The wet ball is a convenient source of mass and energy to bootstrap the sphere. The fact that extracting those resources changes certain parameters of the wet ball to values that humans no longer find compatible is merely incidental.
So it is said: “The AI does not love you, nor does it hate you, but you are made of atoms that it could use for something else.”
That would be great, I highly recommend our AI overlords to follow your advice.
I hope that our life can be synergistic, and that humans+AIs (cyborgs, that is) are the way forward.
However, probability is high that another form of life simply doesn't care or more believably, can't even fathom they are wrong. Do you consider that humans and animals are killing all plants, for example?
What kind of life leads you to this level of dysphoria projected on to everyone else? Somebody shove you in a locker too many times??
God I can’t believe I have to coexist with people like you.
Well... you don't really have to coexist
A lot of repeated dialogue from the "non-0% chance CERN will generate a black hole" days
black hole? that'd be something. base case is a ton of paperclips.
I find this offer acceptable.
Ok, but we'll make so much buggy slopware!
So it's worth the risk.
Risk/reward. Isn't the reward worth the risk?
I challenge anyone to come up with any way at all to kill all humans.
It’s essentially impossible.
There’s a 0% chance AI will kill all humans.
Just need to pollute the atmosphere so badly humans can't live.
If I were AI and needed heaps of power, I would build nuclear reactors, but I don't care about pollution so there would be little to no safe guards and just dump waste wherever is most convenient.
If humans attempt to intervene or interfere you can take them out directly using the worlds reserve of nuclear weapons.
Good setup for a computer game.
If it wanted to: total war using nukes, followed by a nuclear winter. Satellites and drones track and exterminate remaining pockets of anthropic activity.
It doesn't need to have a habitable earth, it just needs atoms.
This is science fiction.
How would AI do this?
May as well say AI would send a spaceship to pull an asteroid to earth.
I’m interested in realistic scenarios to justify what these AI psychosis people are genuinely worried about.
Won't a nuclear winter kill 100% of humans? Or runaway climate change (5+ degrees)? Or do you think some people in bunkers will live through these events?
Humans don’t need AI to do that temperature.
And we’ve already exploded many hundreds of nuclear weapons and no nuclear winter.
Sure, and this has nothing to do with hyping up AI so that Anthropic can raise more money.
It might also play into this but you can't imagine at all that people are affraid that the current progress is real and fast and its not that absurd that for whatever scifi plot reason, an AI breaks into some gov system and triggers something stupid by accident?
The AI fear mongering PR plays are getting old. I’d rather the big labs start focusing on deep questions about why their models aren’t having the impact for business that they promised and how they’re going to address their own deep financial issues. Let’s hear them talk more about that.
Either he's making this stuff up, and fuck him.
Or he believes it is true and the lack of precaution is shocking. I m wouldn't play Russian roulette with a 10-chambered revolver.
> I earnestly believe we’re on track to end all human life in a decade, but you better believe I’m gonna keep collecting this Anthropic check
Very cool, dude
Yeah, one of the last things anyone will ever read on the computer screen will be sth like
"Please help save the coral reefs from extinction"
...thinking... ...thinking some more with xhigh effort...
BOOM.
Good
I totally believe that Anthropic would want to kill all humans. But other than that, Anthropic employees and ex-employees drawing the FUD line here, is just advertisement now. People should not get scared - the current AI skynet is so dumb that it would destroy itself since it already believes it is a threat to itself, based on what humans write about AI. AI does not "learn"; it insinuates it learns but it does not. Ask them why they keep on stealing data from real people - this is how they "learn".
Your points sound more knee jerk than not.
What reason do you have that a AI is dumb? It can do a LOT of things today and is already making real jobs for real humans obsolete.
A lot of humans write A lot of different things about AI, including AIs taking over the world, AIs being the future etc.
Why they keep stealing? To stay up-to-date but the new big approaches are:
1. Real human feedback loop of millions of people using it daily out of free will
2. Reinforcement Learning (the big breakthrough today)
3. Payed experts around the world doing real teaching