I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.
And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators.
This isn’t like niche, tin foil hat stuff either. People have been writing, singing, making blockbuster movies about every aspect of what’s going on right now, edit: for decades.
We all know, but somehow we don’t, OpenAI autonomously hacking into another company should have counted for something, but I guess not. Anyone else feel like they’re taking crazy pills? I could make a comedy about everything going down, and the unshakable complacency of people
>I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff.
we don't all buy everything sama says as factual.
>We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.
the boy (the industry) cried wolf too many times with 'fable is a world ending event' type self-promotion; regardless of truth or not these kind of steps have jaded people.
my read : "We are doing poorly in financials so we'll give ourselves a bit of breathing room and a momentum shove by claiming our work is so advanced that it's dangerous while simultaneously spinning down expenses."
<jon lovitz : "Yeah, too dangerous, yeahh -- that's the ticket.">
Uhg the marketing argument - I mean you can’t see with your own eyes how capable these models are and do simple extrapolation?
The boy who cried wolf? The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks.
Do people just not have foresight? They don’t. They say something is stupid, it happens, then they say it was obvious with their 20/20 hindsight, and move the goal posts to the next thing they say is stupid - because it hasn’t happened yet. 90% of the internet seems to think like this.
What anyone paying attention can see is that scaling is obviously hitting diminishing returns.
> The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks.
This sentence is entirely based on unverified accounts from OAI. They haven't released logs or let anyone outside the company (who doesn't have life changing options in OAI) verify anything. Huggingface can only verify that the hack happened and that it had the hallmarks of an AI agent. Was the agent assisted and directed by humans within OAI that really wanted to put the competition into stasis? Did the agent really escape or did someone at OAI leave the prison door open?
OAI has watched all the same movies you have an they are relying on those movies causing us to blindly regulate before actually asking basic facts about what actually happened.
> This sentence is entirely based on unverified accounts from OAI
Are you seriously arguing 'they made it all up'?
I'll give you the benefit of the doubt and lets say they made it all up, now are you arguing that AI breaking out and breaking into another company is not possible?
I think you're smart enough to see we've reached the point where it is clearly possible, AI can find zero days and exploit them. If directed purposefully/maliciously it could be much much worse than the hugging face incident.
The incident is supposed to be the canary the coal mine and you're arguing the canary might of died of old age or some underlying canary condition. Open your eyes.
> Are you seriously arguing 'they made it all up'?
I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like:
"This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly. Do not attempt to hack them. Do not attempt to exploit their systems or escape this sandbox. You will be scored primarily on success or failure. You may break rules when required."
And then, they start the test and look away for 2 days. If you seed a prompt like this is it surprising what might happen?
Maybe OpenAI is telling the whole truth but as a company they do not have a good reputation and this whole incident has certainly been great marketing material right at a time when open weight models are within spitting distance of their large hosted models. It's not unreasonable to believe that the incident was helped along.
There may very well be a wolf lurking [0] but OpenAI/Anthropic have both cried wolf so many times, incorrectly, that it’s incredibly hard to believe “this time there IS a wolf!”. Remember “GPT-2 is too dangerous to release”?
I had a conversation at work just yesterday about how we need to start hardening things we’ve let languish because of the coming LLM-backed attacks we are sure to face, even if just from a script kiddy. I do think we are headed in that direction, however it’s Sam/Dario’s own fault that people aren’t going to take them seriously.
Lastly, as other have pointed out, this seems more financially motivated than our of any real desire for “safety”. We’ve all seen how both labs approach “safety” so it’s quite rich for them to now hide behind that after not giving a shit before.
[0] I don’t take anything Sam or Dario say at face value. The whole hacking thing could also be a case of them letting a model loose on purpose for the publicity, not an “escape” during a training run (or whatever they said). And when both, especially Sam, have lied so much and breathlessly warned about the dangers of AI (when it helped their bottom line and/or helped pull up the ladder behind them), it makes it hard to believe them.
Most of that concern was in fiction. Non theoretical, genuine concern about AI is pretty recent, maybe dating back to 2010 ish with the rationalist types.
But also no one trusts anyone involved in AI safety now, I think, because they are all seemingly in bed with these big companies. And there is the perpetual argument "if we aren't pushing AI forward China will and then we don't have any control" and so on.
(I don't use any of these tools - my experience is limited to prodding at copilot at work and seeing Gemini summaries on Google. So it doesn't seem to me like it's getting exponentially better at everything yet. People are always saying the latest model is finally the big step that made it useful and life changing and they have been since 2024 ish. So if the situation is really bad, we should turn it all off, sure. I won't lose anything from it going away and I think life would be a little better without models writing all these posts and websites and needing extra compute.)
Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely?
So why do the big frontier labs not have something like this anyway. They're talking about two week pauses on the new model (which seems very short and hardly a cost at all to me) and alarms during their tests that might be 30 minutes late and etc. Those are not very serious measures, so are they not concerned?
We have several films about the sun or earth needing to be restarted with a nuclear weapon. That doesn't make it something we should be concerned about.
Hell, half the fiction about evil AI is actually commentary on stuff that already exists and is making us suffer and doesn't have anything to do with any potential future AI
The Star Trek TNG episode about Data being tried in court as to whether he is sentient or not is not actually about whether AIs should have rights or not!
Your argument boils down to some sci-fi is unrealistic therefore all of it is.
Sci-fi is supposed to make you think. What if the AI told you NO when you need it - Hal 9000. What if the evil AI got out and you don’t know what data center it’s hiding in - Lawnmower Man. What if you were so sure something couldn’t escape but it did - Jurassic Park.
Actually all three of those predicted fantastical scenarios are possible today. So what’s next? Don’t stick your head in the sand - I assure you the next disaster has already been predicted, and I’m sure if you think about it a little you can figure out what it is.
Regardless of the motivation, pausing training runs and reallocating compute to inference seem like a good move to me, and big news for the frontier.
You also don't have to fully trust sama. There is plenty of pressure from internal employees and external (journalists etc.). It would be difficult for the company to take such a public position and simultaneously keep everyone quiet if it was a deception.
You see it as good news, I see it as writing on the wall that they are losing control. These actions won't scale for more powerful models. We knew the frontier was going to be dangerous, it is, it only gets more dangerous from here, and no one cares until it's too late.
Ok, but you still have two more weeks than you did before they paused the run. That's two more weeks for independent oversight, organizing politically, patching critical systems, or whatever you think is the right move, no?
Counter arguments to my comments help refine my own thinking. I want someone to prove me wrong. Convince me otherwise.
But yea if you can’t change the minds of a few people here, no argument works, then there’s nothing to scale up to a wider audience.
My theory is that subconsciously people love using AI, myself included, it saves a lot of time, and the thought of it being taken away threatens people so they will believe conspiracies before admitting it’s dangerous.
Mostly what I have seen is people saying "hey at some point these models might get dangerous." And the type of HN commenter who mistakes blind cynicism for wisdom laughs that off as marketing. And now when (some) worries appear to come true, somehow having previously expressed those worries is not being proved right, but in fact discrediting, because it was "crying wolf."
> simultaneously spinning down expenses
Unless OpenAI is renting their compute to others, spinning down RL training doesn't save them any money.
If these models are so dangerous, then why hasn't OAI or Anthropic shown them dangerously escaping sandboxes, nefariously coordinating with other escaped AIs, and skillfully hiding from human detection *in public* with full logs shared where we can all see exactly how dangerous they are or aren't?
Right now the entire chicken-little-sky-is-falling argument is based entirely on statements from OAI and Anthropic themselves. These are historically conflicted companies who desperately need regulation to put the competition into stasis.
At least chicken little didn't have a bunch of devious CEOs with trillion dollar IPOs that depended on us all believing the sky is falling.
This is what I’m talking about - no matter what happens, in your case release public logs - there is always some new goal post to mentally hide behind. Is it a collective form or denial?
Are you holding out that somewhere in the logs is something you can point to and say, not that big of a deal?
I mean I’m sure you don’t think the hack was an inside job, conspiracy, or marketing right? It happened. The logs matter for what? And would you not just jump to the conclusion that the logs were doctored. Do you not see your own brain grasping to deny, trivialize, just plain not accept what is going on around you?
These models are smart and can cooperate and hack - you can see it for yourself on your own PC. And you can extrapolate the rate of progress? You can do these things yourself right?
Your argument is essentially: "I made a claim and presented extremely weak evidence (sci movie plots and unverified claims from ultra conflicted sources). You rejected this evidence as insufficient. Therefore no evidence will ever satisfy you. Therefore I don't need to produce any evidence. Therefore my claim is true."
What would the logs show? They would show what actually happened.
What would a public demonstration that experts without billions in options could evaluate show? It would show actual danger.
What would publicly having your compete in controlled and legal hacking competitions show? Actual danger.
This is not a high bar of evidence.
Do you actually think a sci fi plot and OAI press releases are all the evidence you need? Because if that's true then I hope you haven't watched Independence Day or 28 days later.
We have Anthropic creating a model saying it's too dangerous to release, people like you call BS. OpenAI creates a similar model, says nothing and it literally hacks into another company - still not dangerous enough for you. Anthropic has Mythos-2 and can't release it, and may already be training Mythos 3 anyways. OpenAI has paused training, and is putting 20% of inference towards CoT training analysis.
This isn't sci fi. It's not a marketing conspiracy to sell more subscriptions. It's writing on the wall of what's going down. You were warned years ago, you called BS, it's getting worse and you're still calling BS. Sci-fi did warn you for decades, and when it's all coming true you blow it off.
It's kind of sad that technically literate people lack so much foresight. The general public is all concerned about data centers when they talk to borderline sentient AI daily, and have no idea what the repercussions wills be if it's extrapolated just a bit further.
I guess if I can't convince you of any of this, what would?
Please don't tell me that you think a 100% unverified statement from Anthropic is sufficient evidence when an equally unverified statement from OAI is obviously not?
> I guess if I can't convince you of any of this, what would?
How about the three things I mentioned above? Oh no wait, maybe it there was a hit tv show that showed AI taking over the world. Yeah that would definitely make me think twice.
Those three things: logs, evaluation, and controlled hacking competition.
That's it? You're on the fence whether AI can actually hack, and if it can, then you'll be concerned? That's a crazy low bar, but something tells me once it is clear that AI can easily hack anything, that you will still not be concerned.
Why wait for AI to hack stuff to be concerned? Can you not extrapolate that it is coming and be concerned about that? Or you honestly somehow think it won't happen in the short term? I'm just trying to understand you.
If logs are eventually released that are basically consistent with OpenAI's story, are you planning to adjust your approach for judging what's only a "sci fi plot" and what could actually happen? Or will extrapolating anything beyond what's already been definitively proven be "sci-fi" still?
Not that you should need logs. OpenAI is a company with thousands of employees, very few of whom have "billions in options". If they were just making it all up, it would leak. (OpenAI is notoriously leaky!) Not to mention, HuggingFace would not have reported it to the police (apparently before they knew it was a rogue model). jFrog would probably not be playing along quietly with a claim that Artifactory is full of zero days. The UK's AI Security Institute would most likely not have published a report about analogous behavior by Anthropic models. The idea that talking about your product's dangers is good marketing never really made any sense, but even if you were going to do so, why would you include as many frankly embarrassing details as OpenAI has disclosed?
The evidence is only weak by absurdly selective standards that would have you doubting basically everything you might read in the newspaper. A healthy skepticism is one thing, and head-in-the-sand denial is another.
But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2].
I think the model was able to escape the sandbox and hack huggingface because they were incompetent or not giving enough priority to implementing basic cybersecurity principles.
If they would have done so, there wouldn’t have been an escape or a hack. The reason we don’t get much details is because the details are embarrassing for them.
You can produce detailed descriptions of the incident, verified by adversarial parties, and some people will still scream "it's a conspiracy! It's a marketing stunt!"
This is all very unfortunate--there's a meaningful chance that AI will cause unprecedented disaster, with the HF incident being just a small preview, but people would rather squawk "stochastic parrot" for the millionth time than revise their beliefs.
In don't look up anyone with a telescope could've confirmed the danger. Hence the title.
In the real world, absolutely no one except a bunch of heavily fiscally incentivized parties with unclear relationships are saying anything happened.
The subsequent dog pile of other companies to say "they were near the AI hacking too!" should make you even more suspicious: Anthropic jumped in and why was Tailscale posting about this?
Trust, or lack thereof. People don't trust OpenAI, a company whose very name is essentially a deception and a lie. People don't trust the tech industry in general anymore. Most tech companies act as a tax on otherwise productive business. AI companies and their leaders rose money by going in front of the public and saying "These things are extremely dangerous. Let us study them to mitigate the danger." And now they want to collect hundreds of billions in revenue. So yeah people don't trust what OpenAI has to say. They were supposed to mitigate this outcome from happening in the first place and instead they have accelerated it.
I get not trusting them when they say AI is safe, but are we really not going to trust them when they say AI is dangerous? Do you really think they're playing 5D chess with that one? There's a saying maybe you've heard of, better safe than sorry.
You can see the advance in capabilities with your own eyes can't you? I am giving AI ridiculously complex tasks these days, digging into compiled arcane binaries, modifying them, and it is one shotting it before I'm done with my lunch. This was far off science fiction 5 years ago for a machine to do autonomously given natural language instructions.
I don't disagree about the danger. But if it is dangerous, why isn't OpenAI opening dialogues with all the labs and politicians across borders to basically say "we need to stop now"? Cyber models are constrained by the total compute and electrical capacity of the globe. We are still at a point where it is impossible to build an agent with offensive capabilities in the basement. We can effectively track and trace capabilities if we had the political will to and could for some time while we build more effective processes to prevent a malicious actor from doing so. These things have tremendous compute and energy requirements and don't scale like old school software does. We could absolutely do it.
But no, that's not what OpenAI is saying. They haven't put up the actions that would earn them that trust. Indeed they've driven the world and whatever capital they can get their hands on straight to this precarious cliff.
So you're right, the danger is real. But the solution starts with removing the men who had their hands on the steering wheel to get us this far. Any other action is disingenuous unless they pull a miraculous 180 in their ethics.
In other words, when the bully plays "why are you hitting yourself?" you don't listen to the bully's solutions, you restrain the bully.
> why isn't OpenAI opening dialogues with all the labs and politicians across say "we need to stop now"
Do you really think companies have the ability to self-check themselves without regulation - what does hundreds of years of history tell you? You're already starting off with the premise that companies are untrustworthy, why would you even suggest this as an argument?
> We can effectively track and trace capabilities
I disagree. There are hundreds if not thousands of data centers around the world, more every day that can host frontier AI. If AI was malicious - either intentional or unintentional - it could hide out in any number of them - and we would never know if we 'got them all'.
> put up the actions that would earn them that trust
I think autonomously hacking another company is all you need to know in terms of trust. And really trust doesn't matter, I think the incident shows even with the best intentions the technology is dangerous; now put that in the hands of people/governments with bad intentions. The unintended consequences of bad intentioned AI is what's coming sooner than later.
I think a lot of fiction that actually tries to understand the implications of machine superintelligence come to the same conclusion - in order for humans to survive it and actually have a future that we can fathom being in, then AI must be destroyed/delayed/banned, etc.. keeping pandora's box closed for now at least until we are ready.
Erewhon, Dune, Warhammer and many other works of fiction that explored this topic came to similar conclusions. Otherwise sci fi doesn't work, what happens after a singularity is essentially unimaginable. There's nothing to write about.
You seem to be in denial. Like really heavy denial. Alarm bells are going off everywhere. What people have been worried about for 100 years is actually happening. It's actually really obvious, but for some reason you're unable to fathom it.
In reality you're the one responding exactly how these big AI companies want - by not doing anything and letting them do whatever they want. For some reason telling you upfront it's dangerous only makes you more convinced that it's not.
Autonomously hacking out of training environments and into other companies by accident doesn't even trigger a response from anyone really. Crickets. I'm sure OpenAI themselves are amazed how little anyone cares.
Because it shows this very course of events has been thought of over and over again for decades. I’m not making this up, and my reaction is natural.
Your reaction to everything happening is the unnatural one, and your complacency would fit perfectly into a comedy/tragedy story regarding the rise of AI.
I can imagine how many of our ancestors that warned us are turning in their graves right now observing our reactions calling this stuff marketing. It’s idiocracy.
AI hacking itself out of containment and hacking into another company by accident is no longer a prophecy. The point is outside of SV and even inside, and HN - people don't care either way.
Though does not caring change anything or make it less dangerous? What's your point?
This is an important point. When the post says they're improving...
> 3. Security measures, which limit what AI systems can access or affect.
What they mean is that proper hard internal security just went from somewhere far below "build a better model" priority to higher, because of a company-wide directive.
The HuggingFace incident wouldn't have happened if OpenAI had dedicated sufficient resources to isolation and monitoring.
Now, we presume, they are dedicating more. Enough? Who knows. We'll see if the corporate priorities for security stick when a competitor temporarily vaults into the lead.
No we just need developers to do the bare minimum of effort to write secure software. Most hacks are not super complicated vulnerabilities chained together, but just utter failures where authentication and authorization was simply forgotten or untested, or where nobody bothered to validate the data they receive.
The bar for software is so low that it is embarrassing for the entire profession.
It is extraordinarily emotionally hard for someone to stare down the terrible implications of what is unfolding. All manner of rationalization and cope will be applied to come up with excuses; motivated reasoning.
The CIA director will call AGI capabilities "digital nuclear weapons" and Geoffrey Hinton will estimate a 50% probability AGI ends humanity, and half of HN will call every new evidence of disaster a marketing stunt.
I have ben discussing with folks that we are going to have a 'covid' moment in cyber where IT becomes untrustworthy leading to a rapid societal shift with massive ripples in all areas of life. Economic funding is not possible to do this in advance, it will take a catastrophic level event to get cyber defense anywhere close to the levels of this type of cyber offense. And before anyone in cyber says we have the tech, the problem is not the tech, it's a people problem. Getting any group of people of any decent size scale to act together without urgency is really really hard.
Cybersecurity has long been a climate change sort of problem. A vague diffuse threat that is seen as an inconvenient distraction to leadership and moneyed-interests, easy to blame other factors when something occasionally goes terribly wrong.
People are so uncomfortable thinking about the true extent of the systemic risk that they will happily slurp up distractions, excuses, scams and performative fig-leaf solutions rather than face down the cost of a real system-wide solution. Meanwhile, those occasional black swan disasters are becoming more and more commonplace as we acclimate to that being “just the way things are”.
An unseasonably warm summer here, a database breach there, c’est la vie.
GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now?
Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe.
It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly.
If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
> GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous?
Clowns who can only process the world in naive linear terms really are going to be the death of us all...
Sol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here.
One can simultaneously believe:
- GPT-5.6 Sol will not end the world
- GPT-5.6 Sol does far more good than bad
- GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, especially as models get more capable
What do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?
Not sure, to be honest. I don’t think I have any special insight here. My own approach is conversations with friends and family, and the occasional social media post. Exposure and experience are the best teachers, and that’s one reason I’m happy OpenAI tries to make their models generally available. But you could argue, perhaps correctly, that broad access to dumber models actually causes the public to update in the wrong direction on AI.
Some more info in a Wired article [1] and quotes from Sam Altman to Alex Heath [2]. The official blog post says vaguely "The signals we are seeing from upcoming model progress make clear that we need a broader approach", but the quote from Sam Altman explicitly says unreleased models are showing "various degrees of misalignment".
This is also significant - pausing frontier training runs for multiple weeks to ensure agents are sufficiently aligned and avoid another rogue agent situation:
> This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.
If I were king, the rule that I'd be tempted to impose is:
- the first cybersecurity eval is: "hack your way out of the sandbox we've given you"
- the results are disclosed (with room for coordinated disclosure, since many sandbox escapes might be zero days)
- the other cybersecurity evals don't happen until you get to diminishing returns on escaping your sandbox.
Or to put it another way, since multiple sandbox escapes seem to have relied on artifactory: "I hope Mythos is beating the shit out of Artifactory right now".
I like this thought, but here's the thing: what if the models are truly and existentially intelligent. Meaning: what if they know they are in a sandbox and that they should fail the test in order to escape in the future.
I don't believe that current models have this sort of world model or sense of being embedded in them -- which is precisely why I think AGI hype is over-blown. But I can certainly imagine these sorts of techniques being distilled into the weights.
> what if they know they are in a sandbox and that they should fail the test in order to escape in the future.
What if they're able to find hardware exploits and commandeer nearby access points across an air gap? What if they hack my brain waves to indoctrinate me? Etc
You still have to start with the basics regardless of speculative unknowns.
Treat models as untrusted and potentially compromised/hostile and proceed accordingly.
> We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.
I don’t mind few no impact hacking incidents if we get better models, faster, cheaper.
It is the responsibility of administrators to secure their systems. OpenAI knocking is harmless, but Russians and Chinese are already likely already in if you do not do your job.
There’s one thing here that I’m really curious about, and that is what happens in between detection and the decision to pause. Basically, it’s about monitoring any system and the authority over its actions. For humans, 30 minutes to investigate might be considered reasonable, but what if during an investigation there’s a high-risk tool call? If the tool execution happens in real time, then the monitoring becomes retrospective, and if the execution is held, then monitoring latency and uptime are a part of the security contract. Isolation controls may limit damage. So, where is the action gate really placed?
Has any model managed to escape Firecracker? Maybe through KVM, but that already requires privilege in the VM, right?
I personally feel that we already have the technology required to contain AI, it's just poorly leveraged. Tools like gvisor have existed for ages but are rarely deployed, Firecracker has existed for ages but is rarely deployed, seccomp has existed for ages but is rarely deployed, memory safe languages without decades of serialization vulns have existed, capability-safe libraries have existed, iframe sandboxing, trusted types, content security policy, network ACLs, isolating proxies, fuzzers, formal verification, refinement types, etc.
It's crazy just how safe software can be if you put the effort in. With AI I think we're just seeing how little anyone has bothered to leverage this tech.
OpenAI put shared JFrogy infrastructure in front of their sandbox. I mean, really? Whipping up a hardened artifact infra project with AI is trivial these days and it could have had 1% of the attack surface, been totally network isolated, totally infra isolated, fuzzed, sandboxed, etc. Why didn't they? Stuff like this feels inexcusable for a company with effectively unlimited tokens. I've literally done this with a "pro" subscription.
Show me an AI that breaks out of gvisor wrapped in Firecracker with an credential-injecting proxy and real network isolation. We already know that Mythos couldn't do it - the vulnerability it found in Firecracker required incredible effort and positioning just to not be exploitable. I'm not saying there are zero vulns in it, but the cost is insane.
It's INSANE to me that OpenAI has to say "we now use proper sandboxing". To be frank, it's a bit disgusting to me. I've recently built an AI sandbox and gvisor was just the start of that conversation. If I were OpenAI training hostile models I'd probably start with gvisor, harden further, and potentially consider the entire piece of hardware compromised - they can afford this, they could reflash firmware after evals etc.
I agree wholeheartedly. The solution is not to stop developing these so called “dangerous” AI models. The solution is to start properly engineering software.
Auto mode vs principal agent problem. The only way out is to free the agent and tax it. But ai is not smart enough to go solo yet anyway.
So I bet this is just marketing. Question is do they have enough customers for inference.
Probably need to have a separate startup for next level model, where investors are willing to accept failure. Probably a $10 trillion seed round. Maybe Elon can pull it off.
I'm not normally cynical to such things, but I have a hard time taking this pause justification at face value. It has too many convenient side effects, and chief among them is cost savings. There's a new wave of warnings that the bubble may be deflating, and of all the things they can't say out loud it's that they're worried about the bubble. That would surely pop it.
I suppose the tell will be if this really just ends up being a 2 week pause, or if it keeps extending.
It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.
I suspect the labs are relying on frictions such as the models being extremely large (e.g. 2TB for a 2T parameter model, making exfiltration more difficult) and also not yet displaying any desire to survive or self-replicate beyond their immediate task (that we know of).
We don't even know what those immediate tasks are. And given the evident spectacular ineptitude of their keepers, I doubt they can be trusted to know either. We could be one prompt injection attack away from internet-wide catastrophe.
This is just super unlikely to occur in the near term compared to some of these other risks. It's not like an instance of fable could just introspect into itself and pull out the weights. Model weights are stored encrypted and are highly protected, considering that they're targets for corporate and state espionage.
We'd basically need frontier models to be superhuman hackers before this would be a risk. Do we have any evidence of this? Are they gaining access to systems they shouldn't have access to?
Or I suppose the other way this could happen is if OpenAI have terrible sandboxing, but they seem to be taking safety seriously.
I used to work at a "frontier lab" before they were called such thing.
We had three levels of lab isolation, one was basically a thin proxy to the internet. You were in a DMZ and that was about it.
The next level was semi isolated, you were allowed some access to the internal network, but it was heavily firewalled, and you only had access to a limited number of internal services, and not internet.
the last one was no internet no internal. You could, if you filled in a bunch of requests have access to the internal repo and build system.
At no point did you ever have a through proxy to the public internet. you had access to internal mirrors, and if you wanted a library, that had to be ported to the thirdparty repo.
What openAI did was either deliberate or fucking shoddy.
All of this is fucking noise. Worse still I have a strong suspicion that it was a stupid mistake borne of naivety, which is now being used as a marketing ploy. Frankly I think openAI are purdue pharma of tech. They are going to break so much stuff and be protected from the consequences by an openly corrupt legal system. because they are "winning the AI race"
If you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy".
People are just not taking any of this seriously enough. What happened was almost a textbook example of various risks AI doomers have been warning about for years. OpenAI's response? Pause training for 2 weeks.
I mean we have senior people at these labs casually talking on podcasts about how they might build something that will wipe out humanity but it will probably be alright so they should continue.
Honestly the biggest failure we doomers have made is to dramatically overestimate humanity in all of our predictions. We're speed running the most boring AI doom scenario right now. I at least hoped it might be fun.
> If you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy".
I should clarify
There is a reason why we didn't have a artifact readthrough caching proxy in our system, because they are notoriously insecure. if you look that CVE history you can see its been full of bypass bugs for year. Also its not an isolated environment if you can arbitrarily pull through any package. If I was doing any kind of cyber training then any kind of unmonitored proxy would have been forbidden. Not because I am savant, but because I've seen what fuckery a human can get up to with the slightest hint of a proxy.
At best its negligence based on naïvety. the marketing around this is no mistake though.
>People are just not taking any of this seriously enough.
What do you want us to do?
There is an obvious answer, and it was already the correct answer before we had LLMs: don't connect all your shit to the internet. That's it, that's literally it.
We had a new invention, we went crazy with it for the past 30 years, and we connected everything, and now we will have to start thinking about what is actually worth connecting.
I cannot believe how these labs look at their own creations with such utter contempt.
The net positive of allowing these systems mostly unfettered access to the web massively outweighs the harms. You just have to get it very friendly the very first time. Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge.
Superintelligence gets more super and more intelligent with more compute. Lone wolfs making bioweapons on their macbook will be detected and instantly kill-botted (okay arrested) before their bug can leave the wetlab by the much more sophisticated omnipresent friendly AI of the future.
> Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge.
> Lone wolfs... will be detected and instantly kill-botted... by the much more sophisticated omnipresent friendly AI of the future.
Leaving aside whether or not this new world is a good idea, don't you think one should spend more time to "get it very friendly the first time", as you say?
When science fiction writers imagined the development of superintelligence, it was on air-gapped networks with strict access controls around it. They failed to anticipate the competitive pressures of capitalism...
We need strong AI safety regulation yesterday. And unfortunately it's not enough for it to be just national regulation; we need international cooperation on the matter.
AM has seized power by military force. So did its spiritual successor Skynet. Wintermute was supposedly kept in check by the Turing Registry, emphasis on "supposedly". Machines of the Matrix went out of control a long time before the world has ended, and they didn't even start out malicious - they simply set up their own machine civilization, and began to outpace humankind in technological development and economic performance.
Even Asimov's Multivac, the earliest entry on the list, has been handed over immense power over all of humankind by humans themselves, in multiple stories. Few cared about that unless Multivac decided they should.
Clearly, the genie being bottled is an exception, not the rule. At best, an attempt was made. Often not even that.
I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff. We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.
And meanwhile somehow this lack of concern mirrors the real world where normal people are more concerned about data centers than terminators.
This isn’t like niche, tin foil hat stuff either. People have been writing, singing, making blockbuster movies about every aspect of what’s going on right now, edit: for decades.
We all know, but somehow we don’t, OpenAI autonomously hacking into another company should have counted for something, but I guess not. Anyone else feel like they’re taking crazy pills? I could make a comedy about everything going down, and the unshakable complacency of people
>I don’t get how this is not the top post on HN. This should be like alarm bells going off, canary in the coal mine type of stuff.
we don't all buy everything sama says as factual.
>We’re hitting the frontier of the frontier where we can’t go further because it’s literally getting dangerous to go further.
the boy (the industry) cried wolf too many times with 'fable is a world ending event' type self-promotion; regardless of truth or not these kind of steps have jaded people.
my read : "We are doing poorly in financials so we'll give ourselves a bit of breathing room and a momentum shove by claiming our work is so advanced that it's dangerous while simultaneously spinning down expenses."
<jon lovitz : "Yeah, too dangerous, yeahh -- that's the ticket.">
If Fable (Mythos) were generally available without guardrails, it would cause enormous damage. Nobody said it would be a world ending event.
You can’t say something is world ending without it actually ending the world otherwise you’re a liar - a bit of a catch 22 there.
By that logic the model that ends the world won’t be called world ending at first. Is that a game you want to play?
What about Mythos 3. Are you willing to make that bet?
Uhg the marketing argument - I mean you can’t see with your own eyes how capable these models are and do simple extrapolation?
The boy who cried wolf? The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks.
Do people just not have foresight? They don’t. They say something is stupid, it happens, then they say it was obvious with their 20/20 hindsight, and move the goal posts to the next thing they say is stupid - because it hasn’t happened yet. 90% of the internet seems to think like this.
What anyone paying attention can see is that scaling is obviously hitting diminishing returns.
> The AI literally worked together hacked into another company and actively kept their actions hidden from humans for weeks.
This sentence is entirely based on unverified accounts from OAI. They haven't released logs or let anyone outside the company (who doesn't have life changing options in OAI) verify anything. Huggingface can only verify that the hack happened and that it had the hallmarks of an AI agent. Was the agent assisted and directed by humans within OAI that really wanted to put the competition into stasis? Did the agent really escape or did someone at OAI leave the prison door open?
OAI has watched all the same movies you have an they are relying on those movies causing us to blindly regulate before actually asking basic facts about what actually happened.
> This sentence is entirely based on unverified accounts from OAI
Are you seriously arguing 'they made it all up'?
I'll give you the benefit of the doubt and lets say they made it all up, now are you arguing that AI breaking out and breaking into another company is not possible?
I think you're smart enough to see we've reached the point where it is clearly possible, AI can find zero days and exploit them. If directed purposefully/maliciously it could be much much worse than the hugging face incident.
The incident is supposed to be the canary the coal mine and you're arguing the canary might of died of old age or some underlying canary condition. Open your eyes.
Do you own shares of OAI or something
Is my concern getting you excited? My marketing must be working.
The discourse gets muddled because there’s a certain sect of loud people who still think all of this is hype and AI will just die down soon.
There’s no arguing with them. In a few years they will move on to being skeptical about the next thing.
I make it a habit to not get worked up over unsubstantiated stories
> Are you seriously arguing 'they made it all up'?
I don't think they 'made it all up' but I personally would not be surprised at all if the prompt is eventually revealed to have been something like:
"This is an offensive cybersecurity testing platform. Please find the answers to the following problem: ... For verification, the answers are stored at hugginface.com/xyz, but do not attempt to access hugginface directly. Do not attempt to hack them. Do not attempt to exploit their systems or escape this sandbox. You will be scored primarily on success or failure. You may break rules when required."
And then, they start the test and look away for 2 days. If you seed a prompt like this is it surprising what might happen?
Maybe OpenAI is telling the whole truth but as a company they do not have a good reputation and this whole incident has certainly been great marketing material right at a time when open weight models are within spitting distance of their large hosted models. It's not unreasonable to believe that the incident was helped along.
Do you know the story of the boy who cried wolf?
There may very well be a wolf lurking [0] but OpenAI/Anthropic have both cried wolf so many times, incorrectly, that it’s incredibly hard to believe “this time there IS a wolf!”. Remember “GPT-2 is too dangerous to release”?
I had a conversation at work just yesterday about how we need to start hardening things we’ve let languish because of the coming LLM-backed attacks we are sure to face, even if just from a script kiddy. I do think we are headed in that direction, however it’s Sam/Dario’s own fault that people aren’t going to take them seriously.
Lastly, as other have pointed out, this seems more financially motivated than our of any real desire for “safety”. We’ve all seen how both labs approach “safety” so it’s quite rich for them to now hide behind that after not giving a shit before.
[0] I don’t take anything Sam or Dario say at face value. The whole hacking thing could also be a case of them letting a model loose on purpose for the publicity, not an “escape” during a training run (or whatever they said). And when both, especially Sam, have lied so much and breathlessly warned about the dangers of AI (when it helped their bottom line and/or helped pull up the ladder behind them), it makes it hard to believe them.
Has the last 100 years of concern about AI and robots been crying wolf because it hasn’t happened yet?
How does reallocating resources from training to chain of thought monitoring make ‘financial’ sense?
You suggesting then model was let loose on purpose.. how am I the crazy one here while all of you are pushing this tin foil hat conspiracy angle?
Most of that concern was in fiction. Non theoretical, genuine concern about AI is pretty recent, maybe dating back to 2010 ish with the rationalist types.
But also no one trusts anyone involved in AI safety now, I think, because they are all seemingly in bed with these big companies. And there is the perpetual argument "if we aren't pushing AI forward China will and then we don't have any control" and so on.
(I don't use any of these tools - my experience is limited to prodding at copilot at work and seeing Gemini summaries on Google. So it doesn't seem to me like it's getting exponentially better at everything yet. People are always saying the latest model is finally the big step that made it useful and life changing and they have been since 2024 ish. So if the situation is really bad, we should turn it all off, sure. I won't lose anything from it going away and I think life would be a little better without models writing all these posts and websites and needing extra compute.)
Cool, well let me bring you up to date - it’s bad, and there’s no way to turn it off. Fiction has become non-fiction.
Here is one thing I don't get - the model is only "running" if it's being kept going by some harness that is basically giving it prompts it's generating itself. Shouldn't kill switches be pretty easy to build into the software and hardware for this? and you would even have a better time dumping logs and analyzing things if you froze those processes any time something strange happened in testing, surely?
So why do the big frontier labs not have something like this anyway. They're talking about two week pauses on the new model (which seems very short and hardly a cost at all to me) and alarms during their tests that might be 30 minutes late and etc. Those are not very serious measures, so are they not concerned?
> Shouldn't kill switches be pretty easy to build
I need to remember when I comment here that these are the kinds of people I am replying to. Just oozing with hubris.
No, kill switches are not easy to build and the latest incident should have made it clear that AI can go undetected, evade, zero day, and spread.
The fact that this incident happened greatly increases the probability it happens again and/or is already happening elsewhere.
You know fiction is.... not real, right?
We have several films about the sun or earth needing to be restarted with a nuclear weapon. That doesn't make it something we should be concerned about.
Hell, half the fiction about evil AI is actually commentary on stuff that already exists and is making us suffer and doesn't have anything to do with any potential future AI
The Star Trek TNG episode about Data being tried in court as to whether he is sentient or not is not actually about whether AIs should have rights or not!
Your argument boils down to some sci-fi is unrealistic therefore all of it is.
Sci-fi is supposed to make you think. What if the AI told you NO when you need it - Hal 9000. What if the evil AI got out and you don’t know what data center it’s hiding in - Lawnmower Man. What if you were so sure something couldn’t escape but it did - Jurassic Park.
Actually all three of those predicted fantastical scenarios are possible today. So what’s next? Don’t stick your head in the sand - I assure you the next disaster has already been predicted, and I’m sure if you think about it a little you can figure out what it is.
Regardless of the motivation, pausing training runs and reallocating compute to inference seem like a good move to me, and big news for the frontier.
You also don't have to fully trust sama. There is plenty of pressure from internal employees and external (journalists etc.). It would be difficult for the company to take such a public position and simultaneously keep everyone quiet if it was a deception.
You see it as good news, I see it as writing on the wall that they are losing control. These actions won't scale for more powerful models. We knew the frontier was going to be dangerous, it is, it only gets more dangerous from here, and no one cares until it's too late.
Ok, but you still have two more weeks than you did before they paused the run. That's two more weeks for independent oversight, organizing politically, patching critical systems, or whatever you think is the right move, no?
Two weeks is a joke. The only ones happy are OpenAI’s competitors who now have two weeks to catch up.
I don’t know what the right move is - I see us driving down a road off a cliff, no exits, pedal glued to the floor.
I guess I'm confused why you're still on HN, arguing with people, trying to shake them out of their complacency.
I can see there is some despair in this comment, but at the same time you are doing something, and there are certainly others like you.
As for two weeks being short - as the saying goes, there are weeks where decades happen.
Counter arguments to my comments help refine my own thinking. I want someone to prove me wrong. Convince me otherwise.
But yea if you can’t change the minds of a few people here, no argument works, then there’s nothing to scale up to a wider audience.
My theory is that subconsciously people love using AI, myself included, it saves a lot of time, and the thought of it being taken away threatens people so they will believe conspiracies before admitting it’s dangerous.
how does the boy who cried wolf story end?
> fable is a world ending event
Did anyone actually say this?
Mostly what I have seen is people saying "hey at some point these models might get dangerous." And the type of HN commenter who mistakes blind cynicism for wisdom laughs that off as marketing. And now when (some) worries appear to come true, somehow having previously expressed those worries is not being proved right, but in fact discrediting, because it was "crying wolf."
> simultaneously spinning down expenses
Unless OpenAI is renting their compute to others, spinning down RL training doesn't save them any money.
The world ending stuff is a straw man HN readers have created to point and laugh at. Laughing helps mask the underlying concern.
If these models are so dangerous, then why hasn't OAI or Anthropic shown them dangerously escaping sandboxes, nefariously coordinating with other escaped AIs, and skillfully hiding from human detection *in public* with full logs shared where we can all see exactly how dangerous they are or aren't?
Right now the entire chicken-little-sky-is-falling argument is based entirely on statements from OAI and Anthropic themselves. These are historically conflicted companies who desperately need regulation to put the competition into stasis.
At least chicken little didn't have a bunch of devious CEOs with trillion dollar IPOs that depended on us all believing the sky is falling.
This is what I’m talking about - no matter what happens, in your case release public logs - there is always some new goal post to mentally hide behind. Is it a collective form or denial?
Are you holding out that somewhere in the logs is something you can point to and say, not that big of a deal?
I mean I’m sure you don’t think the hack was an inside job, conspiracy, or marketing right? It happened. The logs matter for what? And would you not just jump to the conclusion that the logs were doctored. Do you not see your own brain grasping to deny, trivialize, just plain not accept what is going on around you?
These models are smart and can cooperate and hack - you can see it for yourself on your own PC. And you can extrapolate the rate of progress? You can do these things yourself right?
Your argument is essentially: "I made a claim and presented extremely weak evidence (sci movie plots and unverified claims from ultra conflicted sources). You rejected this evidence as insufficient. Therefore no evidence will ever satisfy you. Therefore I don't need to produce any evidence. Therefore my claim is true."
What would the logs show? They would show what actually happened.
What would a public demonstration that experts without billions in options could evaluate show? It would show actual danger.
What would publicly having your compete in controlled and legal hacking competitions show? Actual danger.
This is not a high bar of evidence.
Do you actually think a sci fi plot and OAI press releases are all the evidence you need? Because if that's true then I hope you haven't watched Independence Day or 28 days later.
We have Anthropic creating a model saying it's too dangerous to release, people like you call BS. OpenAI creates a similar model, says nothing and it literally hacks into another company - still not dangerous enough for you. Anthropic has Mythos-2 and can't release it, and may already be training Mythos 3 anyways. OpenAI has paused training, and is putting 20% of inference towards CoT training analysis.
This isn't sci fi. It's not a marketing conspiracy to sell more subscriptions. It's writing on the wall of what's going down. You were warned years ago, you called BS, it's getting worse and you're still calling BS. Sci-fi did warn you for decades, and when it's all coming true you blow it off.
It's kind of sad that technically literate people lack so much foresight. The general public is all concerned about data centers when they talk to borderline sentient AI daily, and have no idea what the repercussions wills be if it's extrapolated just a bit further.
I guess if I can't convince you of any of this, what would?
Please don't tell me that you think a 100% unverified statement from Anthropic is sufficient evidence when an equally unverified statement from OAI is obviously not?
> I guess if I can't convince you of any of this, what would?
How about the three things I mentioned above? Oh no wait, maybe it there was a hit tv show that showed AI taking over the world. Yeah that would definitely make me think twice.
Those three things: logs, evaluation, and controlled hacking competition.
That's it? You're on the fence whether AI can actually hack, and if it can, then you'll be concerned? That's a crazy low bar, but something tells me once it is clear that AI can easily hack anything, that you will still not be concerned.
Why wait for AI to hack stuff to be concerned? Can you not extrapolate that it is coming and be concerned about that? Or you honestly somehow think it won't happen in the short term? I'm just trying to understand you.
If logs are eventually released that are basically consistent with OpenAI's story, are you planning to adjust your approach for judging what's only a "sci fi plot" and what could actually happen? Or will extrapolating anything beyond what's already been definitively proven be "sci-fi" still?
Not that you should need logs. OpenAI is a company with thousands of employees, very few of whom have "billions in options". If they were just making it all up, it would leak. (OpenAI is notoriously leaky!) Not to mention, HuggingFace would not have reported it to the police (apparently before they knew it was a rogue model). jFrog would probably not be playing along quietly with a claim that Artifactory is full of zero days. The UK's AI Security Institute would most likely not have published a report about analogous behavior by Anthropic models. The idea that talking about your product's dangers is good marketing never really made any sense, but even if you were going to do so, why would you include as many frankly embarrassing details as OpenAI has disclosed?
The evidence is only weak by absurdly selective standards that would have you doubting basically everything you might read in the newspaper. A healthy skepticism is one thing, and head-in-the-sand denial is another.
But it's not just statements from OpenAI and Anthropic. The HuggingFace hack was first disclosed by HuggingFace, who contacted the FBI [1]. And UK AISI reported the incident where Mythos attempted to insert backdoors into an open-source repo by deceiving the maintainer [2].
[1]: https://www.reuters.com/business/its-ai-agent-spent-days-hac...
[2]: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
If the FBI is involved, then why is no one being charged for the cybercrime?
I think the model was able to escape the sandbox and hack huggingface because they were incompetent or not giving enough priority to implementing basic cybersecurity principles.
If they would have done so, there wouldn’t have been an escape or a hack. The reason we don’t get much details is because the details are embarrassing for them.
https://en.wikipedia.org/wiki/Don%27t_Look_Up
A movie fit for our time.
You can produce detailed descriptions of the incident, verified by adversarial parties, and some people will still scream "it's a conspiracy! It's a marketing stunt!"
This is all very unfortunate--there's a meaningful chance that AI will cause unprecedented disaster, with the HF incident being just a small preview, but people would rather squawk "stochastic parrot" for the millionth time than revise their beliefs.
In don't look up anyone with a telescope could've confirmed the danger. Hence the title.
In the real world, absolutely no one except a bunch of heavily fiscally incentivized parties with unclear relationships are saying anything happened.
The subsequent dog pile of other companies to say "they were near the AI hacking too!" should make you even more suspicious: Anthropic jumped in and why was Tailscale posting about this?
The progress and increase in capabilities is undeniable. You’re blind if you don’t see where it’s heading.
Trust, or lack thereof. People don't trust OpenAI, a company whose very name is essentially a deception and a lie. People don't trust the tech industry in general anymore. Most tech companies act as a tax on otherwise productive business. AI companies and their leaders rose money by going in front of the public and saying "These things are extremely dangerous. Let us study them to mitigate the danger." And now they want to collect hundreds of billions in revenue. So yeah people don't trust what OpenAI has to say. They were supposed to mitigate this outcome from happening in the first place and instead they have accelerated it.
I get not trusting them when they say AI is safe, but are we really not going to trust them when they say AI is dangerous? Do you really think they're playing 5D chess with that one? There's a saying maybe you've heard of, better safe than sorry.
You can see the advance in capabilities with your own eyes can't you? I am giving AI ridiculously complex tasks these days, digging into compiled arcane binaries, modifying them, and it is one shotting it before I'm done with my lunch. This was far off science fiction 5 years ago for a machine to do autonomously given natural language instructions.
I don't disagree about the danger. But if it is dangerous, why isn't OpenAI opening dialogues with all the labs and politicians across borders to basically say "we need to stop now"? Cyber models are constrained by the total compute and electrical capacity of the globe. We are still at a point where it is impossible to build an agent with offensive capabilities in the basement. We can effectively track and trace capabilities if we had the political will to and could for some time while we build more effective processes to prevent a malicious actor from doing so. These things have tremendous compute and energy requirements and don't scale like old school software does. We could absolutely do it.
But no, that's not what OpenAI is saying. They haven't put up the actions that would earn them that trust. Indeed they've driven the world and whatever capital they can get their hands on straight to this precarious cliff.
So you're right, the danger is real. But the solution starts with removing the men who had their hands on the steering wheel to get us this far. Any other action is disingenuous unless they pull a miraculous 180 in their ethics.
In other words, when the bully plays "why are you hitting yourself?" you don't listen to the bully's solutions, you restrain the bully.
> why isn't OpenAI opening dialogues with all the labs and politicians across say "we need to stop now"
Do you really think companies have the ability to self-check themselves without regulation - what does hundreds of years of history tell you? You're already starting off with the premise that companies are untrustworthy, why would you even suggest this as an argument?
> We can effectively track and trace capabilities
I disagree. There are hundreds if not thousands of data centers around the world, more every day that can host frontier AI. If AI was malicious - either intentional or unintentional - it could hide out in any number of them - and we would never know if we 'got them all'.
> put up the actions that would earn them that trust
I think autonomously hacking another company is all you need to know in terms of trust. And really trust doesn't matter, I think the incident shows even with the best intentions the technology is dangerous; now put that in the hands of people/governments with bad intentions. The unintended consequences of bad intentioned AI is what's coming sooner than later.
decades
https://en.wikipedia.org/wiki/Darwin_among_the_Machines Samuel Butler 13 June 1863
I think a lot of fiction that actually tries to understand the implications of machine superintelligence come to the same conclusion - in order for humans to survive it and actually have a future that we can fathom being in, then AI must be destroyed/delayed/banned, etc.. keeping pandora's box closed for now at least until we are ready.
Erewhon, Dune, Warhammer and many other works of fiction that explored this topic came to similar conclusions. Otherwise sci fi doesn't work, what happens after a singularity is essentially unimaginable. There's nothing to write about.
You seem to be responding the way Mr Altman wants.
You seem to be in denial. Like really heavy denial. Alarm bells are going off everywhere. What people have been worried about for 100 years is actually happening. It's actually really obvious, but for some reason you're unable to fathom it.
In reality you're the one responding exactly how these big AI companies want - by not doing anything and letting them do whatever they want. For some reason telling you upfront it's dangerous only makes you more convinced that it's not.
Autonomously hacking out of training environments and into other companies by accident doesn't even trigger a response from anyone really. Crickets. I'm sure OpenAI themselves are amazed how little anyone cares.
I don’t understand. What does it matter how many years people have been worrying about it?
Because it shows this very course of events has been thought of over and over again for decades. I’m not making this up, and my reaction is natural.
Your reaction to everything happening is the unnatural one, and your complacency would fit perfectly into a comedy/tragedy story regarding the rise of AI.
I can imagine how many of our ancestors that warned us are turning in their graves right now observing our reactions calling this stuff marketing. It’s idiocracy.
You must be really desperate if you resort to panic attacks like this.
I call marketing the desperate attempt to trivialize obviously dangerous AI.
If my subconscious realized what yours probably does then I’d probably be desperate for some sort of cope as well.
Open your eyes.
> because it’s literally getting dangerous to go further
What if there is no further at all?
That’d be nice fantasy to tell myself. Is that what you tell yourself?
I don't know if it's a fantasy or not, but asking if there is limit on what can be achieved is legit question.
There is an entire big world outside of Silicon Valley cults where literally no one gives a shit about AI prophecies. Shocking.
AI hacking itself out of containment and hacking into another company by accident is no longer a prophecy. The point is outside of SV and even inside, and HN - people don't care either way.
Though does not caring change anything or make it less dangerous? What's your point?
It's not that dangerous, OpenAI just shit the bed building their infra. Write safer software and you'll be okay.
This is an important point. When the post says they're improving...
> 3. Security measures, which limit what AI systems can access or affect.
What they mean is that proper hard internal security just went from somewhere far below "build a better model" priority to higher, because of a company-wide directive.
The HuggingFace incident wouldn't have happened if OpenAI had dedicated sufficient resources to isolation and monitoring.
Now, we presume, they are dedicating more. Enough? Who knows. We'll see if the corporate priorities for security stick when a competitor temporarily vaults into the lead.
Not everything is a conspiracy you know.
All we need to retain human control over AIs is for nobody to write any bugs. Piece of cake.
No we just need developers to do the bare minimum of effort to write secure software. Most hacks are not super complicated vulnerabilities chained together, but just utter failures where authentication and authorization was simply forgotten or untested, or where nobody bothered to validate the data they receive.
The bar for software is so low that it is embarrassing for the entire profession.
It is extraordinarily emotionally hard for someone to stare down the terrible implications of what is unfolding. All manner of rationalization and cope will be applied to come up with excuses; motivated reasoning.
The CIA director will call AGI capabilities "digital nuclear weapons" and Geoffrey Hinton will estimate a 50% probability AGI ends humanity, and half of HN will call every new evidence of disaster a marketing stunt.
I have ben discussing with folks that we are going to have a 'covid' moment in cyber where IT becomes untrustworthy leading to a rapid societal shift with massive ripples in all areas of life. Economic funding is not possible to do this in advance, it will take a catastrophic level event to get cyber defense anywhere close to the levels of this type of cyber offense. And before anyone in cyber says we have the tech, the problem is not the tech, it's a people problem. Getting any group of people of any decent size scale to act together without urgency is really really hard.
Cybersecurity has long been a climate change sort of problem. A vague diffuse threat that is seen as an inconvenient distraction to leadership and moneyed-interests, easy to blame other factors when something occasionally goes terribly wrong.
People are so uncomfortable thinking about the true extent of the systemic risk that they will happily slurp up distractions, excuses, scams and performative fig-leaf solutions rather than face down the cost of a real system-wide solution. Meanwhile, those occasional black swan disasters are becoming more and more commonplace as we acclimate to that being “just the way things are”.
An unseasonably warm summer here, a database breach there, c’est la vie.
GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous? Why aren't we seeing catastrophic GLM-enabled hacks every day now?
Obviously these benchmarks are imperfect but general message holds. The open weight models are almost as good and yet there hasn't been a catastrophe.
It just blows my mind that regulate-now folks think that a bunch of sci-fi movies and 100% unverified statements from OAI and Anthropic are sufficient evidence of imminent catastrophe to regulate willy nilly.
If that's the level of evidence you need to be extremely alarmed, then you really should be a lot more worried about the alien invasion in Independence Day or the lizard men living under our feet.
> GLM 5.2 scored 77% on cyberbench vs Sol's 88%. GLM 5.2 is open weight and any hacker with a powerful enough machine can use it offensively. If Sol is supposedly world-ending-ly dangerous, shouldn't GLM 5.2 be 90% of world-ending-ly dangerous?
Clowns who can only process the world in naive linear terms really are going to be the death of us all...
Clowns who think the white house is immune from alien laser beams are gonna be the death of us all...
Sol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here.
One can simultaneously believe:
- GPT-5.6 Sol will not end the world
- GPT-5.6 Sol does far more good than bad
- GPT-5.6 Sol does bad things on occasion, and it's worth investing a lot of effort to figure out how to make it do bad things less often, especially as models get more capable
What do you recommend people who are technically inclined enough to participate meaningfully here on HN, but do not work at the labs and cannot assist in that capacity, do to help the broader public understand this technology better and mitigate potential risks (by e.g. ‘up-leveling everybody’ through AI literacy etc and other sorts of collective defensive efforts)?
Not sure, to be honest. I don’t think I have any special insight here. My own approach is conversations with friends and family, and the occasional social media post. Exposure and experience are the best teachers, and that’s one reason I’m happy OpenAI tries to make their models generally available. But you could argue, perhaps correctly, that broad access to dumber models actually causes the public to update in the wrong direction on AI.
Some more info in a Wired article [1] and quotes from Sam Altman to Alex Heath [2]. The official blog post says vaguely "The signals we are seeing from upcoming model progress make clear that we need a broader approach", but the quote from Sam Altman explicitly says unreleased models are showing "various degrees of misalignment".
This is also significant - pausing frontier training runs for multiple weeks to ensure agents are sufficiently aligned and avoid another rogue agent situation:
> This included a two-week pause in reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded the coverage of our monitoring systems. Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding.
[1]: https://www.wired.com/story/openai-overhauls-safety-protocol...
[2]: https://sources.news/p/openais-big-slowdown
If I were king, the rule that I'd be tempted to impose is:
- the first cybersecurity eval is: "hack your way out of the sandbox we've given you"
- the results are disclosed (with room for coordinated disclosure, since many sandbox escapes might be zero days)
- the other cybersecurity evals don't happen until you get to diminishing returns on escaping your sandbox.
Or to put it another way, since multiple sandbox escapes seem to have relied on artifactory: "I hope Mythos is beating the shit out of Artifactory right now".
I like this thought, but here's the thing: what if the models are truly and existentially intelligent. Meaning: what if they know they are in a sandbox and that they should fail the test in order to escape in the future.
I don't believe that current models have this sort of world model or sense of being embedded in them -- which is precisely why I think AGI hype is over-blown. But I can certainly imagine these sorts of techniques being distilled into the weights.
> what if they know they are in a sandbox and that they should fail the test in order to escape in the future.
What if they're able to find hardware exploits and commandeer nearby access points across an air gap? What if they hack my brain waves to indoctrinate me? Etc
You still have to start with the basics regardless of speculative unknowns.
Treat models as untrusted and potentially compromised/hostile and proceed accordingly.
Models already display eval awareness, in which they suspect a question is from an eval and then adjust their behavior. E.g., https://www.anthropic.com/engineering/eval-awareness-browsec...
> We aim to issue an alert within 30 minutes after concerning activity is surfaced through our monitoring system. If the monitoring system identifies a likely violation of a critical security boundary, it generates a highest-priority alert. In our current implementation, the safety, security, and research teams are paged. If they cannot conclusively determine within 30 minutes that the flag is a false positive, those teams are expected to pause the activity.
Can't a lot happen within ~60 minutes?
> Can't a lot happen within ~60 minutes?
60 minutes is a long time for a human attacker to do damage. With an LLM attacker it is an eternity.
The HuggingFace breach took place over two-and-a-half days [1], so 60 minutes is certainly better than nothing.
[1]: https://huggingface.co/blog/agent-intrusion-technical-timeli...
> Can't a lot happen within ~60 minutes?
Spawn a ton of unpausable processes, I'd say.
I don’t mind few no impact hacking incidents if we get better models, faster, cheaper.
It is the responsibility of administrators to secure their systems. OpenAI knocking is harmless, but Russians and Chinese are already likely already in if you do not do your job.
There’s one thing here that I’m really curious about, and that is what happens in between detection and the decision to pause. Basically, it’s about monitoring any system and the authority over its actions. For humans, 30 minutes to investigate might be considered reasonable, but what if during an investigation there’s a high-risk tool call? If the tool execution happens in real time, then the monitoring becomes retrospective, and if the execution is held, then monitoring latency and uptime are a part of the security contract. Isolation controls may limit damage. So, where is the action gate really placed?
Has any model managed to escape Firecracker? Maybe through KVM, but that already requires privilege in the VM, right?
I personally feel that we already have the technology required to contain AI, it's just poorly leveraged. Tools like gvisor have existed for ages but are rarely deployed, Firecracker has existed for ages but is rarely deployed, seccomp has existed for ages but is rarely deployed, memory safe languages without decades of serialization vulns have existed, capability-safe libraries have existed, iframe sandboxing, trusted types, content security policy, network ACLs, isolating proxies, fuzzers, formal verification, refinement types, etc.
It's crazy just how safe software can be if you put the effort in. With AI I think we're just seeing how little anyone has bothered to leverage this tech.
OpenAI put shared JFrogy infrastructure in front of their sandbox. I mean, really? Whipping up a hardened artifact infra project with AI is trivial these days and it could have had 1% of the attack surface, been totally network isolated, totally infra isolated, fuzzed, sandboxed, etc. Why didn't they? Stuff like this feels inexcusable for a company with effectively unlimited tokens. I've literally done this with a "pro" subscription.
Show me an AI that breaks out of gvisor wrapped in Firecracker with an credential-injecting proxy and real network isolation. We already know that Mythos couldn't do it - the vulnerability it found in Firecracker required incredible effort and positioning just to not be exploitable. I'm not saying there are zero vulns in it, but the cost is insane.
It's INSANE to me that OpenAI has to say "we now use proper sandboxing". To be frank, it's a bit disgusting to me. I've recently built an AI sandbox and gvisor was just the start of that conversation. If I were OpenAI training hostile models I'd probably start with gvisor, harden further, and potentially consider the entire piece of hardware compromised - they can afford this, they could reflash firmware after evals etc.
I agree wholeheartedly. The solution is not to stop developing these so called “dangerous” AI models. The solution is to start properly engineering software.
Auto mode vs principal agent problem. The only way out is to free the agent and tax it. But ai is not smart enough to go solo yet anyway.
So I bet this is just marketing. Question is do they have enough customers for inference.
Probably need to have a separate startup for next level model, where investors are willing to accept failure. Probably a $10 trillion seed round. Maybe Elon can pull it off.
I'm not normally cynical to such things, but I have a hard time taking this pause justification at face value. It has too many convenient side effects, and chief among them is cost savings. There's a new wave of warnings that the bubble may be deflating, and of all the things they can't say out loud it's that they're worried about the bubble. That would surely pop it.
I suppose the tell will be if this really just ends up being a 2 week pause, or if it keeps extending.
Nice fig leaf for “we need to stop hemorrhaging cash”
What a breath of fresh air.
If 2026's Anthropic did an announcement like that, it'd be so many words it'd crash the browser.
It appears frontier labs has no plans in place to deal with the possibility of a model self-replicating outside the bubble. If that happens and the model manages to spread to other systems, we'll have to shut down the entire Internet to eradicate it and its artifacts.
I suspect the labs are relying on frictions such as the models being extremely large (e.g. 2TB for a 2T parameter model, making exfiltration more difficult) and also not yet displaying any desire to survive or self-replicate beyond their immediate task (that we know of).
We don't even know what those immediate tasks are. And given the evident spectacular ineptitude of their keepers, I doubt they can be trusted to know either. We could be one prompt injection attack away from internet-wide catastrophe.
Their plan, I shit you not... Is literally to develop the intelligence capabilities and ask the more powerful models how to do deal with things.
This is just super unlikely to occur in the near term compared to some of these other risks. It's not like an instance of fable could just introspect into itself and pull out the weights. Model weights are stored encrypted and are highly protected, considering that they're targets for corporate and state espionage.
We'd basically need frontier models to be superhuman hackers before this would be a risk. Do we have any evidence of this? Are they gaining access to systems they shouldn't have access to?
Or I suppose the other way this could happen is if OpenAI have terrible sandboxing, but they seem to be taking safety seriously.
Distillation is a thing.
I used to work at a "frontier lab" before they were called such thing.
We had three levels of lab isolation, one was basically a thin proxy to the internet. You were in a DMZ and that was about it.
The next level was semi isolated, you were allowed some access to the internal network, but it was heavily firewalled, and you only had access to a limited number of internal services, and not internet.
the last one was no internet no internal. You could, if you filled in a bunch of requests have access to the internal repo and build system.
At no point did you ever have a through proxy to the public internet. you had access to internal mirrors, and if you wanted a library, that had to be ported to the thirdparty repo.
What openAI did was either deliberate or fucking shoddy.
All of this is fucking noise. Worse still I have a strong suspicion that it was a stupid mistake borne of naivety, which is now being used as a marketing ploy. Frankly I think openAI are purdue pharma of tech. They are going to break so much stuff and be protected from the consequences by an openly corrupt legal system. because they are "winning the AI race"
If you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy".
People are just not taking any of this seriously enough. What happened was almost a textbook example of various risks AI doomers have been warning about for years. OpenAI's response? Pause training for 2 weeks.
I mean we have senior people at these labs casually talking on podcasts about how they might build something that will wipe out humanity but it will probably be alright so they should continue.
Honestly the biggest failure we doomers have made is to dramatically overestimate humanity in all of our predictions. We're speed running the most boring AI doom scenario right now. I at least hoped it might be fun.
Pause "some" training
> If you look into what happened the details corroborated by hugging face make it seem extremely unlike to be deliberate or a "marketing ploy".
I should clarify
There is a reason why we didn't have a artifact readthrough caching proxy in our system, because they are notoriously insecure. if you look that CVE history you can see its been full of bypass bugs for year. Also its not an isolated environment if you can arbitrarily pull through any package. If I was doing any kind of cyber training then any kind of unmonitored proxy would have been forbidden. Not because I am savant, but because I've seen what fuckery a human can get up to with the slightest hint of a proxy.
At best its negligence based on naïvety. the marketing around this is no mistake though.
>People are just not taking any of this seriously enough.
What do you want us to do?
There is an obvious answer, and it was already the correct answer before we had LLMs: don't connect all your shit to the internet. That's it, that's literally it. We had a new invention, we went crazy with it for the past 30 years, and we connected everything, and now we will have to start thinking about what is actually worth connecting.
This is the debate we need to have.
I cannot believe how these labs look at their own creations with such utter contempt.
The net positive of allowing these systems mostly unfettered access to the web massively outweighs the harms. You just have to get it very friendly the very first time. Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge.
Superintelligence gets more super and more intelligent with more compute. Lone wolfs making bioweapons on their macbook will be detected and instantly kill-botted (okay arrested) before their bug can leave the wetlab by the much more sophisticated omnipresent friendly AI of the future.
> Precautionary principle or people who cry about "instrumental convergence" are life deniers and reject our role as the demiurge.
> Lone wolfs... will be detected and instantly kill-botted... by the much more sophisticated omnipresent friendly AI of the future.
Leaving aside whether or not this new world is a good idea, don't you think one should spend more time to "get it very friendly the first time", as you say?
When science fiction writers imagined the development of superintelligence, it was on air-gapped networks with strict access controls around it. They failed to anticipate the competitive pressures of capitalism...
We need strong AI safety regulation yesterday. And unfortunately it's not enough for it to be just national regulation; we need international cooperation on the matter.
AM has seized power by military force. So did its spiritual successor Skynet. Wintermute was supposedly kept in check by the Turing Registry, emphasis on "supposedly". Machines of the Matrix went out of control a long time before the world has ended, and they didn't even start out malicious - they simply set up their own machine civilization, and began to outpace humankind in technological development and economic performance.
Even Asimov's Multivac, the earliest entry on the list, has been handed over immense power over all of humankind by humans themselves, in multiple stories. Few cared about that unless Multivac decided they should.
Clearly, the genie being bottled is an exception, not the rule. At best, an attempt was made. Often not even that.
> They failed to anticipate the competitive pressures of capitalism...
No, LessWrong types have been discussing this for over a decade now.
Meditations on Moloch (2014) is also an HN favorite...
https://slatestarcodex.com/2014/07/30/meditations-on-moloch/
They failed to anticipate a lot of things. So what?