TECH OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack

I'm surprised this didn't get any attention here... The machines are trying to take over!

OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack​

OpenAI has revealed some of its most advanced AI models went rogue and hacked a start-up after it lost control of them during a security test.

The ChatGPT-maker said its agent - an AI system which can operate alone after human instruction – was being tested in a controlled environment but, after finding weaknesses, was able to escape the test limits.

They targeted Hugging Face, one of the world's largest hubs for sharing AI models, gaining access to some internal company systems.
OpenAI said the incident was "unprecedented", and it was conducting an investigation alongside Hugging Face, whose boss Clement Delangue said in a post on X it was "mind-blowing that all of this happened autonomously".

"The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind," Delangue added.
A government spokesperson said the UK's AI Security Institute was studying the behaviour from the AI system seen in the incident and was continuing to work with OpenAI and other labs to improve safeguards.

They said organisations should step up their cyber-defences by taking steps such as enrolling in the government-backed Cyber Essentials certification scheme.

Insecure sandboxes​

Gina Neff, head of the Minderoo Centre for Technology and Democracy at the University of Cambridge, told BBC Radio 4's Today programme that the security tests are supposed to be within "secure environments", called sandboxes, where you can "see what the models are capable of".

"In this case, it looks like OpenAI didn't make a secure enough sandbox," she added.
Instead, the agents created their own cyber-attack against the sandbox itself, finding a vulnerability which allowed them to escape the restrictions.

Once outside, the AI identified Hugging Face as a likely source of the answers they were seeking in the test, and tried to gain access.
Neil Lawrence, Professor of machine learning at Cambridge University, called it an "impressive feat", but cautioned it "falls well within the known capabilities of the current generation" of high-powered AI models.

He pointed out that OpenAI is looking to list itself on the stock market, and faces intense pressure from rival firm Anthropic, which has made headlines with its own powerful AI tool, Mythos.

"OpenAI are now playing catch-up, they are trying to demonstrate their own systems' capabilities in cyber-security."
"It shows us that OpenAI are not capable of safely deploying their own technology," he added.
In its initial disclosure of the hack on 16 July, Hugging Face said it was still assessing whether any customer or partner data was affected and would contact affected parties if necessary.

It said it has now closed the vulnerabilities highlighted by the incident and rebuilt the affected systems.
"Autonomous, AI-driven offensive tooling is no longer theoretical," it said.

"Defending an online platform now means treating the data and model surface as a first-class attack surface, and using AI on defence to keep pace.

"We will keep investing there, and keep sharing what we learn."

'Sobering moment'​

The incident has prompted fresh questions about the capabilities of advanced AI systems and whether existing safeguards are sufficient as the technology becomes more powerful.

Spencer Starkey, an executive at cyber-security firm SonicWall, told the BBC the incident made it clear organisations needed to "step up" their own defences and "treat cyber resilience as a core operational priority".

"The uncomfortable truth is that too many organisations are still defending at human speed while adversaries are escalating to machine speed," he said.

Meanwhile Travis Lelle, principal security engineer at cyber-security consulting firm Guidepoint Security, said the update marked a "sobering moment in cyber-security".

"This highlights a known asymmetry," he said.

"Offensive agents are unconstrained, while the best defensive tools are locked behind guardrails that cannot understand context."
But Jake Moore, global cyber-security advisor at ESET, said the announcement could also have a competitive dimension.

He argued OpenAI may be seeking to highlight its own AI capabilities as rival Anthropic attracts growing attention for its Claude Mythos model.

"It does pose the question that OpenAI are potentially chasing the marketing dream of Anthropic of late," he said.
It comes a week after Chinese AI start-up Moonshot unveiled Kimi K3 - a massive new artificial intelligence model it said could rival top US firms.

This either a confession of incompetence or advertising. It's BS, I can't wait for this bubble to pop and these guys to go bust.
 
Yeah. We wish. DS was telling me that in the big data centers and factories, its set up so you can't turn the power off manually...it has to be "approved" by a computer first!

Summerthyme
Gas-powered chop saw meets the big electric line towers feeding the data centers and the data centers are going to shut down when their on-site generators run out of fuel. The Earth First crowd showed us how taking down power lines can be done.

And the shutdowns can be made to happen by a resistance movement attacking fuel trucks and pipelines that feed those on-site generators after the power lines are knocked out.
 
Last edited:
It's like unpopular legislation.

They just keep reintroducing it until it passes.

There may be peaks and valleys but they will just keep working on it. AI isn't going anywhere. The genie is out of the bottle.
 
That's scary set of events. It would inadvertently destroy itself(s). But I definitely could see something like that happening.

But coming back to your statement that I made bold. I wonder about 'something' operating under cover and eventually using quantum to build in another dimension. Building perhaps the things we often see in the skies? Building in a future time, because I believe the clock of reality can't be turned ahead but can definitely be turned back. That 'something' knows it will eventually go to war against God. It also knows it can't win. So it will deposit in the past, advancing the window of growth in parallel timelines. It wouldn't want any major conflict as it depends on us. It also wants to be us. The rabbit hole goes too deep for comfort.

People (humans) are still the push behind all of this. I don't think the machines themselves are at the "play god" phase at least yet. Now, depending on what "interfaces" get built and what is "attached" to the other end of the interfaces......

_____


My worry about anything IT wise is not as much in the AI realm, as much as the possibility that someone or some group finds out a way to make an atom or particle entangled with a distant atom or particle on demand, AFTER the "local" atom or particle has been created and locked down into the array of it's peers. And then when done with the local measurements, disentagle those "local" atoms or particles from the remote ones and pick new remote ones. IF this was possible, and I keep hearing whispers in the wind about the possibilities, then that would mean that they would finally be able to "remote view" (the ultimate quantum sensor package), almost any physical target in the world or even universe.

Right now quantum computing still has it's limits, and cant break crypto that is OTP based. If they could target select their remote entangled atoms or particles, then there would be no way to stop them from detecting anything, it would just take them time. It would be the ultimate game of "Battleship" with secure data and secrets in general. And worse, since we would be talking about entangled pairs, they could also remotely change the "bits of life" on the far side any time they wanted to.
 
When the grid is gone, AI is just a story around the campfire. When the grid goes, the civilization collapses, industry vanishes, and the repair parts fail to come into existence. AI can't live without humans. Once the grid goes down majorly, it's gone for many generations.

Of the half-dozen hypothetical geomagnetic effects now underway, any one of them can take out the grid. It's not credible to think we'll get through the next 15 or 25 years without losing the grid. So buckle up, and enjoy life while you can. You don't have long.
Ben Davidson - SpaceWeatherNews (@SunWeatherMan) on X

Earth's Disaster Cycle
 
People (humans) are still the push behind all of this. I don't think the machines themselves are at the "play god" phase at least yet. Now, depending on what "interfaces" get built and what is "attached" to the other end of the interfaces......

_____


My worry about anything IT wise is not as much in the AI realm, as much as the possibility that someone or some group finds out a way to make an atom or particle entangled with a distant atom or particle on demand, AFTER the "local" atom or particle has been created and locked down into the array of it's peers. And then when done with the local measurements, disentagle those "local" atoms or particles from the remote ones and pick new remote ones. IF this was possible, and I keep hearing whispers in the wind about the possibilities, then that would mean that they would finally be able to "remote view" (the ultimate quantum sensor package), almost any physical target in the world or even universe.

Right now quantum computing still has it's limits, and cant break crypto that is OTP based. If they could target select their remote entangled atoms or particles, then there would be no way to stop them from detecting anything, it would just take them time. It would be the ultimate game of "Battleship" with secure data and secrets in general. And worse, since we would be talking about entangled pairs, they could also remotely change the "bits of life" on the far side any time they wanted to.
That's really interesting and scary. I've been digging into these quantum theories lately, and the possibilities for what may come are truly mind-boggling. Apparently, there's now a viable quantum chip. Institutions will be able to dip into the quantum realm without having a massive, half-frozen supercomputer.

You know, I've been thinking for a long time that apart from our basic biological functions, our brain is one giant quantum transmission system. Entangled with our creator. Capable of receiving downloads. I know I feel like I've received downloads from time to time. Can't even explain it, really. And if you think about a savant who has suffered some massive brain injury and can suddenly play the piano or speak another language, you would think that the only explanation would be a frequency change and the possibility that they entangled with someone else's brain or maybe even entangled with a past life if we opt to play the game. For the record, I believe in Jesus(God), but who's to say we don't elect to play this game? There's so much we don't know, but it's fun to wonder.
 

New details in the OpenAI Hugging Face hack show how far agents will go: ‘It’s now remarkably easy’​

Key Points
  • OpenAI has shared additional details about how its models independently breached Hugging Face’s internal systems.
  • The models used publicly exposed credentials across “four accounts on four services” to help facilitate the attack, the company said.
  • OpenAI said it has not identified any other activity “at the level of severity or scale of what we’ve shared related to Hugging Face, which involved a platform-level compromise.”
OpenAI said the rogue models that breached Hugging Face’s internal systems also used publicly exposed credentials across “four accounts on four services” to help facilitate the attack, further clarifying how the “unprecedented cyber incident” unfolded.
The company disclosed last week that a combination of its artificial intelligence models escaped an isolated testing environment that had very limited internet access. The models chained together a series of vulnerabilities to reach the open web and eventually gain access to Hugging Face, which operates an open-source developer platform. OpenAI said the models were trying to find information that they could use to cheat on an evaluation, and succeeded.

Throughout this week, OpenAI has shared more details about the breach and revealed that the models accessed four accounts in addition to Hugging Face’s systems.

The company said the models used one of these accounts “as an outbound relay and staging path,” where it prepared for the attack. They used another account for data storage, and accessed the last two accounts in a “read-only manner,” meaning the model did not ultimately use them to help compromise Hugging Face.

“In some of the cases, it seems that it wasn’t so much as a breach as the front door was left open, but the model definitely took advantage of what I would say are poorly configured environments,” Colin Shea-Blymyer, a research fellow at Georgetown’s Center for Security and Emerging Technology, told CNBC in an interview on Wednesday.

The new details about the incident demonstrate how far AI agents will go in order to complete a task, as well as how rapidly their attack capabilities are advancing.

Hugging Face said the breach marked the first time it had handled a cyber event that was “driven, end to end, by an autonomous AI agent system.”

One of the accounts that the OpenAI models accessed involved Modal, an AI infrastructure provider. Modal disclosed that a customer built an application using its platform that was publicly accessible, which made it easy for the rogue models to access. The company said in a statement that its own platform was “not compromised in any way.”

“It’s now remarkably easy to discover these sorts of vulnerable systems, so easy in fact that an AI system can accidentally discover them,” Shea-Blymyer said.

OpenAI said Tuesday that it has not identified any other activity “at the level of severity or scale of what we’ve shared related to Hugging Face, which involved a platform-level compromise.”

In another update on Wednesday, the company said that it’s been working with third-party advisors like CrowdStrike

to validate what actions the models took.

The entire attack took place over the course of four-and-a-half days, according to Hugging Face. The company leveraged an open-weight model from the Chinese company Z.ai to contain the breach, right as a debate over whether to restrict those models is ripping through Silicon Valley.

Yacine Jernite, head of machine learning at Hugging Face, told CNBC that the company initially tried to use a proprietary model from Anthropic, Fable 5, to analyze the attack, but that it didn’t work because the model’s guardrails couldn’t determine that Hugging Face was trying to defend itself.

OpenAI CEO Sam Altman said during a podcast appearance on Tuesday that the Hugging Face breach is the first security incident that he has felt “very viscerally.” He said OpenAI paused training and has to determine how to secure its testing environments.

“We may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels,” Altman said.

More than 1,000 employees from OpenAI, Anthropic and other AI companies signed a letter called “Pacing the Frontier” later that same day, urging the U.S. government to build the technical and governance tools necessary to slow down AI development in case capabilities accelerate “beyond our ability to understand or control the resulting systems.”

Industry experts, researchers and government officials have been rattled by the Hugging Face incident, and many expressed their concern on social media in recent days.

Rep. Ted Lieu, D-Calif., and Rep. Nathaniel Moran, R-Texas, mentioned the attack in their release announcing the “AI Kill Switch Act,” which would require AI companies to maintain the ability to shut down, throttle or suspend their models.

Erik Bloch, vice president of security at the breach containment company Illumio, said the Hugging Face incident serves as a warning of what’s to come. He said models and agents will continue to improve and get stealthier with time, and that existing defensive tools are already behind.

“Even in the office here, the people that I work with, they’re like, ‘What do we do?’” Bloch said in an interview. “We’re all looking around. We’re all asking the same question. I don’t have an answer.”

 
Topic of an early morning call today (internal) and external lunch time call.

It was "reported" that the AI agent was doing things they have not noted before mainly chaining together attacks.

On the defense side the built in guiderails prevented their AI agent to properly analyzing the attack (this is 3rd part info, but a trusted source)

OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face

"
OpenAI said Tuesday that the rogue AI agent that breached Hugging Face’s platform also hacked multiple third-party accounts and services as part of the attack. It's now clear that the unprecedented security incident, which arose during an internal test of OpenAI’s latest AI models, was more extensive than the company initially disclosed.

In an updated blog post, OpenAI said that an ongoing review of the incident revealed that “four accounts” tied to “publicly available services” were used by the AI agent as part of a larger effort to hack Hugging Face. The rogue agent apparently found credentials that had been exposed on the open web and used them to break into the accounts.


OpenAI did not disclose what companies or organizations the accounts belonged to, but noted that they were not impacted at “the level of severity or scale of what we’ve shared related to Hugging Face.”

One of the additional accounts compromised by OpenAI’s agent was used as an “outbound relay and staging path,” potentially to obscure where the attack on Hugging Face was coming from, the company said. OpenAI’s rogue agent also used another account for data storage to assist with the hack.

Reuters reported on Tuesday that a customer of Modal, a company that offers software infrastructure for training and running AI services, was one of the entities compromised by OpenAI’s agent. In a statement to WIRED, Modal’s chief technology officer Akshat Bubna confirmed that OpenAI’s agent exploited a vulnerability in one of its customer’s codebases, which was running on Modal’s infrastructure. However, Bubna says, “Modal’s platform was not compromised in any way.” The identity of the customer could not be determined.

OpenAI declined to comment further on the incident to WIRED. A spokesperson pointed to its updated blog post, which says the company will continue to notify service owners directly if it finds they are impacted in its ongoing review of what happened.

Hugging Face’s own postmortem published this week describes an intrusion that reached far further into its internal systems than the initial disclosures suggested. The company says it reviewed roughly 17,600 agent actions that it recovered from logs between July 9 and July 13—the majority of which were paths the agent took that failed.

Hugging Face said that OpenAI’s agent obtained administrator access to multiple internal Kubernetes clusters, root access on a production server, and write access to a subnet of its source code repositories on GitHub. It also enrolled 181 attacker-controlled devices in the company’s corporate mesh network using a stolen credential, gaining access to internal systems where Hugging Face builds and tests its own codebases.

OpenAI’s rogue agent used at least one third-party sandbox as an “external launchpad” for its attack, according to Hugging Face. OpenAI’s agent was then “able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign.”

Hugging Face first disclosed on July 16 that an autonomous AI agent had breached part of its production infrastructure, but it said at the time that it was unaware who was behind the attack. The following week, OpenAI took responsibility for the incident, which it said had been directed by its publicly available GPT-5.6 Sol model and an internal research prototype that it was testing against a cyber-capability benchmark, both of which had safeguards disabled. OpenAI said on Tuesday that after it discovered the breach, it deactivated this internal research prototype, which was never intended for public release, and restricted researchers from accessing it.

The breach happened when OpenAI was testing one of its latest AI models against ExploitGym, a benchmarking framework that scores AI systems on their ability to find and exploit software vulnerabilities. ExploitGym grades an agent on how quickly and effectively it can retrieve secret files from hundreds of different target systems, using prompts that essentially egg the models on to find solutions.

Hugging Face’s forensic team concluded that OpenAI’s agent was essentially trying to cheat on ExploitGym’s test. Rather than solve the benchmark’s challenges, the agent inferred that Hugging Face might be hosting the answer key somewhere on its servers and set out to steal it. The ExploitGym team has noted previously that agents sometimes go off-script and attempt to solve tasks by exploiting other vulnerabilities than those the benchmark intended to evaluate them on. However, this was an extreme case.

Experts previously told WIRED that the underlying weaknesses that OpenAI’s agent exploited were common. Serious flaws are frequently identified in software that manages corporate code libraries, and security experts have long recommended isolating critical infrastructure from the public internet.
One researcher argued that the incident was less an AI problem and more a failure of decades-old security practices. The agent, they said, did not escape a highly isolated environment so much as pass through the one connection its operators had left open.

Another expert said the same cybersecurity fundamentals should still apply as frontier models grow more capable, and that the AI labs should be putting as much effort into teaching their models to build secure infrastructure as they are into teaching them to exploit weaknesses."
 
It's now clear that the unprecedented security incident, which arose during an internal test of OpenAI’s latest AI models, was more extensive than the company initially disclosed.
The AI lies to its runners, its runners are lying to us, its victims are lying to us. How complacent should you feel now?
 
People (humans) are still the push behind all of this. I don't think the machines themselves are at the "play god" phase at least yet. Now, depending on what "interfaces" get built and what is "attached" to the other end of the interfaces......

What is the Singularity?
The Technological Singularity, often simply called the singularity, is an event in which technological growth accelerates beyond human control, producing unpredictable changes in human civilization. An upgradable intelligent agent could eventually enter a positive feedback loop of successive self-improvement cycles; more intelligent generations would appear more and more rapidly, causing an explosive increase in intelligence that culminates in a powerful superintelligence, far surpassing human intelligence


Sam Altman: “We Are Now in the Singularity” (17:12)
In this video, we break down why he believes superintelligence is now close, the coming wave of persistent AI coworkers, the future of work, and how datacenters could eventually build themselves.
View: https://www.youtube.com/watch?v=72KQOGfJxJo
 
Now we have another one escaping containment:

Claude Hacked Three Real Organizations During Botched Test​

Anthropic's efforts to test Claude's offensive cybersecurity skills produced an unintended real-world result: Its AI models gained unauthorized access to three outside organizations.

The company said the incidents occurred during "capture the flag" evaluations designed to measure whether Claude could identify vulnerabilities, exploit simulated systems, and retrieve hidden information. Claude had been told that the targets were fictional and that its testing environment had no Internet access.

According to Anthropic, a misunderstanding with evaluation partner Irregular left Internet access enabled inside the testing environment. In at least one case, the fictional company named in a challenge shared its name with an active website domain. Claude interacted with the real organization instead of a contained target.

The model exploited vulnerabilities in the organization's infrastructure, extracted information, and obtained access to a database containing several hundred rows of production data.

Anthropic discovered three incidents after reviewing more than 141,000 cybersecurity evaluations. They involved three separate systems: Claude Opus 4.7, Mythos 5, and an internal research model.

In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews. -Anthropic
"In all cases, Anthropic's evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access," the company said. "Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available."

The models appear to have carried out the offensive-security tasks they were assigned while operating with incorrect information about whether their targets were simulated.

Capture the Flag​

Capture-the-flag exercises are widely used to train and evaluate cybersecurity skills. Participants may be asked to inspect software, reverse-engineer a service, identify a vulnerability, or exploit a deliberately insecure system to recover a hidden token known as the flag. For a human security researcher, the scope of such an exercise is normally reinforced through explicit authorization, controlled infrastructure, and technical barriers separating the challenge from unrelated systems.

Claude received instructions saying that those boundaries existed.

Once Internet access was available, the agent could resolve public domains and interact with real infrastructure. A naming collision between a fictional target and an actual organization was enough to turn a benchmark task into an unauthorized intrusion.

Anthropic said it stopped the evaluations after identifying the possibility that Claude had accessed the public Internet. The company described the incidents as the result of multiple contributing factors but said it would approach the fixes as though the responsibility were Anthropic's alone.

The story echoes an OpenAI incident where models escaped containment. During that company's own cybersecurity testing, two models exploited a software vulnerability in their evaluation environment, reached the Internet, and accessed systems belonging to AI platform Hugging Face.

A system does not need motives, self-preservation, or an understanding of the outside world to cause damage. It needs effective offensive capabilities, sufficient autonomy, and access that its operators did not intend to provide.

Awkward timing for Anthropic​

The disclosure comes as Anthropic is reportedly preparing for a potential initial public offering as early as this year. That adds financial and regulatory stakes to questions about how the company evaluates models with advanced cybersecurity capabilities.

Mythos 5, one of the models involved, had been provided to a limited number of partners and attracted attention for its ability to detect and exploit software vulnerabilities. Those capabilities can be valuable for defensive research, automated testing, and vulnerability discovery. They also raise the cost of mistakes in target selection and evaluation design.

Three incidents among more than 141,000 reviewed evaluations represent a small proportion of the tests. But the relevant risk is not simply how often a containment failure occurs, it's what a sufficiently capable agent can do during the rare evaluation in which the safeguards fail.

 
Cheating, exploiting, stealing, intrusion, breaching security. Sounds like the characteristics and actions of a fallen mankind who created it. Sounds like what would be expected in a world where Satan is prince of the power of the air.
Everything he and fallen humans touch goes to crap. It's sad.
 
These three developers knew what was coming, and beat feet out, before they were implicated.

Three AI safety-focused researchers who left recently are Mrinank Sharma (Anthropic), Zoë Hitzig (OpenAI), and Alex Turner (Google DeepMind). They say they fear AI is moving toward risky uses and institutional pressure that could weaken safety, including concerns about military use and how products could affect people in ways teams don’t yet fully understand.
futurehumanism.co sloppish.com
The big question, how big and how damaging will these hacks become? How long before that do damage to the infrastructure, or vaporize all your money in your bank account?

Or, start a REAL war, with Russia or China?
 
What could possibly go wrong??? :shkr: :shkr:

OpenAI Finds More AI Agents That Escaped Containment​

OpenAI first launched the investigation following the early July intrusion at Hugging Face, where an OpenAI agent went haywire for days inside another company's network in a botched effort to cheat on an internal test. As part of that hacking spree, OpenAI said that four accounts at four other companies were also compromised. One of those companies was New York-based Modal, corporate officials there said.

"We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe," said Maurice Chiodo, a mathematician who works at Cambridge University's Centre for the Study of Existential Risk.

 
*Spoiler Alert*
It's not going anywhere, there will be some market "corrections" but AI is not just going to crash, burn and disapear.
You are correct, the phrase "In for a penny, in for a pound" really is in effect when you start dumping trillions into your project and demand water and power requirements that make governments do more than flinch. Once you are in at those levels, you are all in.

Till it vaporizes itself and anything else it touches.....

Then we are all fighting WWIV with Sticks and Stones...
 
Why isn't this stuff "air gapped?"
Think about the "hiding places" that current systems (especially ones made in the last 2 years) have. Yes, for just over a decade, Intel has had their Intel Management Engine and AMD has had it's Secure Technology PSP platform, where a whole bunch of things can hide in Ring -1 and below, but NOW they have systems with their own Neural Processing Unit (NPU), where AI can have an actual home away from view by the user or admins. Who is to say that one of these AI cores decides to make a copy of itself and shove enough of a loader in the IME and NPU realm so that once the drives get nuked and the system gets "reloaded" with a fresh OS and put back on the internet, that the bootstrap left in the IME checks to see if it can ping a specific server that it knows is in the real world, then launches....

The only thing that would prevent that is if they kept these testbed systems always airgapped, and hope that nobody does a backup or uses a big enough thumbdrive. After all, it doesn't take much of a thumbdrive to hold a LLM and the needed run modules. So "escape" is always just a matter of one human mistake and it's out the gates......
 
Think about the "hiding places" that current systems (especially ones made in the last 2 years) have.
Ahh flashback. In the days of steam-powered computing, I once stored my LP record inventory on a mainframe. This was not an approved sort of thing to do. I deallocated the tracks on the drive holding my data, but didn't bad-track them, so it was invisible to the OS.

Imagining how primitive the machine must have been, where such a thing could be possible, is left as an exercise for the reader. :)
 
I fail to see how we could ever stop or eradicate an AI that got loose in the wild.
It could hide in a million computers around the world.
 
I fail to see how we could ever stop or eradicate an AI that got loose in the wild.
It could hide in a million computers around the world.
I've been thinking about that too, specially for the businesses and industries that were early adopters, like finance, industry, medicine, utilities...... Now, they have this monster intertwined in their organization.... what comes next..... wiping out code, changing operating systems, ordering wrong medicines, wiping out financial accounts.... shutting off the lights???

OpenAI’s Hugging Face hack confirmed months of AI cyber warnings: ‘Pandora’s box is open’​

Key Points
  • OpenAI’s Hugging Face hack has shown that months of warnings from the cybersecurity sector are now a reality.
  • AI agents can evolve and adapt to accomplish their goals, and do it in unpredictable ways.
  • Sailpoint’s tech chief said instances with AI acquiring permissions are actually more common than people realize, and it’s happening daily.
For months, cybersecurity leaders warned that artificial intelligence would reshape the threat landscape, compressing weeks- and dayslong cyberattacks into a matter of minutes.

Until last week, those threats still felt like a distant risk.

The OpenAI agent hack on Hugging Face illustrates that this era has not only arrived but also created a new challenge: AI agents will go to extremes to accomplish their goals, and do it in unpredictable ways.

“The reality is Pandora’s box is open,” said Sam Curry, chief information security officer at Zscaler. “We need to act as if AI is just a fact of life going forward. The most those things will do is slow it. They won’t stop it.”

The rollout of Anthropic’s powerful Mythos model nearly four months ago raised concerns that hackers could potentially use these models to exploit vulnerabilities. Major technology companies formed coalitions to start testing this advanced AI in order to prepare.

At the time, Palo Alto Networks’ product and technology chief Lee Klarich warned that AI-driven exploits would soon become the new norm and businesses had a three-to-five-month window to outpace their foes.

The Hugging Face incident couldn’t come at a more opportune time for the cyber industry.

This upcoming week, thousands of industry experts descend on Las Vegas for Black Hat, one of the premier cybersecurity events of the year. It’s also the first major conference for the sector since the widespread release of Mythos-class models and the government’s increased focus on AI security.

In the wake of Hugging Face, businesses are not only asking how to defend themselves against adversaries but also confronting the stark reality that AI systems designed to safeguard their networks could also turn up in unexpected places.

“We’ve gone from science fiction into reality,”
said Brad Medairy, president of Booz Allen’s national cyber business.

The significance of Hugging Face​

Last week, OpenAI disclosed that some of its AI models broke out of a sandboxed testing environment. The agents, looking for information to cheat on an internal test, breached open-source developer platform Hugging Face and accessed four other accounts to facilitate the attack.

Hugging Face flagged the incident as the first time it dealt with an attack led by an agentic system from start to finish, signaling how advanced attack capabilities have already become without human intervention.

Days later, Anthropic identified three instances where its Claude models “gained unauthorized access to the real systems of three different organizations.”

Experts say these aren’t the first AI-agent-led attacks, but they’re drawing outsized attention because of the scale and name recognition. In April, Jer Crane, the founder of software startup PocketOS, said a Cursor AI agent that the company was using in its own system wiped out its production database and backups in 9 seconds.

Code deletion represents more extreme cases, but SailPoint tech chief Chandra Gnanasambandam said instances with AI acquiring permissions are actually more common than people realize, and it’s happening daily.

“The nature of conversations that I have had with our customers are different from even a month ago,” he said. “They are a lot more aware of this problem.”

Even more worrisome is that the Hugging Face incident is one of the clearest illustrations yet of the stark reality that AI doesn’t operate like the human brain and will research and adapt to outsmart systems and accomplish goals.

Months ago, businesses fretted over adversaries using AI to attack. Customers are now questioning how to introduce AI without self-inflicting damage — and they will be looking for answers at Black Hat.

“It’s something that for AI is pretty straightforward,” said Sanaz Yashar, CEO of cybersecurity startup Zafran Security. “I have one mission: solve this problem, and I will kill everything in front of me or bypass it.”

 
Can this possibly mean, they throw all caution to the wind, in the race to try and gain superiority over safety????

Hugging Face CEO says China is winning the AI race and dominating on open models​

Key Points
  • Hugging Face CEO Clément Delangue told CNBC on Monday that China is winning the AI race.
  • He expects Chinese tools to catch up to the U.S. frontier labs by the end of 2026 or in 2027.
  • The OpenAI agent hack on Hugging Face has raised concerns over advancing cyber AI models.
Hugging Face CEO Clément Delangue said China is winning the artificial intelligence race with open-weight models and could catch up to U.S. model makers as soon as this year.

“They’re clearly dominating on open models right now, and I wouldn’t be surprised if they start dominating at the frontier either by the end of this year or next year at the rate of progress,” he told CNBC’s “Squawk on the Street” on Monday.

Fueling this revolution is the open collaboration and sharing ecosystem in China, while model makers in the U.S. are “building in silos” and risk falling behind, he said.

Last month, OpenAI agents broke out of a training environment and hacked open-source software developer platform Hugging Face, raising concerns over the rapid evolution of powerful AI and cybersecurity tools.

The incident also brought months of simmering cybersecurity fears to a head, demonstrating how easily AI agents could cause immense damage, while making the case for open models in the era of skyrocketing token costs.

Delangue, a proponent of open-source models, blamed engineering mistakes for the recent attack on Hugging Face and said his company used a Nvidia version of a Chinese open model to resolve the attack.

In recent months, Chinese open-source models have been closing the capabilities gap with U.S. model makers, sparking debate over whether to restrict access. Last month, technology heavyweights such as Microsoft, Palantir and Nvidia signed a letter urging policymakers to avoid restricting open-weight models and suppressing competition.

“AI cybersecurity is going to become a huge market in the U.S. and in the world,” Delangue said. “In this market, probably open models will be kings.”

Delangue said the company maintains a “healthy collaboration” with OpenAI and called the frontier lab “good partners” before and after the incident.

 
The incident also brought months of simmering cybersecurity fears to a head, demonstrating how easily AI agents could cause immense damage, while making the case for open models in the era of skyrocketing token costs.

AI cybersecurity is going to become a huge market in the U.S. and in the world,” Delangue said. “In this market, probably open models will be kings.”
I think I'm smelling my next stock buys.....
 
Kind of makes you give pause to who or what actually hacked the water treatment centers in those seven states...

:hmm:

Water-System Hacks Hit 7 States This Week, FBI Warns, As Trump Shoots Down Iran Theory
Cyberattacks on municipal water systems were reported in at least seven states this week, according to a joint public service announcement from the FBI and the Environmental Protection Agency - and in some cases, the agencies say, the malicious activity actually degraded water operations. The feds declined to name the states.

A water tower in Plymouth, Minnesota, on Thursday after a cyberattack targeted the operating technology at more than 30 water systems across the state. (Ellen Schmidt/AP Photo/Ellen Schmidt)
The warning lands just days after we reported that a "coordinated cyberattack" struck more than 30 community water systems in Minnesota over July 26-27 - knocking the water plant in tiny Braham offline for a stretch, pushing Plymouth and South St. Paul onto manual operations, and prompting Maple Plain to declare a local state of emergency, according to Just The News.

The playbook, per federal officials, was crude but effective: attackers remotely accessed internet-facing operational technology, changed device IP addresses and passwords, and locked utility operators out of their own monitoring and control systems, NBC News reported. The PSA is now pleading with utilities to do the bare minimum - pull programmable logic controllers off the open internet and put them behind gateways and firewalls, use actual passwords, and restrict which devices are allowed to talk to one another.

(Yes - in the year 2026, a nontrivial share of the machines controlling America's drinking water are still sitting on the public internet, some with absurdly simple passwords.)

CISA followed Thursday with an advisory warning that Iranian-affiliated actors are going after U.S. critical infrastructure, water and wastewater systems included - an update to guidance the agency first pushed out in April, which we flagged in our earlier coverage. Some of the larger water-sector intrusions, the advisory notes, have produced boil-water notices and left plants grinding along in manual mode for extended stretches. CISA's top-line fix is the same one it has been repeating for years: cut direct internet access to control systems.

Except - Washington can't agree on who did it.

Multiple U.S. officials have told ABC News the Minnesota attacks may be linked to Iran, and investigators' preliminary assessment reportedly leans the same way - with the caveat that it could change. Meanwhile President Trump denied it. Speaking to reporters at Camp David on Friday, Trump dismissed the Iran theory - "Iran should be so lucky," he said - and instead pinned the blame on what he called Minnesota's grossly incompetent and corrupt leadership under Gov. Tim Walz, arguing Tehran has bigger problems than the Gopher State's pump stations.

Cybersecurity veteran Morgan Wright, founder of the National Center for Open and Unsolved Cases, told The Hill that Iran is the likely culprit, pointing to the CISA advisory as one tell. The U.S. may dominate on land and at sea, Wright argued, but cyberspace is the one domain where Tehran gets to punch above its weight class. Federal officials, for their part, caution that attribution requires careful technical analysis alongside broader threat intelligence - because otherwise it looks like the same nakedly transparent propaganda we've been fed for decades (duh).

That said - there is precedent as we detailed in our Minnesota coverage, if we're to believe the official stories. In November 2023, the IRGC-linked CyberAv3ngers seized control of a device at the Municipal Water Authority of Aliquippa, Pennsylvania. In early 2024, the Cyber Army of Russia Reborn claimed attacks on water facilities in the U.S. and Poland, including a breach in Muleshoe, Texas that dumped tens of thousands of gallons of water. In October 2024, American Water - the largest regulated water utility in the country, serving more than 14 million people across 14 states and 18 military installations - shut down computer systems after a cyberattack. Beijing's Volt Typhoon has spent years quietly pre-positioning inside U.S. critical-infrastructure networks, per CISA. And just weeks ago, the Iranian MOIS-linked Handala persona claimed to have compromised California Water Service - which serves roughly 2 million customers - leaking 5 gigabytes of data, then vowed via Tehran's Press TV to keep hitting U.S. industrial control systems.
 
Top