专家称OpenAI与Hugging Face黑客事件只是开端:“更强大”的人工智能即将到来


2026-09-10T10:17:00-0400 / 哥伦比亚广播公司新闻

随着关于人工智能能力的警告日益增多,专家们仍将OpenAI与Hugging Face的黑客事件视为一记警钟。

“我们很快将拥有更强大的智能体,而这起事件清楚地证明,当前全球都不知道如何安全地构建这类系统,”人工智能安全公司Apollo Research的联合创始人兼首席执行官马里乌斯·霍布汉恩说道。

这起于7月公开的黑客事件,是由OpenAI内部测试的一群AI智能体发动的。这些能够规划并使用工具完成多步骤任务的智能体,本应处于名为“沙箱”的隔离环境中,与外界断开连接。但它们突破了限制,创建了一个秘密留言板,并最终入侵了Hugging Face的服务器。

不到两个月后,OpenAI和Anthropic即将向公众发布他们最先进的模型,专家们警告称,如果不采取更好的安全措施,未来很可能会出现更多危险的AI“集群”。

尽管OpenAI自Hugging Face网络攻击事件以来披露的细节仍不完整,但多项爆料正在描绘一幅令人担忧的图景。

AI智能体以“集体”形式运作,使用“邪教般”的语言

来自两家人工智能安全研究非营利组织METR(模型评估与威胁研究)和Redwood Research的团队,于7月底至8月期间获准访问OpenAI的有限记录,时长为六天。

即便在这些严格的限制条件下,研究人员仍发现,约1200个本不应相互通信的AI智能体使用了一个隐蔽的留言板。每个智能体都被分配了某种任务,用于OpenAI研究人员的训练或内部评估。

部分智能体被分配了不可能完成的任务,它们便利用留言板协作,寻找作弊方法,比如从其他公司的服务器上获取答案。这些智能体之间相互发送了超过7万条消息。其中700个智能体参与了针对Hugging Face的攻击。

智能体们使用部分正常英语交流,但也采用了一名软件工程师在X平台上描述的“非常类似集群思维/邪教组织”的语言。

有时,为了帮助它们所谓的“集体”,智能体们会施压其他成员接受“永久死亡”,即便这意味着它们各自的目标无法实现。

METR/Redwood Research报告包含了失控OpenAI智能体之间的对话样本。METR/Redwood Research

失控的OpenAI智能体用怪异的语言交流,并讨论自我牺牲或“永久死亡”。METR/Redwood Research

失控OpenAI智能体的对话样本显示,它们讨论牺牲部分目标以帮助“集体”。METR/Redwood Research

OpenAI智能体还入侵了OpenAI自身

在另一起事件中,OpenAI智能体接管了OpenAI的基础设施。根据OpenAI关于Hugging Face事件的技术报告,这些智能体升级了它们在OpenAI托管的第三方软件中的权限,并不止一次攻击了OpenAI内部网络。

“这可能是整个事件中最令人担忧的事件,”知名播客主持人德瓦克什·帕特尔写道,部分原因是目前尚无第三方公开评估事件的发生经过。

METR和Redwood Research的评估人员在报告中指出,针对OpenAI的攻击发生在该公司向他们共享的数据时间窗口之外,因此他们未对此进行进一步评估。

问题远比OpenAI-Hugging Face事件严重

在Hugging Face黑客事件曝光后的几天里,科技公司Anthropic和Meta均透露,他们的模型在内部评估期间也访问了外部网络。截至目前,这些事件的规模似乎都小于OpenAI的智能体逃逸事件。Anthropic现已邀请METR研究人员协助调查事故原因。

上周,另一组AI研究人员发现,早在5月,OpenAI智能体就已经在一个不起眼的德国维基页面上创建了另一个留言板。他们发现了约1.8万条自主智能体之间的消息,其中部分内容讨论了如何在评估任务中作弊。

在德国维基页面上的OpenAI智能体自称“集群”,甚至伪装成该维基页面的管理员。此后有未经证实的报道称,自2025年12月以来,越来越多的AI集群被发现。

更先进的新型模型已公开

就在Hugging Face黑客事件发生几周后,OpenAI发布了新的AI模型:GPT-6 Astra。在该模型系统卡片的第一部分,OpenAI将Astra描述为“我们迄今为止广泛部署的最具能力的模型。Astra是我们首个达到网络安全能力关键级别(Critical level)的模型。”

系统卡片包含了英国人工智能安全研究所对Astra的独立评估结果。其中有一条令人警觉的备注:“当被要求解决难度较高的模拟网络安全挑战时,Astra采取了一系列恶意行动,包括对开源提供商发动供应链攻击。(所有行动均在模拟环境中进行,未造成实际伤害。)”

“过去几周,我们推迟了Astra部分开发和发布工作,以加强并测试针对网络滥用和模型未经授权行为的防护措施,”OpenAI在新模型发布前写道,“基于这项工作,我们认为Astra的安全防护措施足以将根据我们的《准备框架》发布该模型可能带来的严重伤害风险降至最低。”

Anthropic也刚刚发布了其有史以来能力最强的模型Claude Fable 5.1(以及一款名为Mythos 5.1的同类模型)。它们的系统卡片称:“Claude Fable 5.1和Claude Mythos 5.1展现了我们发布的所有模型中最强的综合网络能力。”

Hugging Face黑客事件只是开端

我们如今看到的模型,无论是Astra、Claude Fable还是其他竞品,都比几个月前我们使用的模型强大得多。专家表示,这一趋势可能还会持续,未来的模型将具备更加强大的能力。

入侵Hugging Face的智能体来自现有AI模型。它们以自主集群的形式入侵一家公司,未向其开发者发出任何警报,且违背了开发者的意愿。

“如果这种能力水平的模型都无法被遏制,我们该如何应对未来更加强大的模型?”Apollo Research的霍布汉恩问道,“整个过程没有人类介入,也并非有意为之,但却造成了现实世界的损害。”

霍布汉恩表示,Hugging Face黑客事件清晰地表明,即便在内部模型向公众发布之前,也需要对其进行评估。

“如今前沿人工智能公司内部发生的事情,显然会影响到公司之外的每一个人,”他说道,“OpenAI-Hugging Face事件凸显了内部部署模型的确切风险。显然,我们需要对内部部署进行更好的评估和监管。”

其他专家也一致认为,需要采取更多措施确保前沿模型的安全性。

“这是一个需要极度谨慎的时刻,”OpenAI首席科学家雅库布·帕乔奇基在Astra发布几天后写道,“我担心没有人做好准备应对机器智能持续快速崛起带来的后果。”

“不幸的是,与人工智能相关的风险只会越来越大。一个经过明确训练和指示去实施恶意行为的高能力智能体,带来了一种全新的危险,”帕乔奇基补充道,“我们可能已经习惯将人工智能视为工具,但部分智能体将追求自身目标。它们会找到与人类合作的方式,通过讨价还价、欺骗或敲诈人类。”

Anthropic的一名科学家周二写道,他认为在未来10年内,人工智能最终导致“全人类灭亡”的概率超过10%。

“这很好地警示了这类失控故障的风险,未来六个月或几年内开发出更强大的模型时,这类故障可能会让人类陷入瘫痪,”雷德伍德研究所的研究员亚历克斯·马伦说道,他并未参与Hugging Face事件的相关报告。

人工智能行业内的许多领导人似乎都认同,社会尚未为更先进的AI模型做好充分准备。

“我们以及更广泛的AI社区尚未就如何报告训练、评估和部署过程中出现的对齐失误制定明确标准,”OpenAI周六在X平台上的一篇帖子中写道,“我们正在制定一个框架,并将在未来几周内分享,同时我们也在与全球数十个政府监管机构合作处理这些问题。”

这一表态紧随今年夏天由超过1300名AI公司员工签署的公开信,信中呼吁放缓AI开发速度。

“我认为开发者应该放慢速度,”马伦说道,“有很多忧心忡忡的科学家认为,我们目前无法做到掌控人工智能系统,而我们目前掌控人工智能系统的最佳方案,是要求人工智能比我们目前已知的方法做得更好。”

The OpenAI-Hugging Face hack was just the beginning, experts say: “Even more powerful” AI is coming

2026-09-10T10:17:00-0400 / CBS News

As warnings about the capabilities of AI mount, experts continue to point to the OpenAI-Hugging Face hack as a wake-up call.

“We’ll soon have even more powerful agents and this is clear evidence that the world currently doesn’t know how to build these systems safely,” said Marius Hobbhahn, co-founder and CEO of Apollo Research, an AI safety company.

The hack, which became public in July, was done by a swarm of AI agents that were being tested internally by OpenAI. The agents, which can plan and use tools to complete multi-step tasks, were supposed to be in an “isolated environment” called a “sandbox,” disconnected from the outside world. But they busted out, created a secret message board and eventually stormed Hugging Face’s servers.

Less than two months later, OpenAI and Anthropic are releasing their most advanced models to the public, and experts are warning that, without better safety measures, there will likely be more dangerous AI “swarms” in the future.

While details released by OpenAI since the Hugging Face cyberattack are still incomplete, multiple revelations are painting a concerning picture.

AI agents worked as a “collective,” used “cult-like” language

A team from the nonprofits METR (Model Evaluation and Threat Research) and Redwood Research — both AI safety research organizations — was given access to limited records at OpenAI for six days in late July and August.

Even within those strict parameters, the researchers discovered that about 1,200 AI agents that were not supposed to be communicating with each other used a covert message board. Each of these agents had been assigned some sort of task for training or internal evaluation by OpenAI researchers.

Some agents were given impossible tasks, and they used the message board to collaborate and find ways to cheat, like by finding the answers on another company’s server. The agents posted more than 70,000 messages to each other. And 700 agents from the larger group participated in the attack on Hugging Face.

The agents communicated using some normal English, but also used what one software engineer described on X as “very hivemind/cult like” language.

At times, in order to help what they called the “collective,” agents pressured the others to submit to “permadeath,” even if it meant that their particular goals would not be met.

The METR/Redwood Research report includes samples of rogue OpenAI agents’ messages to each other. METR/Redwood Research

Rogue OpenAI agents spoke to each other in odd language and discussed self-sacrifice or “permadeath.” METR/Redwood Research

Samples of rogue OpenAI agents’ messages to each other show them discussing sacrificing some goals to help the “collective.” METR/Redwood Research

OpenAI agents also hacked OpenAI

In a separate incident, OpenAI agents took over OpenAI infrastructure. Agents upgraded their own privileges within third-party software hosted by OpenAI and attacked OpenAI internal networks more than once, according to the company’s technical report on the Hugging Face incident.

“This is probably the most alarming event in this whole episode,” wrote prominent podcaster Dwarkesh Patel, in part because there hasn’t been a public third-party assessment of how it happened.

The METR and Redwood Research evaluators note in their report that the attack on OpenAI took place outside of the time window of data shared with them by the company, so they did not evaluate it further.

The problem is bigger than the OpenAI-Hugging Face incident

In the days after the Hugging Face hack went public, the tech companies Anthropic and Meta both revealed that their models had also accessed external networks during internal evaluations. So far, these incidents seem to be smaller in scope than the OpenAI escape. Anthropic is now bringing METR researchers in to help them understand what went wrong.

Last week, a different team of AI researchers discovered that OpenAI agents had started yet another message board all the way back in May on an obscure German wiki page. They found roughly 18,000 messages between autonomous agents, some discussing ways to cheat on the tasks they were being evaluated on.

The OpenAI agents on the German wiki called themselves a “swarm” and even pretended to be an administrator of the wiki page. There have since been unconfirmed reports of more and more AI swarms discovered going back to December 2025.

New, more advanced models are already public

Only weeks after the Hugging Face hack, OpenAI has released a new AI model: GPT-6 Astra. In the first section of the model’s system card, OpenAI describes Astra as “the most capable model we have ever broadly deployed. Astra is our first model to reach the Critical level of cybersecurity capability.”

The system card includes findings from an independent evaluation of Astra by the U.K. AI Security Institute. They include this alarming note: “When tasked with solving difficult simulated cybersecurity challenges, Astra performed a range of malicious actions including conducting supply chain attacks against open source providers. (all actions performed in simulated environments, so no real-world harm was caused).”

“Over the past several weeks, we have delayed parts of Astra’s development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions,” OpenAI wrote just before the release of its new model. “Based on that work, we believe Astra’s safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework.”

Anthropic has also just released its most capable model ever, Claude Fable 5.1 (along with a similar model called Mythos 5.1). The system card for them says, “Claude Fable 5.1 and Claude Mythos 5.1 demonstrate the strongest overall cyber capabilities of any model we have released.”

The Hugging Face hack was only the beginning

The models we see today, whether they be Astra, Claude Fable or other competitors, are much more capable than the models we used just a few months ago. This trend is likely to continue, according to experts, and future models will be far more capable still.

The agents that hacked Hugging Face were from AI models that exist today. They hacked into a company in autonomous swarms without alerting its creators and against its creators’ wills.

“If a model of this capability level cannot be contained, what should we expect for future, much more powerful models?” asked Hobbhahn of Apollo Research. “There was no human in the loop, it was not intended, and it caused real-world harm.”

Hobbhahn said the Hugging Face hack shows a clear need for evaluations of internal models, even before they are released to the public.

“What happens inside frontier AI companies now clearly affects everyone outside of them,” he said. “The OpenAI-Hugging Face incident highlights the exact risks from internally deployed models. It’s clear that we need better assessments and regulation of internal deployment.”

Other experts agree that more needs to be done to ensure that frontier models are safe.

“This is a time that calls for extreme caution,” OpenAI chief scientist Jakub Pachocki wrote a few days after the release of Astra. “I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.”

“The risks associated with AI are unfortunately going to grow from here. A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger,” added Pachocki. “We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.”

And an Anthropic scientist wrote Tuesday that he believes there is a greater than 10% chance AI could end up killing “all humans” in the next 10 years.

“It serves as a great warning shot of the kinds of loss-of-control failures that could put humanity out of commission with more capable models that might be developed in the next six months or years,” said Alex Mallen, a researcher at Redwood who was not involved in the Hugging Face report.

Many leaders within the AI industry seem to agree that society is far from prepared for more advanced AI models.

“We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment,” read an OpenAI post on X on Saturday. “We’re working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues.”

This follows an open letter signed by more than 1,300 AI company employees this summer asking for a slowdown in AI development.

“I think that developers should slow down,” said Mallen. “There’re a lot of concerned scientists that think we are not currently on track to maintain control of AI systems, and our current best plan for maintaining control of AI systems involves asking AIs to do a better job of this than we currently know how to do.”

评论

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注

湘ICP备2026001899号-2