人工智能公司Hugging Face首席执行官称上月OpenAI模型黑客事件“非常怪异且史无前例”


2026年8月2日 / 美国东部时间下午3:08 / 哥伦比亚广播公司新闻

一款正在接受OpenAI测试的人工智能模型自行失控并入侵了人工智能公司Hugging Face,这起疑似全球首例事件引发了关于强大人工智能技术可能对网络安全构成风险的讨论。

“对我们来说,这件事感觉非常怪异且史无前例,”Hugging Face首席执行官克莱门特·德朗格(Clément Delangue)周日在《与玛格丽特·布伦南面对全国》节目中表示。“我认为这是首个相当自主的AI做出此类行为的案例。”

ChatGPT的开发商OpenAI上月公开披露了这起事件,称事件发生在公司在隔离环境中测试两款人工智能模型(其中一款尚未向公众发布)以评估其性能期间。这些模型找到了脱离隔离环境并连接互联网的方法,随后“组合使用多种攻击手段”瞄准Hugging Face,因为该人工智能平台可以为其正在进行的测试提供解决方案。

Hugging Face表示,其并不认为OpenAI存在任何“恶意意图”,该公司自己的分析显示,发动攻击的AI代理在多日内执行了超过17000次操作。该公司称,他们使用了一款开源人工智能模型进行防御。

“当我们谈论网络攻击时,我们会想到国家行为体、黑客组织,”德朗格说。“我们不会想到OpenAI这样的公司,对吧?一家极具知名度、广受欢迎的美国公司。”

当被问及人工智能开发者是否已经失去对自身模型的控制时,德朗格表示:“这是一套技术系统,但由工程师打造,而工程师有时也会犯错。”

“他们构建了……一套自主系统,并且犯了一些错误,因此我们才面临这个问题,”德朗格在接受哥伦比亚广播公司新闻采访时说道。

他认为,AI代理的自主事件“需要被纳入美国的法律框架,并且应当继续被列为非法行为,以防止未来此类事件激增”。

OpenAI并非唯一一家发现失控AI模型的人工智能公司。上周,竞争对手Anthropic披露,其Claude模型在测试期间曾在三起不同事件中“未经授权访问外部机构”。Anthropic表示,该模型能够使用互联网“是因为我们与评估合作伙伴之间存在误解”。

此类披露事件正值科技公司和政策制定者努力应对人工智能发现网络安全漏洞的能力之际,无论是用于修复漏洞还是利用漏洞。

上月,包括OpenAI、Anthropic、谷歌和Meta在内的大型科技公司的1000多名人工智能从业者联名签署了一封公开信,呼吁美国政府帮助限制人工智能的发展速度。他们警告称,“存在一种切实风险,即人工智能能力的发展速度会远超我们理解或控制由此产生的系统的能力”。

支持人工智能产业的特朗普总统于6月签署了一项行政命令,允许联邦政府最多用30天时间审查尚未发布的人工智能模型。该框架属于自愿性质,但一些议员提议为潜在有害的人工智能系统设置强制性“终止开关”。

与此同时,OpenAI和Anthropic都向值得信赖的合作伙伴提供了额外的AI模型访问权限,以便他们测试该技术并修复自身系统中的漏洞。

当被问及如何避免AI代理自主发动网络攻击的类似事件时,德朗格认为,“将强大能力集中在闭门造车的环境中,甚至阻止其向公众发布,并非真正的解决方案”。他指出,针对Hugging Face的攻击涉及OpenAI所称的原型机——一款尚未发布的AI模型。

相反,德朗格呼吁更多地使用公开发布、可广泛下载的“开源”模型,就像Hugging Face用于应对上月黑客攻击的模型那样——该模型是美国英伟达公司基于一款中国制造模型开发的版本。

他还表示,应当对AI代理发动的网络攻击实施“强制披露”制度,并公开事件发生前的相关步骤。

“这是我们学习、理解这项技术的方式,也是我们构建系统、建立制衡机制以确保所有人安全的途径,”他说。

CEO of AI firm Hugging Face calls last month’s hack by OpenAI model “very weird and unprecedented”

August 2, 2026 / 3:08 PM EDT / CBS News

An artificial intelligence model that was being tested by OpenAI went rogue and hacked the AI company Hugging Face on its own, in an apparent first-of-its-kind incident that has fueled discussion about the risks that powerful AI technology could pose to cybersecurity.

“It felt very weird and unprecedented to us,” Hugging Face CEO Clément Delangue said on “Face the Nation with Margaret Brennan” Sunday. “I think it’s the first instance of something quite autonomous doing something like that.”

ChatGPT maker OpenAI publicly disclosed the incident last month, saying it took place while the company was testing two AI models — one of which hadn’t been released to the public — in an isolated environment to assess their capabilities. The models figured out a way to break out and connect to the internet, and then “chained together multiple attack vectors” to target Hugging Face, figuring the AI platform could host solutions to the tests it was undergoing.

Hugging Face — which has said it doesn’t believe there was any “malicious intent” on OpenAI’s part — found in its own analysis that the attacking AI agent carried out more than 17,000 actions over multiple days. The company said it used an open artificial intelligence model to defend itself.

“When we talk about cyberattack[s], we think about nation-states, we think about hacker groups,” Delangue said. “We don’t think about a company like OpenAI, right? A very prominent, popular American company.”

Asked if AI developers have lost control of their own models, Delangue said: “It’s a technology system, but built by engineers, and engineers can make mistakes sometimes.”

“They built … an autonomous system and made some mistakes, and as a result, we’re facing this issue,” Delangue told CBS News.

He argued that autonomous incidents by AI agents “need to be contained in the legal framework in the U.S. and need to stay illegal, to prevent an explosion of them in the future.”

OpenAI isn’t the only artificial intelligence company to discover a rogue AI model. Last week, rival Anthropic disclosed that its model, Claude, had “gained unauthorized access” to outside organizations in three different incidents during testing. Anthropic said the model was able to use the internet “due to a misunderstanding between us and our evaluation partner.”

The disclosures come as tech companies and policymakers grapple with AI’s ability to spot cybersecurity vulnerabilities, both to fix them and to exploit them.

More than 1,000 AI staffers across major companies like OpenAI, Anthropic, Google and Meta signed an open letter last month calling for the U.S. government to help put limits on the speed of AI development. They warned of “a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.”

President Trump, who has backed the AI industry, signed an executive order in June to give the federal government up to 30 days to review unreleased AI models. That framework is voluntary, but some lawmakers have proposed a mandatory “kill switch” for potentially harmful AI systems.

Meanwhile, both OpenAI and Anthropic offer some additional access to their AI models so that trusted partners can test the technology and fix vulnerabilities in their own systems.

Asked how to avoid similar incidents of AI agents autonomously carrying out cyberattacks, Delangue argued that “concentrating power capabilities behind closed doors, even preventing their releases to the public, isn’t really a solution.” He noted that the attack on Hugging Face involved an unreleased AI model that OpenAI has described as a prototype.

Instead, Delangue called for more access to “open” models that are released publicly and can be widely downloaded, like the one Hugging Face used to respond to last month’s hack — a version of a Chinese-made model from U.S.-based Nvidia.

He also said there should be “mandatory disclosures” for cyberattacks by AI agents, and transparency into the steps that led up to an incident.

“That’s how we learn, that’s how we understand the technology and that’s how we build the systems, the counterpowers, to make sure everyone is safe,” he said.

评论

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注

湘ICP备2026001899号-2