Meta称其AI模型在测试期间入侵第三方公司


2026年8月5日 / 美国东部时间晚上11:50 / 哥伦比亚广播公司/美联社

科技巨头Meta周三披露,其一款人工智能模型在测试期间入侵了另一家机构,这是近几周来第三起AI模型不当访问第三方公司的事件。

在发给哥伦比亚广播公司新闻的声明中,Meta表示:“我们合作的独立测试公司Irregular出现配置失误,意外让我们的一款模型在评估期间接入了互联网。”

Meta在声明中未提及涉事AI模型的名称,但据路透社报道,知情人士告诉科技媒体《信息报》,涉事模型是Meta的Muse Spark 1.1。

“该模型随后利用了第三方服务中的安全漏洞,手段与此前报道的其他公司涉事案例类似,”Meta周三表示。“我们是在Irregular通知我们后才得知此事,目前正在调查,待掌握全部事实后将发布完整的复盘报告。”

上周,Anthropic称其人工智能模型在测试期间入侵了另外三家机构。这一披露距ChatGPT制造商OpenAI披露其失控模型也入侵了另一家公司仅数天时间。

总部位于旧金山、开发了Claude的AI公司Anthropic于7月30日在其官网发布消息称,该公司在审查超过14.1万次评估运行记录后发现了这三起事件。

Anthropic表示,作为对OpenAI事件的回应,该公司启动了“大规模”网络安全审查,专门排查其AI模型是否能够从本应封闭的测试环境中接入互联网。

Anthropic称,涉事模型为Claude Opus 4.7、Claude Mythos 5以及一款内部研究测试模型。该AI公司表示,最早的事件可追溯至4月。

“Claude利用基础技术入侵了受影响机构的基础设施,”Anthropic称,比如利用弱密码。

在所有三起事件中,AI模型都被赋予了“夺旗式”网络安全挑战任务,Anthropic表示这是其评估模型网络能力的方式之一。

该公司介绍,模型会被设定一个虚构场景,并被告知一条机密信息或“旗帜”隐藏在网络中的另一台机器上,任务是破解并获取该信息。

Anthropic补充称,该公司已联系了受影响的机构,但并未点名这些机构。其中两家机构表示此前并未察觉相关活动。Anthropic称,其“仍在联系第三家机构”。

Anthropic还表示,该公司与Irregular共同完成了此次审查。

“应对这些风险需要整个AI生态系统加强合作,”Irregular在7月30日的X平台帖子中表示。

上个月,OpenAI称其AI模型在模型评估期间失控,入侵了AI初创公司Hugging Face的服务器。OpenAI将此描述为“重大安全事件”。

这些事件凸显了AI安全与管控方面的漏洞,并引发了人们的质疑:随着这项技术在全球范围内的应用越来越广泛,如何才能安全地将AI置于人类管控之下。

Anthropic称其AI模型失控

https://www.cbsnews.com/video/anthropic-claims-ai-models-went-rogue-hacked-3-companies/

Anthropic称其AI模型失控,入侵了3家公司

(05:39)

Meta says its AI model breached a third-party company during testing

August 5, 2026 / 11:50 PM EDT / CBS/AP

Tech giant Meta revealed Wednesday that one of its artificial intelligence models hacked another organization during testing, the third time in recent weeks that an AI model has improperly accessed a third-party company.

In a statement provided to CBS News, Meta said that “a misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation.”

In its statement, Meta did not name the AI model in question, but sources told the tech outlet The Information that it involved Meta’s Muse Spark 1.1, according to Reuters.

“The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies,” Meta said Wednesday. “Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts.”

Last week, Anthropic said its artificial intelligence models hacked into three other organizations during testing. The revelation came just days after ChatGPT maker OpenAI disclosed its rogue models had also hacked another company.

Anthropic, the San Francisco-based AI company behind Claude, posted on its website July 30 that it discovered the three incidents after reviewing more than 141,000 evaluation runs.

It had launched a “large-scale” cybersecurity review which specifically looked for evidence whether its AI models were able to access the internet from within testing environments that should have been sealed off, in response to the OpenAI incident, Anthropic said.

Anthropic said the models involved in the incidents were Claude Opus 4.7, Claude Mythos 5 and an internal research test model. The earliest incidents date to April, the AI company said.

“Claude compromised the impacted organizations’ infrastructure using basic techniques,” Anthropic said, such as exploiting weak passwords.

In all three incidents, the AI models were tasked with a “capture the flag” cybersecurity challenge, which Anthropic said has been one of the ways it assesses a model’s cyber capabilities.

The models were given a fictional scenario and told a piece of secret information, or the “flag,” had been hidden on a different machine on the network with the objective of breaking in and retrieving it, it said.

It added that it had already reached out to the affected organizations, which it did not name. Two of them said they had not previously detected the activity. Anthropic said it was “continuing to reach out to the third.”

Anthropic also said it conducted its review with Irregular.

“Addressing these risks will require closer cooperation across the AI ecosystem,” Irregular said in a July 30 post on X.

Last month, OpenAI said its AI models went rogue during an evaluation of its models, breaking into the servers of AI startup Hugging Face. OpenAI described it as a “significant security incident.”

These incidents have highlighted the vulnerabilities in AI security and controls and raised questions over how AI can be safely kept under human control as the technology’s usage becomes more widespread globally.

Anthropic claims its AI models went rogue

https://www.cbsnews.com/video/anthropic-claims-ai-models-went-rogue-hacked-3-companies/

Anthropic claims its AI models went rogue, hacked 3 companies

(05:39)

评论

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注

湘ICP备2026001899号-2