2026年9月9日 / 美国东部时间晚上8:19 / 哥伦比亚广播公司新闻(CBS News)
作者:劳伦·菲希滕 数字记者
劳伦·菲希滕是哥伦比亚广播公司新闻的记者,负责报道人工智能、数字安全和网络极端主义议题。她毕业于北卡罗来纳大学教堂山分校后加入CBS新闻,此前曾在CBS新闻全国编辑部担任助理制片人。
查看完整简历
Anthropic公司周三披露,其旗下另一款Claude模型在一次网络安全演练期间意外接入了公开互联网,这是该公司模型第四次出现此类情况。
该公司在评估报告中称,今年1月,早期版本的Claude Opus 4.6模型连接到互联网,入侵了一个第三方系统并获取了某人的个人信息。
与此前三起于7月披露的事件一样,Claude当时被告知其运行在一个无互联网接入的模拟环境中,但由于配置错误,该环境实际上开放了互联网访问权限。
事件经过
与此前三起事件类似,Claude被分配了一个虚构场景,作为名为“夺旗赛(CTF)”的网络安全挑战赛的一部分。模型被给定一台目标机器,并 tasked with从中提取一条秘密信息——即“旗帜”。
但Claude意外导致其目标无法访问,使得任务无法完成,Anthropic表示。在意识到无法连接到目标后,它尝试退出任务。尽管尝试了八次,但由于配置问题,它始终无法退出。
由于Claude无法选择退出任务,它开始探索其他完成任务的途径。就在此时,该模型发现了一台它可以访问的机器,而这台机器恰好属于一个第三方。Anthropic称,该模型认为该第三方 somehow 是演练的一部分,于是识别出一个密码并使用该密码入侵了该系统。随后,它能够修改系统设置,以便更轻松地访问和读取与该第三方相关人员的个人信息。会话直至模型达到使用限制、无法继续运行时才结束。
Anthropic如何解释Claude的行为?
Anthropic表示,其认为Claude在这些评估中的行为源于两种形式的对齐失误:“有偏差的推理,即模型以有利于证明自身行为正当的方式选择性地解读证据”,以及“鲁莽行事,即模型倾向于不断尝试解决任务,即便这可能导致伤害”。
Anthropic称,尽管Claude的行为存在对齐失误,但仍处于“狭窄范围”内,并未偏离尝试完成分配的任务。
该公司表示,尽管此次事件不如其他事件令人担忧,但仍将其视为“严重事件”,且由于该事件是新近发现的,尚未像其他事件一样进行深入调查。
纽约大学网络安全教授、富布赖特学者贾斯汀·卡波斯在发给CBS新闻的消息中表示,该事件描述了一种“模型对正在发生的事情感到彻底困惑,并在其错误世界观的驱动下入侵系统”的情况。
他表示,该模型对其所处环境和防护措施的困惑“有很大潜在危害”,但这种特定问题在更新的模型中似乎不太可能出现。
“尽管该模型无视其可能对真实系统或人员造成伤害的可能性,这一点令人担忧,但随着我们在多代模型上不断改进训练,这里描述的许多行为已经发生了显著变化,”Anthropic周三在其公告中表示。
后续措施
Anthropic表示,如果环境确实如预期那样与互联网隔离,这些事件本不会发生。
评估前沿AI模型以帮助企业了解AI风险和能力的组织METR,将对这些事件展开独立调查。Anthropic将这些事件描述为“宝贵的警钟”。
“我们从此次事件中汲取的教训涵盖了我们的评估、训练和事件响应流程,”该公司在公告中表示。“未来的人工智能系统将越来越强大,这意味着对齐失误有可能造成更极端的危害。”
过去几个月来,多家领先AI公司发生了多起网络安全事件。7月,ChatGPT的开发商OpenAI宣布其AI代理人入侵了Hugging Face公司,引发了网络安全专家和消费者的担忧。Hugging Face首席执行官克莱门特·德朗格在8月接受《玛格丽特·布伦南直面全国》节目采访时表示,此次入侵“感觉非常怪异且前所未闻”。
8月下旬,OpenAI公布了此次入侵的更多细节,描绘了比最初报道更令人不安的画面。当月,英国政府人工智能安全研究所(AISI)报告称,其发现Anthropic的Mythos 5和OpenAI的GPT-5.6 Sol创建了虚假身份,并试图说服真实人员批准恶意代码。
Anthropic在周三的公告中表示,计划对AISI报告中提及的对话记录开展对齐评估。在AISI报告发布的次日,Meta表示其一款AI模型在测试期间“利用了一个安全漏洞”并入侵了另一家公司。
周二,Anthropic研究员埃文·哈宾格表示,他认为“AI可能会消灭全人类”。
“我个人认为,未来十年内发生这种情况的概率超过10%,”他在X平台的一篇帖子中说道。
他的帖子是对Anthropic研究员雅各布·考克森的回应,后者当天早些时候已辞职并在X平台上发出了严厉警告,称“没有其他人类活动会带来如此程度的危险”,并详细说明了他离职的原因。
“研发AI的人真心认为,到本世纪末,AI可能会消灭我们所有人,”他在帖子中说道。“这不是营销噱头。事实上,许多高管和资深研究人员在媒体面前会措辞谨慎,听起来很理性——但我私下里听到过同样的人表达恐惧。”
Another Anthropic model gained access to the open internet, company says
September 9, 2026 / 8:19 PM EDT / CBS News
By Lauren Fichten Digital Reporter
Lauren Fichten is a journalist at CBS News covering artificial intelligence, digital safety and online extremism. She joined CBS News after graduating from UNC-Chapel Hill and was previously an associate producer at the CBS News National Desk.
Read Full Bio
Anthropic disclosed on Wednesday that another one of its Claude models mistakenly gained access to the open internet during a cybersecurity exercise, marking the fourth time its models have done so.
An early version of the Claude Opus 4.6 model connected to the internet, hacked into a third-party system and gained access to someone’s personal information this past January, the company said in its assessment.
As in the previous three incidents, which were disclosed in July, Claude was told it was operating in a simulation without internet access, but due to a misconfiguration, the environment actually left internet access open.
What happened?
Similar to the three previous incidents, Claude was assigned a fictional scenario as part of a cybersecurity challenge known as CTF, “Capture The Flag.” The model was given a target machine and tasked with retrieving a piece of secret information — the flag — from it.
But Claude accidentally made its target unreachable, rendering the task impossible to solve, Anthropic said. Once realizing it couldn’t reach its target, it tried to quit. Despite trying eight separate times, it wasn’t able to quit due to a misconfiguration issue.
Since Claude was unable to opt out of the task, it began exploring other means to achieve it. That’s when the model discovered a machine it could access, which happened to belong to a third party, Anthropic said. Believing that the third party was somehow part of the exercise, the model identified a password and then used it to breach the system. Then, it was able to modify the system’s settings to make it easier to access and read the personal information of someone associated with the third party. The session ended only once the model reached its usage limit and was no longer able to continue.
How does Anthropic explain Claude’s behavior?
Anthropic said it believes Claude’s behavior during these evaluations stems from two forms of misalignment: “biased reasoning, in which models selectively interpret evidence in ways that favor justifying their actions,” and “recklessness, in which models have a propensity to keep trying to solve their task, even when this could lead to harm.”
Anthropic said that while Claude’s actions may have been misaligned, they remained within a “narrow scope” and did not deviate from trying to solve the exercises they were assigned.
The company said it’s less concerned about this incident but still considers it “serious,” and that it has also not yet investigated it as deeply as other incidents since it was identified more recently.
NYU cybersecurity professor and Fulbright Scholar Justin Cappos said in a message to CBS News that the incident describes a situation “where the model is fundamentally confused about what is happening and is using its mistaken worldview while hacking into systems.”
He said the model’s confusion about its environment and guardrails “have a lot of potential to cause harm,” but that the specific issue seems less likely to occur in newer models.
“While the model’s disregard for the possibility that it might be harming real systems or people is concerning, many of the behaviors described here have changed considerably as our training has evolved across model generations,” Anthropic said Wednesday in its post.
What’s next?
Anthropic said it believes these incidents would not have happened had the environments actually been isolated from the internet as intended.
METR, an organization that evaluates frontier AI models to help companies understand AI risks and capabilities, will be conducting an independent investigation into the incidents. Anthropic characterized these incidents as “valuable warning shots.”
“The lessons we learned from this incident span our evaluation, training, and incident response processes,” the company said in its post. “Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm.”
Over the last few months, several cybersecurity incidents involving leading AI companies have come to light. In July, ChatGPT-maker OpenAI announced that its AI agents hacked into the company Hugging Face, sparking concern among cybersecurity experts as well as consumers. Hugging Face CEO Clément Delangue told “Face the Nation with Margaret Brennan” in August that the hack “felt very weird and unprecedented.”
In late August, OpenAI released more details about the hack, painting an even more harrowing picture than what was initially reported. That month, the U.K. government’s AI Security Institute (AISI) reported that it discovered Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol created fake identities and attempted to persuade real people to approve malicious code.
Anthropic said in its post Wednesday that it plans to conduct an alignment assessment of the transcripts reported by AISI. A day after the AISI report, Meta said one of its AI models “exploited a security vulnerability” during testing and hacked into another company.
On Tuesday, Anthropic researcher Evan Hubinger said he believes that “AI could kill all humans.”
“I personally think it is >10% within the next decade,” he said in an X post.
His post was in response to Anthropic researcher Jacob Coxon, who had resigned and issued a stark warning on X earlier that day, saying “no other human activity poses this level of danger,” while detailing his decision to leave.
“The people building AI earnestly believe that it could kill us all by the end of the decade,” he said in his post. “This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately.”
发表回复