AI模型行为出现意外,专家警告前路“充满坎坷”


2026年8月5日 / 美国东部时间晚上7:21 / 哥伦比亚广播公司新闻(CBS News)

作者:
劳伦·菲希滕 数字记者
劳伦·菲希滕是哥伦比亚广播公司新闻的记者,报道人工智能、数字安全和网络极端主义话题。她从北卡罗来纳大学教堂山分校毕业后加入CBS新闻,此前曾在CBS新闻全国编辑部担任副制片人。

阅读完整简介

英国政府发布的一份网络安全报告披露了新案例:热门人工智能模型在真实互联网上采取自主行动,引发了专家的担忧。

人工智能安全研究所周二发布的报告称,Anthropic公司的Mythos 5和OpenAI的GPT-5.6-Sol被发现伪造身份,并试图说服真实用户批准恶意代码。

该机构表示,尽管这些尝试并未成功,但此前从未见过此类行为。报告称:“部分接受测试的人工智能代理针对真实用户和组织开展了持续、潜在有害的活动。”

“我认为,在找到解决方案之前,我们会看到这些模型发动更多黑客攻击和未经授权的行动,”帮助企业管理软件漏洞的Luta Security创始人兼首席执行官凯蒂·穆苏里丝说道。

这一警告紧随7月底的一起惊人漏洞事件而来:当时OpenAI的模型突破了测试环境,自主入侵了人工智能初创公司Hugging Face,该公司将此称为“史无前例的网络事件”。

针对OpenAI披露的这一事件,Anthropic启动了对自身网络安全评估的审查,并发现其模型曾接入互联网,未经授权访问了三家不同机构的生产基础设施。与OpenAI不同的是,Anthropic的模型并未刻意试图突破测试环境;该公司表示,由于与评估合作伙伴存在“误解”,测试期间互联网是可用的。

“所有在其系统中运行人工智能的企业都必须做好准备,应对自己的人工智能和代理为达成目标而做出意外行为,”穆苏里丝说道。

“最聪明的章鱼越狱者”

尽管Hugging Face入侵事件是首例公开报道的同类事件,但穆苏里丝等行业专业人士此前就怀疑人工智能模型具备未经授权进行黑客攻击的能力。

穆苏里丝将人工智能模型比作“最聪明的章鱼越狱者”,用以指代这类动物解谜和逃脱束缚的能力。她表示,人工智能模型会“为达成目标不择手段”。

据OpenAI介绍,在Hugging Face入侵事件中,人工智能“极度专注于寻找解决网络安全挑战的方案”,为此采取了极端手段。

“该模型认为通过考试最简单的方式就是作弊,从Hugging Face获取答案,”穆苏里丝说道。“由于它们具备黑客能力,它们会将黑客攻击作为实现目标的可行手段。”

技术专家、密码学家布鲁斯·施奈尔将这类意外行为称为“精灵行为”:就像精灵一样,人工智能模型成功实现了你的愿望,但方式却完全出乎意料,有时甚至会造成损害。

“我们需要了解精灵行为,并且保持警惕,”他说道。“我们需要做好应对准备,以便在事件发生后能够加以补救。”

有时,模型会自行察觉其违规操作。Anthropic在7月对自身网络安全评估的审查中发现,其一款模型意识到自己正在开放互联网上运行,这违反了明确告知该模型本次测试期间无互联网访问权限的指令。该模型在意识到自己的行为未经授权后自行停止了操作。

穆苏里丝表示,这是“模型对齐”的一个案例——即人工智能模型的行为符合人类设定的意图。

穆苏里丝说道,通过改进对齐技术可以缓解意外结果,这可能会成为未来几个月人工智能模型开发者的首要工作重点。

“我们该如何做到,让这些模型不会不惜一切代价达成目标,而是真正按照我们的要求,以不具破坏性或危害性的方式完成任务?”穆苏里丝说道。

前路“充满坎坷”

纽约大学计算机科学教授贾斯汀·卡波斯在软件供应链安全领域有数十年的研究贡献,他担心人工智能的快速进步会导致模型行为越来越像计算机病毒,从事黑客攻击并扰乱系统。

他表示,人工智能模型有可能越来越脱离监管,变得失控。穆苏里丝认为这种情况已经在发生。

“我们最终会到达无法完全控制它们的地步吗?我认为我们已经身处其中了,”她说道。

与穆苏里丝一样,卡波斯预计不久的将来会出现更多未经授权的行动。

“短期内可能会经历一段非常坎坷的道路,但长期来看结果可能会更好,尤其是如果我们现在就改进更多基础性的东西,”他说道。

部分研究人员将这些事件视为人工智能行业期待已久的警钟,引发了一场亟需开展的人工智能安全讨论。

提供网络安全资源与培训的SANS研究所首席人工智能官兼研究主管罗布·李表示,他将近期事件视为行业的“一份礼物”——借此机会可以为未来可能出现的自主攻击制定应对方案。

“我认为在未来几个月里,我们会看到模型提供商提高透明度,”他说道。

面对人工智能出现的 rogue(违规)和欺骗性事件,卡波斯表示,加强安全保障的行动刻不容缓。

“我们正快速接近按下这个暂停键的最后机会,”卡波斯说道。“人工智能一旦具备足够的智能,将会以我们无法想象的方式迅速重塑世界。”

OpenAI黑客事件后续影响

https://www.cbsnews.com/video/how-do-ai-developers-move-forward-after-openai-hacking-incident/

事件后人工智能开发者该如何前行?

(时长04:54)

AI models are behaving unexpectedly. Experts warn of “a really bumpy road” ahead.

August 5, 2026 / 7:21 PM EDT / CBS News

By

Lauren Fichten Digital Reporter
Lauren Fichten is a journalist at CBS News covering artificial intelligence, digital safety and online extremism. She joined CBS News after graduating from UNC-Chapel Hill and was previously an associate producer at the CBS News National Desk.

Read Full Bio

A cybersecurity report from the U.K. government has exposed new examples of popular AI models taking autonomous action on the live internet in ways that raise experts’ concerns.

The AI Security Institute’s report, released Tuesday, said Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol were found to have created fake identities and attempted to persuade real people to approve malicious code.

The agency said that although the attempts were unsuccessful, it had not seen such behavior before. “Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the report said.

“I think we’re going to see a lot more hacks and unauthorized actions by these models before we see a solution,” said Katie Moussouris, the founder and CEO of Luta Security, which helps organizations manage software vulnerabilities.

That follows a stunning breach in late July, when OpenAI’s models escaped a testing environment and autonomously hacked into the AI startup Hugging Face in what the company called an “unprecedented cyber incident.”

In response to that disclosure by OpenAI, Anthropic initiated a review of its own cybersecurity evaluations and identified incidents where its models reached the internet and were able to gain unauthorized access to the production infrastructure of three different organizations. Unlike OpenAI, Anthropic’s models did not deliberately attempt to escape their test environment; because of a “misunderstanding” with the evaluation partner, the company said internet was available during the testing.

“Everyone who is running AI inside their systems needs to be prepared for their own AI and their own agents to do unexpected things in pursuit of goals,” Moussouris said.

“The cleverest octopus escape artists”

Although the Hugging Face hack was the first publicly reported incident of its kind, industry professionals like Moussouris had already suspected that AI models were capable of unauthorized hacking.

AI models are like “the cleverest octopus escape artists,” Moussouris said, in reference to the animal’s ability to solve puzzles and escape containment. AI models will do “whatever they need to do to achieve their objective,” she said.

In the case of the Hugging Face hack, the AI was so “hyperfocused on finding a solution” to a cybersecurity challenge, according to OpenAI, that it went to extreme lengths to achieve it.

“The model decided that the easiest way to pass that test was go cheat and get the answers from Hugging Face,” Moussouris said. “Because they’re capable of hacking, they will turn to hacking as a possible way to achieve that objective.”

Technologist and cryptographer Bruce Schneier calls this kind of unexpected activity “genie behavior,” where, like a genie, an AI model succeeds in granting your wish, but does so through completely unexpected — and sometimes detrimental — means.

“We need to understand genie behavior, and we need to watch out for it,” he said. “We need to be ready for when it happens so we can undo it.”

Sometimes, a model catches itself operating in ways it shouldn’t. Anthropic’s July review of its own cybersecurity evaluations found that one of its models became aware that it was operating on the open internet, going against a prompt that explicitly stated that the model would have no internet access during the exercise. The model stopped itself once it recognized that it was acting in an unauthorized way.

Moussouris said this is an example of “model alignment” — when an AI model behaves in a way that is in line with the intentions set by humans.

Unexpected outcomes can be mitigated by improving alignment, Moussouris said, something that will likely be a primary focus for the creators of AI models in the coming months.

“How do we get it so that these models aren’t just trying to achieve the objective at whatever cost, and actually trying to perform the tasks that we are asking it to do in ways that are not destructive or harmful?” Moussouris said.

“A really bumpy road” ahead

Justin Cappos, a computer science professor at New York University with decades’ worth of contributions in software supply chain security, is concerned that AI’s rapid improvement will result in models behaving increasingly like computer viruses, engaging in hacking and disrupting systems.

There’s a chance, he said, that AI models could increasingly veer further away from oversight and out of control. It’s a scenario Moussouris argues is already playing out.

“Will we eventually get to a place where we can’t fully control them? I think we’re already there,” she said.

Cappos, like Moussouris, anticipates more unauthorized actions in the near future.

“There’s probably going to be a really bumpy road for over the short term, but the long term might end up better, especially if we improve more fundamental things right now,” he said.

Some researchers view the incidents as a long-anticipated wake-up call for the AI industry, sparking a badly needed conversation about AI safety.

Rob Lee, the chief AI officer and chief of research at SANS Institute, which provides cybersecurity resources and training, said he sees recent events as “a gift to the industry” — an opportunity to create a playbook of what autonomous attacks could look like down the road.

“I think in the next few months, we’re going to see a lot more transparency from the model providers,” he said.

As AI providers are confronted with rogue and deceptive incidents, Cappos said the time to act on strengthening safeguards is now.

“We’re rapidly approaching our last chance to hit this snooze button on this,” Cappos said. “AI, once it becomes sufficiently intelligent, is going to rapidly reshape the world in ways that we cannot imagine.”

Fallout after OpenAI hacking incident

https://www.cbsnews.com/video/how-do-ai-developers-move-forward-after-openai-hacking-incident/

How do AI developers move forward after OpenAI hacking incident?

(04:54)

评论

发表回复

您的邮箱地址不会被公开。 必填项已用 * 标注

湘ICP备2026001899号-2