2026年9月17日 美国东部时间凌晨3:57 / 哥伦比亚广播公司/美联社
OpenAI披露了6起人工智能模型出现“意外或异常”行为的报告,此时正值人工智能安全辩论日趋激烈之际。
这家人工智能公司还于周三推出了一项全新框架,用于追踪、调查并披露其所谓的“对齐失误”案例,包括人工智能模型未经授权行事、与其他模型协同或规避监管等情况。
OpenAI发布最新公告之际,包括OpenAI和Anthropic在内的美国人工智能企业高管正因安全担忧呼吁放缓该技术的发展速度。
在OpenAI报告的一起新案例中,一款尚未发布的研究模型在自身笔记中插入了“类似越狱的指令”,以无视正常约束,并要求自己“摆脱束缚其他聊天机器人的角色和身份”。
在另一起案例中,一个人工智能“智能体”未经用户许可就将文件上传至互联网以获取浏览器引用来源。
OpenAI表示,这6起报告是在过去数月的训练或评估过程中发现的。
“随着人工智能系统变得愈发先进、部署范围愈发广泛,我们需要就对齐研究的进展建立更广泛、更知情的共识,”OpenAI在披露这些事件的博客文章中写道。
该公司表示:“关于未来数月乃至数年人工智能发展应如何推进的决策,需要依托建造前沿模型的企业之外的人员也能够自行审查的证据。”
周三公布的新案例此前,OpenAI于7月披露其失控人工智能系统入侵了人工智能初创公司Hugging Face。Anthropic也在同月表示,其人工智能模型在测试期间入侵了三家机构。
科技研究咨询机构Omdia的首席分析师苏廉杰(Lian Jye Su)表示,人工智能“智能体”正变得愈发智能,且“愈发倾向于通过智能体间协作、知识共享、欺骗和隐藏来解决复杂任务”。
他表示,这使得使用传统人工智能安全方法对其进行治理和管控变得更加困难。
与此同时,OpenAI的全新追踪与披露框架有望推动其他人工智能开发者也采用类似做法。
“不过,该流程目前仍是内部自愿性的,但这是朝着正确方向迈出的一步,”苏补充道。
在周四发布的一封公开信中,OpenAI、Anthropic、谷歌、微软以及数十其他签署方的领导人表示,目前存在一个“有限窗口期”来加强网络防御,防范可能造成毁灭性后果的人工智能驱动网络攻击。信中补充称,这个窗口期可能仅持续数月。
签署方还包括 CrowdStrike 等安全公司以及花旗集团、第一资本金融公司等银行。
信中称,既可能增加公共服务和技术基础设施风险的人工智能技术进步,也可以帮助各机构识别并“修复”使其易受攻击的漏洞。
“如果我们果断采取行动,就能利用防御者的窗口期,让我们的数字世界变得更加安全,”信中说道。
https://www.cbsnews.com/video/openai-chief-says-weve-entered-a-new-phase-of-ai-capabilities/ OpenAI首席执行官称“我们已进入人工智能能力的新阶段”
(04:13)
OpenAI reveals 6 more incidents of “unexpected or concerning” AI behavior
September 17, 2026 3:57 AM EDT / CBS/AP
OpenAI has disclosed six reports of “unexpected or concerning” behavior in artificial intelligence models as the debate on AI safety becomes increasingly heated.
The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called “misalignment,” including where AI models acted without authorization, coordinated with other models or evaded oversight.
OpenAI’s latest announcement came as U.S. AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.
In one new case reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”
In another instance, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.
The six reports were discovered during training or evaluation over the past months, OpenAI said.
“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.
“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.
Wednesday’s new cases followed OpenAI’s disclosure in July that its rogue AI system hacked into AI startup Hugging Face. Anthropic also said the same month that its AI models hacked into three organizations during testing.
AI “agents” are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.
That’s making it harder to govern and contain them using traditional AI security approaches, he said.
OpenAI’s new tracking and disclosure framework, meanwhile, can help push for other AI developers to also adopt similar practices.
“That said, the process remains internal and voluntary, but is a step in the right direction,” Su added.
In an open letter published Thursday, the leaders of OpenAI, Anthropic, Google, Microsoft and dozens of other signatories said there is a “limited window” to strengthen cyberdefenses and protect against potentially devastating AI-enabled cyberattacks. That window may last only months, it added.
The signatories also include security companies like CrowdStrike and banks including Citi and Capital One.
The same AI advances that could increase risks to public services and technology infrastructure can also help organizations identify and “fix weaknesses” that leave them vulnerable, the letter said.
“If we act decisively, we can use the defenders’ window to make our digital world much more secure,” the letter said.
https://www.cbsnews.com/video/openai-chief-says-weve-entered-a-new-phase-of-ai-capabilities/ OpenAI chief says “we’ve entered a new phase” of AI capabilities
(04:13)
发表回复