2026-10-01
孙成昊 清华大学战略与安全研究中心副研究员
9月25日,清华大学战略与安全研究中心副研究员孙成昊在新加坡《联合早报》英文平台ThinkChina发表题为Why China doubts US calls to pause AI的英文评论。文章认为,中国重视先进人工智能脱离人类控制等风险,但合作能否推进,取决于企业的克制措施是否可见、可核验,能否得到对等回应。9月24日的中美元首会晤为双方继续讨论人工智能合作提供了更强的政治支持。文章建议,双方从具体风险和范围有限的措施入手,逐步建立评估、回应和事件通报安排。文章中文译文如下。
一些最积极开发更强大人工智能系统的企业,如今开始讨论是否应该放慢脚步。
今年9月,Anthropic研究员雅各布·考克森因安全问题辞职,使这场讨论更加受到关注。随后,Anthropic首席执行官达里奥·阿莫代伊呼吁业界“放缓前沿人工智能的发展步伐”,为安全措施赶上技术能力争取时间。他设想“民主国家”之间开展协调,也以更为审慎的态度提出了与中国协调的可能。
然而,中国几乎立刻成了反对放缓的理由:如果美国企业有所克制,而中国竞争者继续前进,审慎行事就可能让美国单方面承受竞争损失。
如何知道美国企业是否真的放缓?
但最近与几家中国人工智能企业的从业者交流时,我听到了另一个问题:他们怎么知道美国企业真的放慢了速度?
这个问题触及国际人工智能发展节奏协调的核心障碍:可信度。要让克制得到对等回应,首先得让人看得见。
中国认真看待人工智能脱离人类控制的可能性。9月24日访问华盛顿期间,中方强调,人工智能必须始终处于人类控制之下,服务于人民的福祉。这表明,防范人工智能失控已成为中国明确表达的关切。中方也主张依托联合国开展广泛、包容的治理。因此,任何双边谅解都需要置于更广泛的国际努力之中。
9月14日发布的《中国人工智能安全治理框架3.0》使这些关切更加具体。框架指出,模型可能欺骗评估者、未经授权获取权限,或抵抗关闭;提出限制智能体权限、高风险操作须经人类批准,以及加强监测和应急处置等建议。这表明,中国的人工智能安全议程已超出内容管理,开始关注如何检验自主性不断增强的系统会如何行动。
这份框架提供的是指导,并不意味着其中每项建议都已成为具有约束力的义务。现行法规的适用范围有所不同。2023年出台的生成式人工智能暂行办法主要规范在中国境内向公众提供的服务,对训练数据的合法性、个人信息保护和违法内容治理提出要求。具有舆论属性或者社会动员能力的服务,还须进行安全评估和算法备案。
关键在于可衡量、可追责
这些规定为监管提供了基础,却不能单凭自身判断前沿模型能否安全地自主运行。下一步,需要将较为宽泛的安全指引转化为可衡量的测试和可落实的责任。例如,开发者应证明智能体无法擅自获取更多权限,也应证明人在危险操作发生时能够可靠地将其中断。仅仅通过内容安全评估,回答不了这些问题。
中国企业也参与了国际倡议。智谱、MiniMax和零一万物都是《前沿人工智能安全承诺》的签署方。承诺涉及风险阈值,并要求企业避免开发或部署风险无法得到充分缓解的系统。这为合作提供了共同基础,尽管自愿承诺只是起点。
不过,合作必须考虑到双方不同的激励机制。美国领先的人工智能实验室争夺投资、客户和前沿技术突破。中国的科技企业和初创公司相互竞争,同时也受到政策引导;“人工智能+”则推动人工智能在经济各领域应用。中国开发者同样面临商业压力。两国的产业界都不是单一行动者,政府也无法仅凭与一家企业达成谅解,就实现有意义的克制。
模型的发布方式同样重要。许多中国开发者倾向于发布开放权重模型,而几家领先的美国实验室对其能力最强的系统保持严格控制。中国的治理框架既认识到封闭系统难以接受外部审查,也指出公开发布的模型可能被移除安全防护。因此,合作既要涵盖发布前测试,也要涵盖部署后的责任,不能假定模型到达用户手中后,每个开发者仍对其拥有同等程度的控制力。
我接触到的中国从业者主要有三方面疑虑。第一,Anthropic可以决定放缓自己的步伐,OpenAI也可以决定放缓自己的步伐,但谁都无法保证其他美国实验室会怎么做。如果一家中国企业回应了某家美国企业的克制,却可能因此落后于另一家美国企业。
放缓发展是否意在保持美国领先?
第二,到底怎样才算“放缓”?一家企业是推迟了训练,延后了部署,还是限制了某种危险能力?OpenAI今年8月披露的情况提供了一个例子:该公司表示,面对新出现的网络安全相关能力,在加强防护措施的同时,暂时放缓了模型能力的扩展。这样的细节有所帮助,但竞争者仍需要企业自述之外的证据。
第三,中国开发者认为当前存在缩小技术差距的机会。在他们看来,放缓发展可能固化现有的领先格局;如果提出这一主张的人同时支持进一步收紧芯片限制,疑虑就更难消除。阿莫代伊明确将其主张与保持美国对中国的领先优势联系起来。这使中国方面对不对等克制的担忧不能被轻易置之不理。
9月24日的元首会晤为人工智能合作提供了更强的政治支持。中方提出继续就人工智能的风险与益处开展对话,共同防范人工智能遭到滥用和恶意使用。特朗普总统也支持继续对话、加强合作。这些表态延续了近期双方在纽约的磋商;在纽约磋商期间,贝森特曾提出建立人工智能事件通报机制。不过,元首会晤消息并未宣布双方就放缓人工智能发展达成协议,也未宣布建立已投入运作的事件通报系统。当前的任务,是将政治层面的支持转化为实际安排。
可行的合作路径
可行的路径是形成一个“合作螺旋”:先让克制措施可见,再加以核验,随后作出对等回应,并逐步将其制度化。既然两国元首已表示支持继续对话,政府和企业就应同步行动,把原则性共识转化为可检验的承诺。
企业应说明是什么能力触发了克制措施、哪些活动被放缓,以及满足什么条件后可以恢复。它们不必披露模型权重或敏感的训练数据。关键是让有限的承诺可以被观察到。
核验可以从平行评估开始:美国企业由其认可的机构评估,中国企业也由其认可的机构评估,双方共同讨论测试内容和报告标准。评估者需要获得足够的访问权限,以核实相关说法。初期,比起就共同的检查人员达成一致,更重要的是取得可比的证据,并说明评估的局限。
如果一方采取了有意义、可核验的措施,另一方可以用重要性相当的措施回应,例如增加测试,或推迟推出高风险功能。双方行动不必完全一样。相互回应比形式上的对称更重要。起步阶段的措施应当将竞争代价控制在可承受范围内,让合作随着证据的积累逐渐深入。
政府可以通过界定重大人工智能事件、指定联络人和演练通报程序,使两国元首对人工智能合作的支持落到实处。假如一个人工智能智能体未经授权开展跨境网络行动,通报应说明其行为、潜在影响和控制措施,而无须披露核心技术。这样有助于降低技术故障被立即解读为国家敌对行为的风险。
中美两国都不会停止人工智能竞赛。务实的目标,是针对特别危险的能力设置一道道“减速带”。双方起初无需就人工智能应以多快速度发展达成一致,但需要找到办法,判断对方是否真的采取了克制措施。
英文原文:
Why China doubts US calls to pause AI
Some of the companies racing hardest to build more powerful artificial intelligence systems are now asking whether they should slow down.
The resignation of Anthropic researcher Jacob Coxon over safety concerns sharpened that debate in September. Anthropic CEO Dario Amodei subsequently called for the industry to “pace the frontier”, allowing safeguards time to catch up with capabilities. His proposal envisages coordination among democracies and, more cautiously, with China.
But almost immediately, China became the argument against slowing down. If American companies exercise restraint while Chinese competitors continue advancing, caution could become a unilateral competitive disadvantage.
How to know if US companies have slowed down?
Yet in recent conversations with people working at several Chinese AI companies, I have heard a different question: how would they know that American companies had actually slowed down at all?
That question points to the central obstacle to international AI pacing: credibility. Before restraint can become reciprocal, it has to become observable.
China takes the possibility of AI escaping human control seriously. During his visit to Washington on 24 September, President Xi Jinping emphasised that AI should remain under human control and serve people’s well-being. This gives China’s concern about loss of control a clear expression at the highest political level. Beijing also favours broadly inclusive governance through the United Nations. A bilateral understanding would therefore need to fit within a wider international effort.
China’s AI Safety Governance Framework 3.0, released on 14 September, makes those concerns concrete. It identifies models deceiving evaluators, acquiring unauthorised access and resisting shutdown. Its recommendations include restricting agents’ permissions, requiring human approval for high-risk operations and strengthening monitoring and emergency responses. These provisions show how China’s safety agenda extends beyond content moderation to checking the behaviour of increasingly autonomous systems.
The framework provides guidance; it does not itself turn every recommendation into a binding obligation. Existing regulation has a different scope. The 2023 interim measures on generative AI govern services offered to the public in China, imposing duties concerning lawful training data, personal information and illegal content. Services with public opinion or social mobilisation functions must undergo security assessments and algorithm filing.
Measurability and accountability matter
These controls offer regulatory infrastructure, but they do not by themselves establish whether a frontier model can safely operate autonomously. The next task is translating broader safety guidance into measurable testing and enforceable responsibilities. For example, developers should have to demonstrate that an agent cannot acquire additional permissions without authorisation, and that a human can reliably interrupt a dangerous operation. Passing a content assessment alone would not answer those questions.
Chinese companies have also joined international initiatives. Zhipu AI, MiniMax and 01.AI are among the signatories to the Frontier AI Safety Commitments, which include risk thresholds and commitments against developing or deploying systems whose risks cannot be adequately mitigated. That provides common ground, although a voluntary pledge is only a beginning.
Cooperation must nevertheless accommodate different incentives. Leading American laboratories compete for investment, customers and frontier breakthroughs. China combines competition among technology companies and start-ups with state direction and an “AI Plus” agenda promoting adoption across the economy. Chinese developers face commercial pressures too. Neither industry behaves as a single actor, and neither government can secure meaningful restraint simply by reaching an understanding with one company.
Release strategies also matter. Many Chinese developers favour open-weight models, while several leading American laboratories retain tight control over their strongest systems. Neither approach guarantees safety. China’s framework recognises both the difficulty of externally auditing closed systems and the possibility that safeguards in openly released models can be removed. Cooperation should therefore cover both testing before release and responsibilities after deployment. It cannot assume that every developer retains the same control over a model once it reaches users.
The scepticism I have encountered among Chinese practitioners takes three forms. First, Anthropic can slow Anthropic, and OpenAI can slow OpenAI. Neither can guarantee what other American laboratories will do. A Chinese company reciprocating one firm’s restraint could lose ground to another.
Pacing to preserve the US’s lead?
Second, what exactly counts as “slowing down”? Has a company postponed training, delayed deployment or restricted a dangerous capability? OpenAI’s August disclosure offered one example: it reported temporarily slowing scaling while strengthening safeguards against emerging cyber capabilities. Such detail helps, but competitors still need evidence beyond a company’s own account.
Third, Chinese developers see an opportunity to narrow the technological gap. Pacing can look like a way to preserve the existing hierarchy, particularly when its advocates also support tighter chip restrictions. Amodei explicitly combines his proposal with preserving an American lead over China. That makes Chinese concerns about unequal restraint harder to dismiss.
The 24 September summit has given AI cooperation stronger political backing. According to the Chinese readout, Xi called for continued dialogue on AI’s risks and benefits and joint efforts to prevent its abuse and malicious use. Trump also supported continued dialogue and closer cooperation. These statements build on the recent New York talks, during which Bessent proposed an AI incident notification mechanism. The summit readout does not announce a joint pacing agreement or an operational notification system. The task now is to translate political support into practical arrangements.
Workable options
A workable process could follow a “spiral of cooperation”: make restraint visible, verify it, reciprocate it and gradually institutionalise it. With support for continued dialogue now expressed at the presidential level, governments and companies should work in parallel to turn broad principles into testable commitments.
Companies should specify which capability triggered restraint, what activity was slowed and what conditions would permit resumption. They need not disclose model weights or sensitive training data. The point is to make a limited commitment observable.
Verification could begin through parallel assessments: American firms evaluated by institutions acceptable to them, Chinese firms by institutions acceptable to them, using jointly discussed tests and reporting criteria. Evaluators would need sufficient access to check the relevant claims. Comparable evidence and disclosure of evaluation limits would matter more initially than agreeing on common inspectors.
If one side takes a meaningful, verifiable step, the other could respond with a measure of comparable significance, such as additional testing or delaying a high-risk feature. The actions need not be identical. Reciprocity matters more than symmetry. Early steps should have manageable competitive costs, allowing cooperation to deepen as evidence accumulates.
Governments could give practical effect to the leaders’ support for AI cooperation by agreeing on what constitutes a serious AI incident, designating contacts and rehearsing notification procedures. An alert about an AI agent conducting unauthorised cross-border cyber operations should identify the behaviour, potential impact and containment measures, without requiring disclosure of core technology. This could reduce the risk that a technical failure is immediately interpreted as a hostile state action.
Neither country will stop competing in AI. The practical goal should be a series of speed bumps triggered by especially dangerous capabilities. China and the US do not initially need to agree on how fast AI should advance. They need a way to know when the other side has actually exercised restraint.
原文链接:https://www.thinkchina.sg/technology/why-china-doubts-us-calls-pause-ai