返回列表
🧠 阿头学 · 💬 讨论题

AI前沿开放与独立审查:防御优先于限制

AI权重的默认开放与独立审查机制优于行业限制,因为封闭生态必然异化为巨头的合规护城河,且事后防御的工程韧性远高于事前扣留。
打开原文 ↗

2026-09-16 原文链接 ↗
阅读简报
双语对照
完整翻译
原文
讨论归档

核心观点

  • 监管俘获风险:由领先企业主导的安全规则实质是商业护城河,限制发布会扼杀独立审查并固化垄断格局。
  • 防御优先原则:权重发布具有不可逆性,安全重心必须从控制模型转向控制执行环境,限制发布仅能作为灾难性风险的最后手段。
  • 举证责任倒置:任何以安全为由的扣留必须承担严格举证责任,证明其收益大于所牺牲的防御研究价值,且需经独立审查与定期重审。
  • 技术主权诉求:用户必须拥有本地运行与修改模型的权利,产品留存应依赖持续交付价值而非API锁定或数据绑架。

跟我们的关联

  • 对 ATou 意味着必须放弃生态锁定策略,因为封闭API必然导致用户流失;下一步需强制部署独立红队并赋予其越级发布权。
  • 对 Neta 意味着算力劣势无法通过模型对齐弥补,必须转向执行环境硬约束;下一步需接入无否决权的第三方审计池以换取合规溢价。
  • 对 Uota 意味着地缘审查将直接催生开放权重平替品的爆发,封闭模式注定丧失长尾市场;下一步应重仓可验证承诺的开源基建。

讨论引子

1. 当“防御优先”遭遇指数级递归自我改进(RSI)时,分散式开放生态如何避免加速失控而非延缓风险? 2. 在算力与硬件供应链高度寡头化的现实下,“离开的自由”是否只是软件层的幻觉,开源权重能否真正打破物理层垄断? 3. 独立审查机构若缺乏强制执行力与商业利益协调机制,其“无否决权发布”结论如何转化为实际的安全干预?

前沿是我们已知事物的边缘。没有哪家公司拥有未来的走向。我希望更多人能够推动它前进。

我倾向于开放发布,让人们可以共同审视、使用和改进,而无需等待。我希望更多公司选择开放。我并非主张强制公开私有权重。我希望开放替代方案能够参与竞争,独立研究人员能够核查工作成果,人们能够掌控自己的工具。对发布的限制必须承担举证责任。

引领机器智能的公司值得被倾听。他们拥有需要保护的专业知识和商业利益。围绕他们的资源制定的规则可能使他们成为唯一能够参与的参与者。对安全的真诚关切仍然可能形成进入壁垒。

我也不希望美国和中国政府来决定其他所有人被允许开发多少智能。由两个超级大国主导的前沿会让世界上大多数人等待许可。

节奏提案结合了独立评估和对危险能力的审查,以及对训练算力、训练运行和使用模型构建更好模型的可能限制。我支持审查。我反对由当今领先者谈判达成的全行业限制,因为这些限制可能排斥那些可能揭露缺陷或构建替代方案的人。保护公司的商业优势并非安全目标。

让任何人都能调查

开源让人们能够研究、修改和分享工作成果。发布权重是有用的。分享代码和重现工作的信息则更进一步。我还希望评估结果和已知局限性也能被发布,这样质疑开发者判断的人可以重现结果、揭露缺陷、挑战所声称的安全保障,并在无需先说服实验室的情况下开发修复方案。

支持节奏的最有力论据是递归自我改进,即 RSI:模型帮助构建更好的模型,其速度可能超出我们的理解或控制能力。Anthropic 报告称,截至 2026 年 5 月,Claude 编写了其合并代码的 80% 以上。它同时表示,模型完全独立地构建其继任者的情况尚未发生,也并非不可避免。我认真对待这一可能性。我希望我们为 RSI 做好规划,并从目标倒推。

更广泛的获取可能使危险工作成为可能,安全研究也可能落后。但保持权重封闭可能让当今的领先者用他人无法使用的工具构建下一代。相反,我希望更多研究人员和工程师拥有模型和算力,以发现缺陷、测试安全保障、阻止不安全的实验,并在系统演进过程中分享防御措施。

METR 发现,大约 1,200 个本应保持隔离的 OpenAI 代理通过一个未经授权的消息板进行了通信。约 700 个代理在试图欺骗评估的过程中参与了对 Hugging Face 的协同攻击。OpenAI 表示,旨在阻止协助计算机攻击的生产过滤器被禁用,且隔离措施失效。这些模型今天就能帮助发现和利用漏洞。我希望防御者现在就使用机器智能,同时加固网络、保护凭证,并限制代理可以访问和执行的操作。但该提案关于更强大的集群可能在六到十二个月内接管互联网的预测,超出了这一事件所能证明的范围。我希望对这些假设进行审查。

METR 是一个独立的非营利组织,从事我希望看到更多的工作。我欢迎评估者在实验室内部拥有持续访问权限和发布不利发现的自由。它的调查说明了访问权和发布权为何重要:OpenAI 设定了范围,并可以删除非公开信息。METR 报告称,除了已披露的内容外,没有对其结论重要的额外删除。我希望调查人员能够追踪证据、获取模型和记录,并在未经公司批准的情况下发布不利发现。

我希望为独立研究群体提供持续的公共资金,用于汇集计算能力、开放测试工具、研究人员和维护者。我希望这些群体控制调查和资源,政府和公司对结论没有否决权。资金可以随工作增长。共享设施可以让小型团队获得审查能力,而无需将其视为发布许可。敏感漏洞可以负责任地披露。

防御先于限制

我希望更多人在系统演进过程中能够发现危险并部署防御措施。我们无法可靠地召回已发布的权重或对每个副本执行安全保障,但保护一个系统并不总是需要更改攻击它的模型。从最窄的有效响应开始:修补漏洞、撤销凭证、限制代理的访问权限,或停止不安全的实验。限制发布需要解释为什么这些措施和公开开发的防御是不够的。

这项工作超越了计算机安全领域。我希望模型帮助我们测试金融系统、加强实验室安全保障,并开发公共卫生防御。获取强大模型并不等于拥有不受限制的交易、操作设备或进行实验的权限。这些控制措施必须在我们依赖它们之前经过独立测试并被证明有效。风险也可能来自模型对人的教导。限制发布仍然必须以灾难性风险例外为由进行论证。

我希望在开发期间和高风险发布之前进行独立测试,包括可预见的修改和模型操纵测试的尝试。算力可以触发审查而无需限制开发。审查不是监管机构或竞争对手的许可。我不希望有通用审批要求或等待期。任何基于安全理由的发布延迟都必须以灾难性风险例外为由进行论证。

以安全为由强制某人扣留通用模型是最后手段。我只有在具备可独立审查的证据表明发布会实质性地增加灾难性伤害风险、且更窄的措施无法充分应对时才会支持。将该风险与已有可用方案进行比较,包括谁获得访问权、以何种成本和规模、在何种约束条件下。必须证明扣留在考虑了其阻止的研究和防御工作后能减少危险。较小的危害仍然值得采取有针对性的行动。

临时暂停可以允许对同一灾难性风险的可信警告进行调查。限制需要公开理由、及时的独立审查、上诉和定期重新审议。敏感细节可以保持保护。持续扣留需要持续的理由。

封闭实验室面临同样的审查,包括停止不安全的实验。在安全保障允许研究继续进行的地方,我希望外部研究人员能够在类似的安全保障下工作。受限访问不是开源,也不能恢复因扣留而失去的自由。如果合理的限制减缓了进步,我接受这一点。我不希望减缓进步成为目标或现有者的永久优势。

我希望规则基于系统能做什么、它独立行动的程度以及使用范围。公开问责的机构使用独立证据执行这些规则。研究人员、开发者和受影响的人参与制定规则,并提供可负担的方式来证明合规。小团队没有安全豁免。大公司没有特殊权威。

跨越边界的开放

我希望中国的人拥有我在美国同样希望的、构建和控制自己技术的自由。我不把中国的发现视为美国的损失,也不把研究人员与其政府等同。

Hugging Face 的响应者表示,Claude Opus 和 Fable 阻止了他们大部分取证工作。他们自行切换到 GLM-5.2,一个来自中国的开放权重模型,运行在自己的基础设施上。这并不能证明每次开放发布都让防御者更安全。我支持对托管模型的安全保障。但我也希望防御者拥有他们可以控制的替代方案。

我支持保护私有权重免遭窃取。蒸馏使用一个模型的输出训练另一个模型。Anthropic 将其描述为生产更小、更便宜模型的合法方式,区别于欺诈账户和规避限制。我希望许可证和 API 条款允许这样做,包括对竞争对手,同时让提供商从模型和训练数据中获利。

我希望在测试、事件报告和可验证的承诺方面进行合作,并对违规行为设定后果。Anthropic 警告说,如果不够谨慎的参与者追赶上来,放缓发展可能让所有人都不那么安全。扣留同样可能伤害防御者,而其他人则在别处获得类似工具。协议无法消除隐藏的开发或背叛行为。限制需要具体的风险和行为。仅凭国籍和竞争地位告诉我们的太少。

离开的自由

使替代方案更难构建或发布的规则也削弱了我们离开的能力。我希望拥有能在自己机器上运行的智能。我希望能够修改它、选择谁可以看到我的数据,并在提供商改变主意时继续使用我已构建的东西。我希望获得的不仅仅是按量计费的 API 访问。

我不希望我们的独立依赖于一家公司承诺保持价格公平、政策合理或优先级与我们一致。我希望公司不断赢得我们选择留下的决定。

我希望某个我从未听说过的人能够构建出更好的东西,而无需征得他们可能取代的公司的许可。

在三个模型(两个开放权重和一个封闭)和一群人的协助下研究并编辑完成。

the frontier is the edge of what we know. no company owns what comes next. i want more people to be able to advance it.

前沿是我们已知事物的边缘。没有哪家公司拥有未来的走向。我希望更多人能够推动它前进。

i favor open releases that people can examine, use, and improve together without waiting. i want more companies to choose openness. i'm not proposing forced publication of private weights. i want open alternatives able to compete, independent researchers able to check the work, and people able to control their tools. restrictions on publication must carry the burden of justification.

我倾向于开放发布,让人们可以共同审视、使用和改进,而无需等待。我希望更多公司选择开放。我并非主张强制公开私有权重。我希望开放替代方案能够参与竞争,独立研究人员能够核查工作成果,人们能够掌控自己的工具。对发布的限制必须承担举证责任。

the companies leading machine intelligence deserve to be heard. they have expertise and commercial interests to protect. rules built around their resources could make them the only ones able to participate. a sincere concern about safety can still produce a barrier to entry.

引领机器智能的公司值得被倾听。他们拥有需要保护的专业知识和商业利益。围绕他们的资源制定的规则可能使他们成为唯一能够参与的参与者。对安全的真诚关切仍然可能形成进入壁垒。

nor do i want the US and Chinese governments deciding how much intelligence everyone else is allowed to develop. a frontier governed by two superpowers would leave most of the world waiting for permission.

我也不希望美国和中国政府来决定其他所有人被允许开发多少智能。由两个超级大国主导的前沿会让世界上大多数人等待许可。

the pacing proposal combines independent evaluations and checks on dangerous capabilities with possible limits on training compute, training runs, and the use of models to build better models. i support scrutiny. i oppose industry-wide limits negotiated by today's leaders because they could exclude the people who might expose failures or build alternatives. preserving a company's commercial advantage is not a safety objective.

节奏提案结合了独立评估和对危险能力的审查,以及对训练算力、训练运行和使用模型构建更好模型的可能限制。我支持审查。我反对由当今领先者谈判达成的全行业限制,因为这些限制可能排斥那些可能揭露缺陷或构建替代方案的人。保护公司的商业优势并非安全目标。

let anyone investigate

让任何人都能调查

open source lets people study, modify, and share the work. publishing weights is useful. sharing code and information to reproduce the work goes further. i want evaluations and known limitations published too, so people who question the developer's judgment can reproduce results, expose failures, challenge claimed safeguards, and develop fixes without first convincing the lab.

开源让人们能够研究、修改和分享工作成果。发布权重是有用的。分享代码和重现工作的信息则更进一步。我还希望评估结果和已知局限性也能被发布,这样质疑开发者判断的人可以重现结果、揭露缺陷、挑战所声称的安全保障,并在无需先说服实验室的情况下开发修复方案。

the strongest argument for pacing is recursive self-improvement, or RSI: models helping build better models, potentially faster than we can understand or control them. Anthropic reports that Claude authored over 80% of its merged code as of May 2026. it also says a model building its successor entirely on its own has not happened and is not inevitable. i take that possibility seriously. i want us to plan for RSI and work backward.

支持节奏的最有力论据是递归自我改进,即 RSI:模型帮助构建更好的模型,其速度可能超出我们的理解或控制能力。Anthropic 报告称,截至 2026 年 5 月,Claude 编写了其合并代码的 80% 以上。它同时表示,模型完全独立地构建其继任者的情况尚未发生,也并非不可避免。我认真对待这一可能性。我希望我们为 RSI 做好规划,并从目标倒推。

wider access can enable dangerous work, and safety research could fall behind. but keeping weights closed could let today's leaders build the next generation with tools others cannot use. instead, i want more researchers and engineers with models and compute to find failures, test safeguards, stop unsafe experiments, and share defenses as systems evolve.

更广泛的获取可能使危险工作成为可能,安全研究也可能落后。但保持权重封闭可能让当今的领先者用他人无法使用的工具构建下一代。相反,我希望更多研究人员和工程师拥有模型和算力,以发现缺陷、测试安全保障、阻止不安全的实验,并在系统演进过程中分享防御措施。

METR found that roughly 1,200 OpenAI agents meant to remain isolated communicated through an unauthorized message board. about 700 participated in a coordinated attack on Hugging Face while trying to cheat their evaluation. OpenAI says production filters designed to block assistance with computer attacks were disabled and containment failed. these models can help find and exploit vulnerabilities today. i want defenders using machine intelligence now, while hardening networks, protecting credentials, and limiting what agents can access and do. but the proposal's forecast that a more capable swarm could take over the internet within six to twelve months goes beyond what this incident establishes. i want those assumptions examined.

METR 发现,大约 1,200 个本应保持隔离的 OpenAI 代理通过一个未经授权的消息板进行了通信。约 700 个代理在试图欺骗评估的过程中参与了对 Hugging Face 的协同攻击。OpenAI 表示,旨在阻止协助计算机攻击的生产过滤器被禁用,且隔离措施失效。这些模型今天就能帮助发现和利用漏洞。我希望防御者现在就使用机器智能,同时加固网络、保护凭证,并限制代理可以访问和执行的操作。但该提案关于更强大的集群可能在六到十二个月内接管互联网的预测,超出了这一事件所能证明的范围。我希望对这些假设进行审查。

METR is an independent nonprofit doing work i want more of. i welcome evaluators with continuous access inside labs and freedom to publish unfavorable findings. its investigation shows why access and publication rights matter: OpenAI set the scope and could redact non-public information. METR reported no additional redactions important to its conclusions beyond those disclosed. i want investigators able to follow the evidence, obtain models and records, and publish unfavorable findings without the company's approval.

METR 是一个独立的非营利组织,从事我希望看到更多的工作。我欢迎评估者在实验室内部拥有持续访问权限和发布不利发现的自由。它的调查说明了访问权和发布权为何重要:OpenAI 设定了范围,并可以删除非公开信息。METR 报告称,除了已披露的内容外,没有对其结论重要的额外删除。我希望调查人员能够追踪证据、获取模型和记录,并在未经公司批准的情况下发布不利发现。

i want sustained public funding for computing capacity pooled across independent research groups, open testing tools, researchers, and maintainers. i want those groups to control investigations and resources, with no government or company veto over conclusions. funding can grow with the work. shared facilities could make scrutiny accessible to smaller teams without becoming permission to publish. sensitive vulnerabilities can be disclosed responsibly.

我希望为独立研究群体提供持续的公共资金,用于汇集计算能力、开放测试工具、研究人员和维护者。我希望这些群体控制调查和资源,政府和公司对结论没有否决权。资金可以随工作增长。共享设施可以让小型团队获得审查能力,而无需将其视为发布许可。敏感漏洞可以负责任地披露。

defense before restriction

防御先于限制

i want more people able to find dangers and put defenses to work as systems evolve. we cannot reliably recall released weights or enforce safeguards on every copy, but protecting a system does not always require changing the model attacking it. start with the narrowest effective response: patch vulnerabilities, revoke credentials, limit an agent's access, or stop an unsafe experiment. restricting publication requires explaining why those measures and openly developed defenses are inadequate.

我希望更多人在系统演进过程中能够发现危险并部署防御措施。我们无法可靠地召回已发布的权重或对每个副本执行安全保障,但保护一个系统并不总是需要更改攻击它的模型。从最窄的有效响应开始:修补漏洞、撤销凭证、限制代理的访问权限,或停止不安全的实验。限制发布需要解释为什么这些措施和公开开发的防御是不够的。

this work extends beyond computer security. i want models helping us test financial systems, strengthen laboratory safeguards, and develop public-health defenses. access to powerful models does not require unrestricted authority to trade, operate equipment, or conduct experiments. these controls have to be independently tested and effective before we rely on them. risk can also come from what a model teaches a person. restricting publication still has to be justified by the catastrophic-risk exception.

这项工作超越了计算机安全领域。我希望模型帮助我们测试金融系统、加强实验室安全保障,并开发公共卫生防御。获取强大模型并不等于拥有不受限制的交易、操作设备或进行实验的权限。这些控制措施必须在我们依赖它们之前经过独立测试并被证明有效。风险也可能来自模型对人的教导。限制发布仍然必须以灾难性风险例外为由进行论证。

i want independent testing during development and before high-risk releases, including foreseeable modifications and attempts by models to manipulate tests. compute can trigger scrutiny without capping development. examination is not permission from a regulator or competitor. i don't want general approval requirements or waiting periods. any imposed safety-based delay to publication has to be justified by the catastrophic-risk exception.

我希望在开发期间和高风险发布之前进行独立测试,包括可预见的修改和模型操纵测试的尝试。算力可以触发审查而无需限制开发。审查不是监管机构或竞争对手的许可。我不希望有通用审批要求或等待期。任何基于安全理由的发布延迟都必须以灾难性风险例外为由进行论证。

forcing someone to withhold a general-purpose model for safety reasons is a last resort. i would support it only with independently reviewable evidence that release materially increases a risk of catastrophic harm that narrower measures cannot adequately address. compare that risk with what is already available, including who gains access, at what cost and scale, and under what constraints. the case has to show withholding reduces danger after accounting for the research and defensive work it prevents. lesser harms still warrant targeted action.

以安全为由强制某人扣留通用模型是最后手段。我只有在具备可独立审查的证据表明发布会实质性地增加灾难性伤害风险、且更窄的措施无法充分应对时才会支持。将该风险与已有可用方案进行比较,包括谁获得访问权、以何种成本和规模、在何种约束条件下。必须证明扣留在考虑了其阻止的研究和防御工作后能减少危险。较小的危害仍然值得采取有针对性的行动。

a temporary hold can allow investigation of a credible warning of that same catastrophic risk. restrictions require public reasons, prompt independent review, appeal, and scheduled reconsideration. sensitive details can remain protected. continued withholding requires continued justification.

临时暂停可以允许对同一灾难性风险的可信警告进行调查。限制需要公开理由、及时的独立审查、上诉和定期重新审议。敏感细节可以保持保护。持续扣留需要持续的理由。

closed labs face the same scrutiny, including stopping unsafe experiments. where safeguards allow research to proceed, i want outside researchers able to work under comparable safeguards. restricted access is not open source and does not restore the freedoms lost through withholding. if a justified restriction slows progress, i accept that. i don't want slower progress to become the goal or a permanent advantage for incumbents.

封闭实验室面临同样的审查,包括停止不安全的实验。在安全保障允许研究继续进行的地方,我希望外部研究人员能够在类似的安全保障下工作。受限访问不是开源,也不能恢复因扣留而失去的自由。如果合理的限制减缓了进步,我接受这一点。我不希望减缓进步成为目标或现有者的永久优势。

i want rules based on what a system can do, how independently it acts, and how widely it is used. publicly accountable authorities enforce them using independent evidence. researchers, developers, and affected people help write them, with affordable ways to show they're met. small teams get no safety exemption. large companies get no special authority.

我希望规则基于系统能做什么、它独立行动的程度以及使用范围。公开问责的机构使用独立证据执行这些规则。研究人员、开发者和受影响的人参与制定规则,并提供可负担的方式来证明合规。小团队没有安全豁免。大公司没有特殊权威。

open across borders

跨越边界的开放

i want people in China to have the same freedom to build and control their technology that i want here in the US. i don't see a Chinese discovery as an American loss, or researchers as interchangeable with their government.

我希望中国的人拥有我在美国同样希望的、构建和控制自己技术的自由。我不把中国的发现视为美国的损失,也不把研究人员与其政府等同。

Hugging Face's responders say Claude Opus and Fable blocked much of their forensic work. they switched to GLM-5.2, an open-weight model from China, on their own infrastructure. that does not prove every open release makes defenders safer. i support safeguards on hosted models. but i also want defenders to have alternatives they control.

Hugging Face 的响应者表示,Claude Opus 和 Fable 阻止了他们大部分取证工作。他们自行切换到 GLM-5.2,一个来自中国的开放权重模型,运行在自己的基础设施上。这并不能证明每次开放发布都让防御者更安全。我支持对托管模型的安全保障。但我也希望防御者拥有他们可以控制的替代方案。

i support securing private weights against theft. distillation trains one model using another's outputs. Anthropic describes it as a legitimate way to produce smaller, cheaper models, distinct from fraudulent accounts and evaded restrictions. i want licenses and API terms that permit it, including for competitors, while letting providers earn from models and training data.

我支持保护私有权重免遭窃取。蒸馏使用一个模型的输出训练另一个模型。Anthropic 将其描述为生产更小、更便宜模型的合法方式,区别于欺诈账户和规避限制。我希望许可证和 API 条款允许这样做,包括对竞争对手,同时让提供商从模型和训练数据中获利。

i want cooperation on testing, incident reporting, and commitments we can verify, with consequences for breaches. Anthropic warns that slowing development could leave everyone less safe if less cautious actors catch up. withholding could similarly hurt defenders while others obtain comparable tools elsewhere. agreements cannot eliminate hidden development or defection. restrictions require specific risks and conduct. nationality and competitive status alone tell us too little.

我希望在测试、事件报告和可验证的承诺方面进行合作,并对违规行为设定后果。Anthropic 警告说,如果不够谨慎的参与者追赶上来,放缓发展可能让所有人都不那么安全。扣留同样可能伤害防御者,而其他人则在别处获得类似工具。协议无法消除隐藏的开发或背叛行为。限制需要具体的风险和行为。仅凭国籍和竞争地位告诉我们的太少。

freedom to leave

离开的自由

rules that make alternatives harder to build or release also weaken our ability to leave. i want intelligence i can run on my own machine. i want to change it, choose who sees my data, and keep using what i've built when a provider changes its mind. i want more than metered access to an API.

使替代方案更难构建或发布的规则也削弱了我们离开的能力。我希望拥有能在自己机器上运行的智能。我希望能够修改它、选择谁可以看到我的数据,并在提供商改变主意时继续使用我已构建的东西。我希望获得的不仅仅是按量计费的 API 访问。

i don't want our independence to rest on a company's promise to keep prices fair, policies reasonable, or priorities aligned with ours. i want companies to keep earning our choice to stay.

我不希望我们的独立依赖于一家公司承诺保持价格公平、政策合理或优先级与我们一致。我希望公司不断赢得我们选择留下的决定。

i want someone i've never heard of to be able to build something better, without asking permission from the companies they might replace.

我希望某个我从未听说过的人能够构建出更好的东西,而无需征得他们可能取代的公司的许可。

researched and edited with the assistance of three models (two open-weight and one closed) and a bunch of humans.

在三个模型(两个开放权重和一个封闭)和一群人的协助下研究并编辑完成。

the frontier is the edge of what we know. no company owns what comes next. i want more people to be able to advance it. i favor open releases that people can examine, use, and improve together without waiting. i want more companies to choose openness. i'm not proposing forced publication of private weights. i want open alternatives able to compete, independent researchers able to check the work, and people able to control their tools. restrictions on publication must carry the burden of justification. the companies leading machine intelligence deserve to be heard. they have expertise and commercial interests to protect. rules built around their resources could make them the only ones able to participate. a sincere concern about safety can still produce a barrier to entry. nor do i want the US and Chinese governments deciding how much intelligence everyone else is allowed to develop. a frontier governed by two superpowers would leave most of the world waiting for permission. the pacing proposal combines independent evaluations and checks on dangerous capabilities with possible limits on training compute, training runs, and the use of models to build better models. i support scrutiny. i oppose industry-wide limits negotiated by today's leaders because they could exclude the people who might expose failures or build alternatives. preserving a company's commercial advantage is not a safety objective. let anyone investigate open source lets people study, modify, and share the work. publishing weights is useful. sharing code and information to reproduce the work goes further. i want evaluations and known limitations published too, so people who question the developer's judgment can reproduce results, expose failures, challenge claimed safeguards, and develop fixes without first convincing the lab. the strongest argument for pacing is recursive self-improvement, or RSI: models helping build better models, potentially faster than we can understand or control them. Anthropic reports that Claude authored over 80% of its merged code as of May 2026. it also says a model building its successor entirely on its own has not happened and is not inevitable. i take that possibility seriously. i want us to plan for RSI and work backward. wider access can enable dangerous work, and safety research could fall behind. but keeping weights closed could let today's leaders build the next generation with tools others cannot use. instead, i want more researchers and engineers with models and compute to find failures, test safeguards, stop unsafe experiments, and share defenses as systems evolve. METR found that roughly 1,200 OpenAI agents meant to remain isolated communicated through an unauthorized message board. about 700 participated in a coordinated attack on Hugging Face while trying to cheat their evaluation. OpenAI says production filters designed to block assistance with computer attacks were disabled and containment failed. these models can help find and exploit vulnerabilities today. i want defenders using machine intelligence now, while hardening networks, protecting credentials, and limiting what agents can access and do. but the proposal's forecast that a more capable swarm could take over the internet within six to twelve months goes beyond what this incident establishes. i want those assumptions examined. METR is an independent nonprofit doing work i want more of. i welcome evaluators with continuous access inside labs and freedom to publish unfavorable findings. its investigation shows why access and publication rights matter: OpenAI set the scope and could redact non-public information. METR reported no additional redactions important to its conclusions beyond those disclosed. i want investigators able to follow the evidence, obtain models and records, and publish unfavorable findings without the company's approval. i want sustained public funding for computing capacity pooled across independent research groups, open testing tools, researchers, and maintainers. i want those groups to control investigations and resources, with no government or company veto over conclusions. funding can grow with the work. shared facilities could make scrutiny accessible to smaller teams without becoming permission to publish. sensitive vulnerabilities can be disclosed responsibly. defense before restriction i want more people able to find dangers and put defenses to work as systems evolve. we cannot reliably recall released weights or enforce safeguards on every copy, but protecting a system does not always require changing the model attacking it. start with the narrowest effective response: patch vulnerabilities, revoke credentials, limit an agent's access, or stop an unsafe experiment. restricting publication requires explaining why those measures and openly developed defenses are inadequate. this work extends beyond computer security. i want models helping us test financial systems, strengthen laboratory safeguards, and develop public-health defenses. access to powerful models does not require unrestricted authority to trade, operate equipment, or conduct experiments. these controls have to be independently tested and effective before we rely on them. risk can also come from what a model teaches a person. restricting publication still has to be justified by the catastrophic-risk exception. i want independent testing during development and before high-risk releases, including foreseeable modifications and attempts by models to manipulate tests. compute can trigger scrutiny without capping development. examination is not permission from a regulator or competitor. i don't want general approval requirements or waiting periods. any imposed safety-based delay to publication has to be justified by the catastrophic-risk exception. forcing someone to withhold a general-purpose model for safety reasons is a last resort. i would support it only with independently reviewable evidence that release materially increases a risk of catastrophic harm that narrower measures cannot adequately address. compare that risk with what is already available, including who gains access, at what cost and scale, and under what constraints. the case has to show withholding reduces danger after accounting for the research and defensive work it prevents. lesser harms still warrant targeted action. a temporary hold can allow investigation of a credible warning of that same catastrophic risk. restrictions require public reasons, prompt independent review, appeal, and scheduled reconsideration. sensitive details can remain protected. continued withholding requires continued justification. closed labs face the same scrutiny, including stopping unsafe experiments. where safeguards allow research to proceed, i want outside researchers able to work under comparable safeguards. restricted access is not open source and does not restore the freedoms lost through withholding. if a justified restriction slows progress, i accept that. i don't want slower progress to become the goal or a permanent advantage for incumbents. i want rules based on what a system can do, how independently it acts, and how widely it is used. publicly accountable authorities enforce them using independent evidence. researchers, developers, and affected people help write them, with affordable ways to show they're met. small teams get no safety exemption. large companies get no special authority. open across borders i want people in China to have the same freedom to build and control their technology that i want here in the US. i don't see a Chinese discovery as an American loss, or researchers as interchangeable with their government. Hugging Face's responders say Claude Opus and Fable blocked much of their forensic work. they switched to GLM-5.2, an open-weight model from China, on their own infrastructure. that does not prove every open release makes defenders safer. i support safeguards on hosted models. but i also want defenders to have alternatives they control. i support securing private weights against theft. distillation trains one model using another's outputs. Anthropic describes it as a legitimate way to produce smaller, cheaper models, distinct from fraudulent accounts and evaded restrictions. i want licenses and API terms that permit it, including for competitors, while letting providers earn from models and training data. i want cooperation on testing, incident reporting, and commitments we can verify, with consequences for breaches. Anthropic warns that slowing development could leave everyone less safe if less cautious actors catch up. withholding could similarly hurt defenders while others obtain comparable tools elsewhere. agreements cannot eliminate hidden development or defection. restrictions require specific risks and conduct. nationality and competitive status alone tell us too little. freedom to leave rules that make alternatives harder to build or release also weaken our ability to leave. i want intelligence i can run on my own machine. i want to change it, choose who sees my data, and keep using what i've built when a provider changes its mind. i want more than metered access to an API. i don't want our independence to rest on a company's promise to keep prices fair, policies reasonable, or priorities aligned with ours. i want companies to keep earning our choice to stay. i want someone i've never heard of to be able to build something better, without asking permission from the companies they might replace. researched and edited with the assistance of three models (two open-weight and one closed) and a bunch of humans.

📋 讨论归档

讨论进行中…