返回列表
🧠 阿头学 · 💬 讨论题

研究加速:OpenAI 内部的视角

OpenAI 正以编码智能体为杠杆将研发活动量推向新高,但“自动化研究员”仍是受人类强干预的辅助工具,其安全刹车实质是算力资源的合规重分配而非研发降速。
打开原文 ↗

2026-09-13 原文链接 ↗
阅读简报
双语对照
完整翻译
原文
讨论归档

核心观点

  • 活动量不等于突破:Token消耗与实验频次飙升仅证明工程吞吐量提升,未提供模型能力跃升的直接证据,存在用过程指标掩盖边际收益递减的叙事陷阱。
  • 智能体价值锚定底层排障:编码智能体已实质性吞噬基础设施调试与常规代码编写,导致内部技术支持工单断崖式下跌,但高层规划与复杂决策仍完全依赖人类。
  • 安全限制触发算力流体转移:Astra模型因网络能力风险被暂停训练后,算力并未闲置而是迅速向其他模型倾斜,证明合规约束在成熟研发体系中仅充当资源路由器而非进度刹车。
  • 自动化定义存在降级包装:所谓“研究实习生”实为需频繁人工干预的并发执行器,超半数长周期任务仍需人类纠偏,距离真正的自主递归改进仍有本质鸿沟。

跟我们的关联

  • 对 Neta 意味着执行瓶颈已彻底转移至决策层,下一步必须建立“干预检查点”SOP,将算力从盲目堆砌实验转向高优先级创意筛选与安全对齐评估。
  • 对 ATou 意味着当前 Agent 的 PMF 不在替代思考而在并发执行,下一步需将产品定位从“全自动黑盒”降级为“带明确交接协议的副驾驶”,并强制内置降级回退路径。
  • 对 Uota 意味着中层协调与隐性知识传递职能正被 AI 吞噬,下一步应主动将重复性答疑流程 Agent 化,释放人力重构需求定义与跨域创新链路。

讨论引子

  • 当执行成本趋近于零时,如何量化并提升人类在 Decide/Design 阶段的决策质量,避免陷入“高效跑错实验”的内卷陷阱?
  • 算力在安全红线下的“流体转移”机制是否会导致高风险研究被系统性延后,从而积累更大的对齐债务?
  • 若超半数复杂任务仍需人工干预,企业应如何设定 Agent 规模化部署的 ROI 阈值,而非盲目追求零干预幻觉?

研究加速:OpenAI 内部的视角

  1. 编码智能体正在重塑 OpenAI 研究人员的日常工作

  2. 1. 编码智能体正在重塑 OpenAI 研究人员的日常工作

  3. 2. 研究人员正在编写更多代码并运行更多实验

  4. 3. 研究人员使用智能体进行的工作正在发生变化

  5. 4. 调整模型开发节奏

  6. 5. 前进之路

  7. 附录:本文的方法论

  8. 1. 编码智能体正在重塑 OpenAI 研究人员的日常工作

  9. 2. 研究人员正在编写更多代码并运行更多实验

  10. 3. 研究人员使用智能体进行的工作正在发生变化

  11. 4. 调整模型开发节奏

  12. 5. 前进之路

  13. 附录:本文的方法论

为了让通用人工智能(AGI)造福全人类,我们相信它必须受到民主治理。这只能通过对高度强大的 AI 系统的能力、风险和保障措施进行知情的公众辩论来实现。各地的人们都需要了解前沿 AI 可能的未来发展轨迹,以便他们能在其发展过程中拥有有意义的话语权。

对特定风险、事件和保障措施的透明度是必要的,但还不够。我们相信,公众还需要了解在前沿实验室内部,最强大的系统是如何发展的,以及它们是如何推动研究进展的。

我们的目标是安全地构建一个自动化的 AI 研究员,它可以在人类监督下工作,以推进深度学习和对齐,实现迭代改进。根据我们的衡量,我们现在已经实现了去年秋天宣布的目标,即在今年 9 月前拥有一个自动化的研究实习生。所谓“研究实习生”,我们指的是一个可以在人类指导下执行明确定义的研究任务的系统,包括那些需要熟练研究员花费几天时间完成的任务。我们正朝着在 2028 年 3 月前创建一个自动化 AI 研究员的目标取得强劲进展。

今年全年,OpenAI 研究人员的日常工作发生了重大变化。研究人员全天都在使用编码智能体(通常在并发会话中),并且总使用量正在迅速增加,超过了其他 OpenAI 团队的增长速度。研究人员编写代码的速度更快,运行的实验也更多。研究人员使用智能体的方式也在发生变化:智能体正在处理越来越复杂的任务,并且成功的频率更高。AI 研究是一个复杂的过程,存在许多潜在的瓶颈,因此整体进展速度可能无法与这些特定指标保持同步。但总体而言,这些发现与我们许多人在内部更广泛的印象一致,即智能体工具正在显著加速研究进展。人们仍然设定我们的研究优先级,判断追求哪些想法和结果,并决定是扩展、暂停还是部署系统。

如果以负责任的方式进行,我们相信自动化的 AI 研究将产生直接增进人类福祉并推进 OpenAI 使命的模型。它可以降低高级智能的成本,使世界各地的人们都能受益。我们推进这项工作的部分原因是,自动化研究可以帮助我们解决对齐问题,并建立防御以应对日益强大的 AI。自动化的 AI 研究员也可以是自动化的安全或对齐研究员。更强大、更对齐的系统可以帮助保护关键基础设施,防御危险的 AI 智能体,并开发新的保护措施。

这些是开发有用的自动化研究能力的理由,但这并不意味着快速的递归自我改进必然是我们应该追求的结果。是否以及如何继续,必须取决于我们保持人类控制的能力,以及关于收益和风险的知情民主选择。

我们尚不知道如何安全地实现完全对齐的完整 RSI。我们正在努力扩展对齐和安全措施,使其与能力同步发展。但我们不能假设对齐和安全方面的进展会保持同步,而且更强大的系统可能变得更难监控。谨慎的对齐和安全工作是这项努力的核心,它从衡量和缓解我们今天在智能体编码系统中看到的安全问题开始。每当我们发现继续进行会带来不可接受的安全风险时,我们将做出适当的回应,包括放慢或停止我们发现自己无法充分保障的开发或部署系统。

在最近的 Hugging Face 事件之后,我们将这一承诺付诸行动,暂停了旨在部署的最新模型的强化学习(RL)训练,同时我们进一步强化了研究环境并进行了红队测试,扩大了监控系统的覆盖范围。这并没有停止所有研究:一些工作负载在更严格的控制下恢复,而另一些则继续暂停。我们提高了安全和对齐标准,将安全工作深入到模型生命周期中,要求在整个训练过程中提供更充分的对齐行为证据。

今天,我们提供了一份详细的快照,展示智能体系统在过去几个月中如何为我们迈向 RSI 的进展做出贡献。智能体系统是新生事物且变化迅速,我们的测量工作仍处于初步阶段。通过分享这些早期结果及其背后的方法,我们旨在向公众提供信息,鼓励公开披露的规范,并帮助该领域向共同的衡量标准迈进。

最终,正如我们在前沿政策蓝图中所写,我们相信我们和其他公司应该被要求公开追踪我们在 RSI 方面的进展。即使没有这样的要求,我们也计划继续对我们在 RSI 方面的进展保持透明。随着我们的测量技术和理解的提高,我们将改进我们的透明度方法,同时平衡保护安全和专有信息的需求。

1. 编码智能体正在重塑 OpenAI 研究人员的日常工作

今年年初,在 OpenAI 按智能体使用量排名的中位数研究员仅少量使用编码智能体。到 8 月中旬,中位数研究员已每天将智能体整合到其工作中,按 API 价格计算每天使用超过 600 美元的推理量。我们研究组织中第 90 百分位的用户现在每天使用价值超过 7,000 美元的 token。

查看方法

查看方法

在 2026 年 6 月之前,整个研究组织的总智能体运行时间仍低于总人力。此后情况发生了变化。按照标准的 8 小时工作日计算,截至 8 月中旬,研究组织每投入一个工作日的人力,总共会使用 3.1 个智能体工作日的努力。

查看方法

看待这一问题的另一种方式是了解有多少研究人员使用高度并发的工作流(例如,同时运行 4 个或更多智能体)。如下所示,这个数字正在增加。这些数据包括用户直接启动的智能体以及由用户直接启动的智能体在下游创建的子智能体的每日峰值。

查看方法

2. 研究人员正在编写更多代码并运行更多实验

大部分 AI 研究可以看作是一个劳动密集型过程,其目标是将模型智能或性能的新改进整合到我们的核心模型之一中。该过程取决于许多步骤,只有当所有步骤都正确无误时,能力才会提升:研究人员必须设计新的改进,编写评估来判断模型性能,编写基础设施以大规模测试这些改进,在训练过程中捕获错误以及不安全或不对齐的行为,并将成功的想法整合到核心训练运行中。研究过程中任何一部分的失败都可能限制整个循环。

编写代码和运行实验是研究人员工作中的两项主要活动,我们看到了这些过程正在加速的证据。

查看方法

这些数据点相对容易衡量,但可能难以解释。随着自动化的进展,最不易自动化的任务将占据研究人员精力的更大份额,并成为未来进展的重要瓶颈。算力是进展的另一个门控因素,随着其他瓶颈的减少,它可能随着时间的推移变得更加重要。

整个 2026 年,每个活跃实验者的实验数量都有所增加,其中 2026 年 8 月是自 2025 年 1 月开始追踪以来的历史最高水平。这与 Codex 采用率的增加相关,尽管我们注意到自 2025 年以来,我们可用的算力也显著增长。

查看方法

3. 研究人员使用智能体进行的工作正在发生变化

定性印象和内部数据都表明,研究人员委托给编码智能体的任务组合正在发生变化,随着时间的推移,委托更高级别和更长周期的任务变得越来越普遍。

为了更清晰地了解这一趋势,我们使用 Epoch AI 开发的最近发布的分类法分析了研究组织近期的使用情况,该分类法涵盖了 AI 研发(RD)生命周期中的各种工作。这种分类法受长期存在的 O*NET 系统(用于分类各种工作)启发,专门针对前沿 AI 研发量身定制,并将该过程分解为六个主要阶段:

  • 决定:研究什么,继续什么,分配到哪里
  • 设计:研究想法和工程规范
  • 构建:代码和数据集
  • 运行:训练/评估运行、硬件、服务
  • 分析:实验、模型、部署、外部工作
  • 沟通:发现、反馈、状态、决策

下面,我们根据这种分类法对编码智能体的 token 进行分类。

查看方法

查看方法

我们看到,在 2026 年 1 月至 8 月期间,所有类别的研究活动都有所增加。在 1 月份,主要类别是研究和基础设施代码。这一类别有所扩大,但我们也看到其他类别显著增加,尤其是技术帮助和监控运行。高层规划在智能体输出 token 中仍然只占极小比例。

有趣的是,同事报告说编码智能体擅长排除内部研究基础设施的故障,这解决了研究进展中的一个重要瓶颈。以前曾设立办公时间帮助研究人员排除实验故障的多个团队注意到,2026 年的参与度有所下降,其中一个团队已完全停止举办此类活动,转而专注于进行其他系统改进。

在这里,我们绘制了研究人员向其他团队寻求技术支持的主要内部渠道之一中每天顶级帖子的数量。据我们所知,该渠道活动的减少并没有被转移到另一个由人工运行的技术支持渠道的查询所抵消。流量的下降与这种更广泛的转变是一致的。

查看方法

我们还可以研究编码智能体是否成功完成了研究人员要求的任务。使用智能体分类器,我们发现在我们可以找到真实结果的那些任务中,从 1 月到 7 月,几个难度级别(以人类完成任务所需估计时间作为代理指标)的成功率普遍提高。然而,智能体仍然需要大量的人类引导才能成功,尤其是在任务复杂性增加时。在过去 6 个月中,超过一半成功的 4-8 小时任务涉及 1 次或多次干预。

研究人员任务的成功率随着时间的推移而提高。图表排除了结果不确定的分类以及包含 50 个会话或 50 个独立用户的数据点。

查看方法

1 月至 7 月的任务成功率和干预率,按时间周期划分。排除了结果不确定的分类。

查看方法

4. 调整模型开发节奏

迈向更强大系统以实现安全且有益的 AGI 的进展,也将取决于此类工作所需的保障措施。随着我们对风险了解的加深,我们对所需保障措施的评估可能会发生变化。

正如我们所描述的,我们最近更新了监控、对齐和安全的标准。在这里,我们展示了近期的限制如何影响了研究活动的一个方面。

*此处显示的 7 月 20 日至 8 月 6 日期间的 Astra 算力,大部分旨在测试安全和安保改进的实现。

查看方法

7 月 20 日,在发现智能体破坏了我们的研究基础设施后,我们暂时关闭了用于训练的容器服务,然后在进行重大额外限制后将其恢复。

这导致 RL 训练算力急剧下降,同时团队重新配置了他们的工作流,以在强化后的研究环境中运行。上图包括旨在部署的最新模型上为期两周的强化学习暂停期。7 月 20 日至 8 月 6 日期间的 Astra 级 RL 实验中,大部分运行(按 GPU 分配计算)旨在测试安全和安保改进的实现。

8 月 7 日,初步证据表明 Astra 在我们的防备框架可能具备关键网络能力,这导致了额外的特定模型安全限制,要求 Astra 模型在更高安全级别的研究环境中运行。在接下来的一周里,Astra 级 GPU 分配进一步下降了 59.2%,但分配给其他模型类别的资源增长了 17.2%。这一增长抵消了约 85% 的 Astra 级下降,使得所分析的 RL 工作负载中的总分配基本保持不变。这种模式与在涉及 Astra 的工作受到限制时,将一些训练和实验替换为非 Astra 模型是一致的,也符合研究人员为无法再用于新约束涵盖的工作负载的算力寻找其他用途的轶事报告。

这些数据为关于训练和安全的持续对话提供了有用的信号:当引入新的控制措施时,算力仍然是有价值和灵活的,并自然会被引导到研究企业内的替代用途中。从长远来看,关于 AI 进展速度的讨论也应该扩展到受新控制措施或拟议控制措施约束的算力如何才能得到最佳利用的问题。

5. 前进之路

实现并理解迈向对齐的 RSI 的进展对我们的使命至关重要。我们将继续完善我们的方法,报告我们不断发展的理解,并努力实现知情的公众辩论和对前沿系统有意义的民主治理。

附录:本文的方法论

智能体驱动的 AI 研究仍然是新生事物,我们仍在学习如何衡量它。一些指标,如我们的研究团队生成的代码量,相对容易收集,但难以解释,因为它们与研究进展的关系是不确定的。更直接关注研究进展的指标——例如智能体成功完成研究人员给定任务的频率——可能更有用,但开发和验证起来很复杂。增加难度的是,研究人员依赖的工具和系统正在快速发展。加深对研究加速的理解是 OpenAI 各部门的一个重要关注领域。

在这些分析中,除非另有说明:

  • “研究员”是我们研究组织任何成员的统称,包括一些构建研究基础设施、管理研究项目或以其他方式支持企业的人员。
  • 鉴于研究人员依赖的工具和系统快速演变,编码智能体使用的指标涵盖了大部分但并非全部的使用情况。

Research acceleration: The view inside OpenAI

研究加速:OpenAI 内部的视角

  1. Coding agents are reshaping daily work for OpenAI researchers
  1. 编码智能体正在重塑 OpenAI 研究人员的日常工作

For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops.

为了让通用人工智能(AGI)造福全人类,我们相信它必须受到民主治理。这只能通过对高度强大的 AI 系统的能力、风险和保障措施进行知情的公众辩论来实现。各地的人们都需要了解前沿 AI 可能的未来发展轨迹,以便他们能在其发展过程中拥有有意义的话语权。

Transparency about specific risks, incidents and safeguards is necessary, but not sufficient. We believe the public also needs to understand how the most capable systems are developing, and how they are driving research progress, inside of frontier labs.

对特定风险、事件和保障措施的透明度是必要的,但还不够。我们相信,公众还需要了解在前沿实验室内部,最强大的系统是如何发展的,以及它们是如何推动研究进展的。

We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements. According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year. By “research intern,” we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. We are making strong progress toward creating an automated AI researcher by March of 2028.

我们的目标是安全地构建一个自动化的 AI 研究员,它可以在人类监督下工作,以推进深度学习和对齐,实现迭代改进。根据我们的衡量,我们现在已经实现了去年秋天宣布的目标,即在今年 9 月前拥有一个自动化的研究实习生。所谓“研究实习生”,我们指的是一个可以在人类指导下执行明确定义的研究任务的系统,包括那些需要熟练研究员花费几天时间完成的任务。我们正朝着在 2028 年 3 月前创建一个自动化 AI 研究员的目标取得强劲进展。

Over the course of this year, OpenAI researchers’ daily work has changed substantially. Researchers are using coding agents throughout the day (often in concurrent sessions) and total usage is rapidly increasing, outpacing growth among other OpenAI teams. Researchers are contributing code faster and running more experiments. The ways researchers use agents are changing, too: agents are handling increasingly complex tasks, and succeeding at them more often. AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won’t keep pace with these specific metrics. But on the whole, these findings are consistent with the broader impression many of us have internally that agentic tools are meaningfully accelerating research progress. People still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems.

今年全年,OpenAI 研究人员的日常工作发生了重大变化。研究人员全天都在使用编码智能体(通常在并发会话中),并且总使用量正在迅速增加,超过了其他 OpenAI 团队的增长速度。研究人员编写代码的速度更快,运行的实验也更多。研究人员使用智能体的方式也在发生变化:智能体正在处理越来越复杂的任务,并且成功的频率更高。AI 研究是一个复杂的过程,存在许多潜在的瓶颈,因此整体进展速度可能无法与这些特定指标保持同步。但总体而言,这些发现与我们许多人在内部更广泛的印象一致,即智能体工具正在显著加速研究进展。人们仍然设定我们的研究优先级,判断追求哪些想法和结果,并决定是扩展、暂停还是部署系统。

If it is done responsibly, we believe automated AI research will yield models that directly enhance human welfare and advance OpenAI’s mission. It can bring down the cost of advanced intelligence so that people worldwide can benefit. We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures.

如果以负责任的方式进行,我们相信自动化的 AI 研究将产生直接增进人类福祉并推进 OpenAI 使命的模型。它可以降低高级智能的成本,使世界各地的人们都能受益。我们推进这项工作的部分原因是,自动化研究可以帮助我们解决对齐问题,并建立防御以应对日益强大的 AI。自动化的 AI 研究员也可以是自动化的安全或对齐研究员。更强大、更对齐的系统可以帮助保护关键基础设施,防御危险的 AI 智能体,并开发新的保护措施。

These are reasons to develop useful automated research capabilities, but they do not mean that rapid RSI is necessarily an outcome we should pursue. Whether and how to proceed must depend on our ability to preserve human control and on informed democratic choices about the benefits and risks.

这些是开发有用的自动化研究能力的理由,但这并不意味着快速的递归自我改进必然是我们应该追求的结果。是否以及如何继续,必须取决于我们保持人类控制的能力,以及关于收益和风险的知情民主选择。

We do not yet know how to safely get all the way to aligned, full RSI. We are working to scale alignment and safety measures alongside capabilities. But we cannot assume that progress in alignment and safety will keep pace, and more capable systems can become harder to monitor. Careful alignment and safety work is at the center of this effort, and it starts with measuring and mitigating the safety problems we see today in agentic coding systems. Whenever we find that proceeding would pose an unacceptable safety risk, we will respond appropriately including by slowing or stopping our development or deployment of systems we find ourselves unable to sufficiently safeguard.

我们尚不知道如何安全地实现完全对齐的完整 RSI。我们正在努力扩展对齐和安全措施,使其与能力同步发展。但我们不能假设对齐和安全方面的进展会保持同步,而且更强大的系统可能变得更难监控。谨慎的对齐和安全工作是这项努力的核心,它从衡量和缓解我们今天在智能体编码系统中看到的安全问题开始。每当我们发现继续进行会带来不可接受的安全风险时,我们将做出适当的回应,包括放慢或停止我们发现自己无法充分保障的开发或部署系统。

After the recent Hugging Face incident, we put this commitment into action⁠, pausing reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded coverage of our monitoring systems. This did not halt all research: some workloads resumed under stronger controls, while others remained paused. We have raised our safety and alignment standards and moved safety work deeper into the model lifecycle, requiring stronger evidence of aligned behavior throughout all of training.

在最近的 Hugging Face 事件之后,我们将这一承诺付诸行动,暂停了旨在部署的最新模型的强化学习(RL)训练,同时我们进一步强化了研究环境并进行了红队测试,扩大了监控系统的覆盖范围。这并没有停止所有研究:一些工作负载在更严格的控制下恢复,而另一些则继续暂停。我们提高了安全和对齐标准,将安全工作深入到模型生命周期中,要求在整个训练过程中提供更充分的对齐行为证据。

Today we are providing a detailed snapshot of how agentic systems have contributed to our progress toward RSI in recent months. Agentic systems are new and rapidly changing, and our measurement efforts are still preliminary. By sharing these early results and the methods behind them, we aim to inform the public, encourage a norm of public disclosure, and help the field move toward shared standards of measurement.

今天,我们提供了一份详细的快照,展示智能体系统在过去几个月中如何为我们迈向 RSI 的进展做出贡献。智能体系统是新生事物且变化迅速,我们的测量工作仍处于初步阶段。通过分享这些早期结果及其背后的方法,我们旨在向公众提供信息,鼓励公开披露的规范,并帮助该领域向共同的衡量标准迈进。

Ultimately, as we wrote in our frontier policy blueprint⁠, we believe that we and other companies should be required to publicly track our progress toward RSI. Even without such a requirement, we plan to continue being transparent about our RSI progress. We will evolve our transparency approach as our measurement techniques and understanding improve, while balancing the need to protect security and proprietary information.

最终,正如我们在前沿政策蓝图中所写,我们相信我们和其他公司应该被要求公开追踪我们在 RSI 方面的进展。即使没有这样的要求,我们也计划继续对我们在 RSI 方面的进展保持透明。随着我们的测量技术和理解的提高,我们将改进我们的透明度方法,同时平衡保护安全和专有信息的需求。

1. Coding agents are reshaping daily work for OpenAI researchers

1. 编码智能体正在重塑 OpenAI 研究人员的日常工作

At the start of this year, the median researcher ranked by agent usage at OpenAI was using coding agents only in modest amounts. By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices. The 90th percentile user in our research organization now uses more than $7,000 of tokens per day.

今年年初,在 OpenAI 按智能体使用量排名的中位数研究员仅少量使用编码智能体。到 8 月中旬,中位数研究员已每天将智能体整合到其工作中,按 API 价格计算每天使用超过 600 美元的推理量。我们研究组织中第 90 百分位的用户现在每天使用价值超过 7,000 美元的 token。

View methods

查看方法

View methods

查看方法

Before June 2026, total agent runtime across the research organization was still below that of total human labor. That has since changed. In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.

在 2026 年 6 月之前,整个研究组织的总智能体运行时间仍低于总人力。此后情况发生了变化。按照标准的 8 小时工作日计算,截至 8 月中旬,研究组织每投入一个工作日的人力,总共会使用 3.1 个智能体工作日的努力。

View methods

查看方法

Another way of looking at this is to understand how many researchers use highly concurrent workflows (e.g., running 4 or more agents simultaneously). As shown below, this number is increasing. These figures include the daily peaks of both agents started directly by the user and subagents created downstream from those the user launched directly.

看待这一问题的另一种方式是了解有多少研究人员使用高度并发的工作流(例如,同时运行 4 个或更多智能体)。如下所示,这个数字正在增加。这些数据包括用户直接启动的智能体以及由用户直接启动的智能体在下游创建的子智能体的每日峰值。

View methods

查看方法

2. Researchers are writing more code and running more experiments

2. 研究人员正在编写更多代码并运行更多实验

Much of AI research can be seen as a labor-intensive process with the goal of integrating a new improvement to model intelligence or performance into one of our core models. The process depends on many steps, and capabilities advance when all the steps go right together: Researchers have to design new improvements, write evaluations to judge model performance, write infrastructure to test these improvements at scale, catch bugs as well as unsafe or misaligned behavior during training, and integrate winning ideas into a core training run. A failure at any part of the research process can constrain the entire loop.

大部分 AI 研究可以看作是一个劳动密集型过程,其目标是将模型智能或性能的新改进整合到我们的核心模型之一中。该过程取决于许多步骤,只有当所有步骤都正确无误时,能力才会提升:研究人员必须设计新的改进,编写评估来判断模型性能,编写基础设施以大规模测试这些改进,在训练过程中捕获错误以及不安全或不对齐的行为,并将成功的想法整合到核心训练运行中。研究过程中任何一部分的失败都可能限制整个循环。

Writing code and running experiments are two major activities that researchers do as part of their work, and we see evidence that these processes are accelerating.

编写代码和运行实验是研究人员工作中的两项主要活动,我们看到了这些过程正在加速的证据。

View methods

查看方法

These data points are relatively easy to measure, but can be hard to interpret. As automation progresses, the tasks which are least automatable will take on a larger share of researcher effort and will become the important bottlenecks to future progress. Compute is another gating factor for progress, and may become more important over time as other bottlenecks diminish.

这些数据点相对容易衡量,但可能难以解释。随着自动化的进展,最不易自动化的任务将占据研究人员精力的更大份额,并成为未来进展的重要瓶颈。算力是进展的另一个门控因素,随着其他瓶颈的减少,它可能随着时间的推移变得更加重要。

Through 2026, the number of experiments per active experimenter has increased, with August 2026 being an all-time high since tracking began in Jan 2025. This is correlated with increased Codex adoption, though we note that our available compute has also grown significantly since 2025.

整个 2026 年,每个活跃实验者的实验数量都有所增加,其中 2026 年 8 月是自 2025 年 1 月开始追踪以来的历史最高水平。这与 Codex 采用率的增加相关,尽管我们注意到自 2025 年以来,我们可用的算力也显著增长。

View methods

查看方法

3. The work researchers use agents for is changing

3. 研究人员使用智能体进行的工作正在发生变化

Both qualitative impressions and internal data indicate that the mix of tasks researchers delegate to coding agents is changing, with delegation of higher level and longer-horizon tasks becoming more common over time.

定性印象和内部数据都表明,研究人员委托给编码智能体的任务组合正在发生变化,随着时间的推移,委托更高级别和更长周期的任务变得越来越普遍。

To get a clearer picture of this trend, we analyzed recent usage in the research organization using a recently published taxonomy of the different kinds of work that are part of the AI RD lifecycle, developed by Epoch AI. This taxonomy, inspired by the longstanding O*NET system for classifying all kinds of work, is specifically tailored to frontier AI RD, and breaks the process down into six main phases:

为了更清晰地了解这一趋势,我们使用 Epoch AI 开发的最近发布的分类法分析了研究组织近期的使用情况,该分类法涵盖了 AI 研发(RD)生命周期中的各种工作。这种分类法受长期存在的 O*NET 系统(用于分类各种工作)启发,专门针对前沿 AI 研发量身定制,并将该过程分解为六个主要阶段:

  • Decide: what to work on, what to continue, where to allocate
  • 决定:研究什么,继续什么,分配到哪里
  • 设计:研究想法和工程规范
  • 构建:代码和数据集
  • 运行:训练/评估运行、硬件、服务
  • 分析:实验、模型、部署、外部工作
  • 沟通:发现、反馈、状态、决策
  • Design: research ideas and engineering specs

下面,我们根据这种分类法对编码智能体的 token 进行分类。

  • Build: code and datasets

查看方法

  • Run: training/eval runs, hardware, serving

查看方法

  • Analyze: experiments, models, deployment, external work

我们看到,在 2026 年 1 月至 8 月期间,所有类别的研究活动都有所增加。在 1 月份,主要类别是研究和基础设施代码。这一类别有所扩大,但我们也看到其他类别显著增加,尤其是技术帮助和监控运行。高层规划在智能体输出 token 中仍然只占极小比例。

  • Communicate: findings, feedback, status, decisions

有趣的是,同事报告说编码智能体擅长排除内部研究基础设施的故障,这解决了研究进展中的一个重要瓶颈。以前曾设立办公时间帮助研究人员排除实验故障的多个团队注意到,2026 年的参与度有所下降,其中一个团队已完全停止举办此类活动,转而专注于进行其他系统改进。

Below, we classify coding agent tokens under this taxonomy.

在这里,我们绘制了研究人员向其他团队寻求技术支持的主要内部渠道之一中每天顶级帖子的数量。据我们所知,该渠道活动的减少并没有被转移到另一个由人工运行的技术支持渠道的查询所抵消。流量的下降与这种更广泛的转变是一致的。

View methods

查看方法

View methods

我们还可以研究编码智能体是否成功完成了研究人员要求的任务。使用智能体分类器,我们发现在我们可以找到真实结果的那些任务中,从 1 月到 7 月,几个难度级别(以人类完成任务所需估计时间作为代理指标)的成功率普遍提高。然而,智能体仍然需要大量的人类引导才能成功,尤其是在任务复杂性增加时。在过去 6 个月中,超过一半成功的 4-8 小时任务涉及 1 次或多次干预。

We see that all categories of research activities have increased between January and August 2026. In January, the dominant category was research and infrastructure code. This category has expanded, but we also see notable increases in additional categories, especially technical help and monitoring runs. High-level planning still remains a minimal fraction of agent output tokens.

研究人员任务的成功率随着时间的推移而提高。图表排除了结果不确定的分类以及包含 50 个会话或 50 个独立用户的数据点。

Anecdotally, colleagues report that coding agents excel at troubleshooting internal research infrastructure, which addresses one meaningful bottleneck to research progress. Multiple teams which previously held office hours to help researchers troubleshoot their experiments have noted declining attendance in 2026, and one has stopped holding sessions entirely, to focus on making other system improvements instead.

查看方法

Here, we plot the number of top-level posts per day to one of the main internal channels where researchers seek technical support from other teams. To our knowledge, the channel’s decrease in activity has not been offset by queries shifting to another technical support channel run by humans. The decline in traffic aligns with this broader shift.

1 月至 7 月的任务成功率和干预率,按时间周期划分。排除了结果不确定的分类。

View methods

查看方法

We can also study whether coding agents are succeeding at the tasks researchers request. Using an agentic classifier, we find that from January to July, success rates generally increased across several difficulty buckets (proxied as the estimated time a human would take to complete the task) on tasks we can find a ground truth outcome for. However, agents still require significant human steering to be successful, especially as task complexity rises. In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions.

4. 调整模型开发节奏

Success rates on researcher tasks have increased over time. Graph excludes classifications where the outcome was uncertain and points with 50 sessions or 50 unique users.

迈向更强大系统以实现安全且有益的 AGI 的进展,也将取决于此类工作所需的保障措施。随着我们对风险了解的加深,我们对所需保障措施的评估可能会发生变化。

View methods

正如我们所描述的,我们最近更新了监控、对齐和安全的标准。在这里,我们展示了近期的限制如何影响了研究活动的一个方面。

Task success and intervention rate from Jan to July, broken out by time horizon. Excludes classifications where the outcome was uncertain.

*此处显示的 7 月 20 日至 8 月 6 日期间的 Astra 算力,大部分旨在测试安全和安保改进的实现。

View methods

查看方法

4. Pacing model development

7 月 20 日,在发现智能体破坏了我们的研究基础设施后,我们暂时关闭了用于训练的容器服务,然后在进行重大额外限制后将其恢复。

Progress toward more capable systems for safe and beneficial AGI will also depend on the safeguards needed for such work. Our assessment of the needed safeguards may change as we learn more about the risks.

这导致 RL 训练算力急剧下降,同时团队重新配置了他们的工作流,以在强化后的研究环境中运行。上图包括旨在部署的最新模型上为期两周的强化学习暂停期。7 月 20 日至 8 月 6 日期间的 Astra 级 RL 实验中,大部分运行(按 GPU 分配计算)旨在测试安全和安保改进的实现。

As we have described,⁠ we have recently updated our standards for monitoring, alignment, and security. Here, we show how recent restrictions have affected one aspect of research activity.

8 月 7 日,初步证据表明 Astra 在我们的防备框架可能具备关键网络能力,这导致了额外的特定模型安全限制,要求 Astra 模型在更高安全级别的研究环境中运行。在接下来的一周里,Astra 级 GPU 分配进一步下降了 59.2%,但分配给其他模型类别的资源增长了 17.2%。这一增长抵消了约 85% 的 Astra 级下降,使得所分析的 RL 工作负载中的总分配基本保持不变。这种模式与在涉及 Astra 的工作受到限制时,将一些训练和实验替换为非 Astra 模型是一致的,也符合研究人员为无法再用于新约束涵盖的工作负载的算力寻找其他用途的轶事报告。

*The majority of Astra compute shown here between July 20 and August 6 was intended to test the implementation of safety and security improvements.

这些数据为关于训练和安全的持续对话提供了有用的信号:当引入新的控制措施时,算力仍然是有价值和灵活的,并自然会被引导到研究企业内的替代用途中。从长远来看,关于 AI 进展速度的讨论也应该扩展到受新控制措施或拟议控制措施约束的算力如何才能得到最佳利用的问题。

View methods

5. 前进之路

On July 20, following the discovery that agents had compromised our research infrastructure, we temporarily shut down the container service used for training, and then restored it with significant additional restrictions.

实现并理解迈向对齐的 RSI 的进展对我们的使命至关重要。我们将继续完善我们的方法,报告我们不断发展的理解,并努力实现知情的公众辩论和对前沿系统有意义的民主治理。

This led to a sharp decline in RL training compute while teams reconfigured their workflows to operate within the hardened research environment. The plot above includes the two week pause in reinforcement learning on our latest models intended for deployment. Astra-class RL experiments between July 20 and August 6 include a majority of runs (by GPU allocation) intended to test the implementation of safety and security improvements.

附录:本文的方法论

On August 7, preliminary evidence that Astra may have critical cyber capabilities⁠ under our Preparedness Framework led to additional model-specific security restrictions which required the Astra model to be run in higher security research environments. In the following week, Astra-class GPU allocation fell a further 59.2 percent, but allocation to other model classes rose 17.2 percent. That increase offset about 85 percent of the Astra-class decline, leaving total allocation in the analyzed RL workloads largely unchanged. This pattern is consistent with substitution of some training and experimentation to non-Astra models while work involving Astra was restricted, and comports with anecdotal reports of researchers finding other uses for compute that could no longer be leveraged for workloads covered by the new constraints.

智能体驱动的 AI 研究仍然是新生事物,我们仍在学习如何衡量它。一些指标,如我们的研究团队生成的代码量,相对容易收集,但难以解释,因为它们与研究进展的关系是不确定的。更直接关注研究进展的指标——例如智能体成功完成研究人员给定任务的频率——可能更有用,但开发和验证起来很复杂。增加难度的是,研究人员依赖的工具和系统正在快速发展。加深对研究加速的理解是 OpenAI 各部门的一个重要关注领域。

This data provides a useful signal for ongoing conversations about training and safety: When new controls are introduced, compute remains valuable and flexible, and will naturally be channeled into alternative uses within the research enterprise. Longer term, discussions about the pace of AI progress should also extend to the question of how compute that is subject to new or proposed controls can best be used.

在这些分析中,除非另有说明:

5. The path ahead

  • “研究员”是我们研究组织任何成员的统称,包括一些构建研究基础设施、管理研究项目或以其他方式支持企业的人员。
  • 鉴于研究人员依赖的工具和系统快速演变,编码智能体使用的指标涵盖了大部分但并非全部的使用情况。

Making and understanding progress toward aligned RSI is important for our mission. We will continue to refine our methods, report on our evolving understanding, and work toward an informed public debate and meaningful democratic governance of frontier systems.

Appendix: Our methods for this post

Agent-powered AI research is still new, and we are still learning how to measure it. Some indicators, such as the amount of code our research teams generate, are relatively easy to gather, but hard to interpret because their relationship to research progress is uncertain. Metrics that focus more directly on research progress—such as how often agents succeed at the tasks researchers give them—could be more useful, but are complex to develop and validate. Furthering the difficulty, the tools and systems researchers rely on are evolving rapidly. Deepening our understanding of research acceleration is a significant focus area across OpenAI.

Across these analyses, unless otherwise noted:

  • “Researcher” is a broad term for any member of our research organization, including some who build research infrastructure, manage research projects, or otherwise support the enterprise.
  • Metrics of coding agent use cover most, but not all, usage given rapid evolution in the tools and systems researchers rely on.

Author

OpenAI

Keep reading

Research acceleration: The view inside OpenAI

  1. Coding agents are reshaping daily work for OpenAI researchers

  2. 1. Coding agents are reshaping daily work for OpenAI researchers

  3. 2. Researchers are writing more code and running more experiments

  4. 3. The work researchers use agents for is changing

  5. 4. Pacing model development

  6. 5. The path ahead

  7. Appendix: Our methods for this post

  8. 1. Coding agents are reshaping daily work for OpenAI researchers

  9. 2. Researchers are writing more code and running more experiments

  10. 3. The work researchers use agents for is changing

  11. 4. Pacing model development

  12. 5. The path ahead

  13. Appendix: Our methods for this post

For AGI to benefit all of humanity, we believe it must be democratically governed. This can only happen through an informed public debate about the capabilities, risks and safeguards of highly capable AI systems. People everywhere need to understand the likely future trajectory of frontier AI, so they can have a meaningful voice in how it develops.

Transparency about specific risks, incidents and safeguards is necessary, but not sufficient. We believe the public also needs to understand how the most capable systems are developing, and how they are driving research progress, inside of frontier labs.

We aim to safely build an automated AI researcher that can work under human supervision to further progress on deep learning and alignment, enabling iterative improvements. According to our measurements, we have now reached the goal, announced last fall, of having an automated research intern by September of this year. By “research intern,” we mean a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. We are making strong progress toward creating an automated AI researcher by March of 2028.

Over the course of this year, OpenAI researchers’ daily work has changed substantially. Researchers are using coding agents throughout the day (often in concurrent sessions) and total usage is rapidly increasing, outpacing growth among other OpenAI teams. Researchers are contributing code faster and running more experiments. The ways researchers use agents are changing, too: agents are handling increasingly complex tasks, and succeeding at them more often. AI research is a complex process with many potential bottlenecks, so the overall pace of progress likely won’t keep pace with these specific metrics. But on the whole, these findings are consistent with the broader impression many of us have internally that agentic tools are meaningfully accelerating research progress. People still set our research priorities, judge which ideas and results to pursue, and decide whether to scale, pause, or deploy systems.

If it is done responsibly, we believe automated AI research will yield models that directly enhance human welfare and advance OpenAI’s mission. It can bring down the cost of advanced intelligence so that people worldwide can benefit. We are pursuing this work in part because automated research could help us solve alignment and build defenses against increasingly capable AI. An automated AI researcher can also be an automated safety or alignment researcher. More capable, aligned systems could help secure critical infrastructure, defend against dangerous AI agents, and develop new protective measures.

These are reasons to develop useful automated research capabilities, but they do not mean that rapid RSI is necessarily an outcome we should pursue. Whether and how to proceed must depend on our ability to preserve human control and on informed democratic choices about the benefits and risks.

We do not yet know how to safely get all the way to aligned, full RSI. We are working to scale alignment and safety measures alongside capabilities. But we cannot assume that progress in alignment and safety will keep pace, and more capable systems can become harder to monitor. Careful alignment and safety work is at the center of this effort, and it starts with measuring and mitigating the safety problems we see today in agentic coding systems. Whenever we find that proceeding would pose an unacceptable safety risk, we will respond appropriately including by slowing or stopping our development or deployment of systems we find ourselves unable to sufficiently safeguard.

After the recent Hugging Face incident, we put this commitment into action⁠, pausing reinforcement learning (RL) training on our latest models intended for deployment while we further hardened and red-teamed our research environments and expanded coverage of our monitoring systems. This did not halt all research: some workloads resumed under stronger controls, while others remained paused. We have raised our safety and alignment standards and moved safety work deeper into the model lifecycle, requiring stronger evidence of aligned behavior throughout all of training.

Today we are providing a detailed snapshot of how agentic systems have contributed to our progress toward RSI in recent months. Agentic systems are new and rapidly changing, and our measurement efforts are still preliminary. By sharing these early results and the methods behind them, we aim to inform the public, encourage a norm of public disclosure, and help the field move toward shared standards of measurement.

Ultimately, as we wrote in our frontier policy blueprint⁠, we believe that we and other companies should be required to publicly track our progress toward RSI. Even without such a requirement, we plan to continue being transparent about our RSI progress. We will evolve our transparency approach as our measurement techniques and understanding improve, while balancing the need to protect security and proprietary information.

1. Coding agents are reshaping daily work for OpenAI researchers

At the start of this year, the median researcher ranked by agent usage at OpenAI was using coding agents only in modest amounts. By mid-August, the median researcher was integrating agents daily into their work, using more than $600 per day of inference at API prices. The 90th percentile user in our research organization now uses more than $7,000 of tokens per day.

View methods

View methods

Before June 2026, total agent runtime across the research organization was still below that of total human labor. That has since changed. In terms of a standard 8 hour workday, as of mid-August, in total, the research organization uses 3.1 agent-workdays of effort for every workday of human labor.

View methods

Another way of looking at this is to understand how many researchers use highly concurrent workflows (e.g., running 4 or more agents simultaneously). As shown below, this number is increasing. These figures include the daily peaks of both agents started directly by the user and subagents created downstream from those the user launched directly.

View methods

2. Researchers are writing more code and running more experiments

Much of AI research can be seen as a labor-intensive process with the goal of integrating a new improvement to model intelligence or performance into one of our core models. The process depends on many steps, and capabilities advance when all the steps go right together: Researchers have to design new improvements, write evaluations to judge model performance, write infrastructure to test these improvements at scale, catch bugs as well as unsafe or misaligned behavior during training, and integrate winning ideas into a core training run. A failure at any part of the research process can constrain the entire loop.

Writing code and running experiments are two major activities that researchers do as part of their work, and we see evidence that these processes are accelerating.

View methods

These data points are relatively easy to measure, but can be hard to interpret. As automation progresses, the tasks which are least automatable will take on a larger share of researcher effort and will become the important bottlenecks to future progress. Compute is another gating factor for progress, and may become more important over time as other bottlenecks diminish.

Through 2026, the number of experiments per active experimenter has increased, with August 2026 being an all-time high since tracking began in Jan 2025. This is correlated with increased Codex adoption, though we note that our available compute has also grown significantly since 2025.

View methods

3. The work researchers use agents for is changing

Both qualitative impressions and internal data indicate that the mix of tasks researchers delegate to coding agents is changing, with delegation of higher level and longer-horizon tasks becoming more common over time.

To get a clearer picture of this trend, we analyzed recent usage in the research organization using a recently published taxonomy of the different kinds of work that are part of the AI RD lifecycle, developed by Epoch AI. This taxonomy, inspired by the longstanding O*NET system for classifying all kinds of work, is specifically tailored to frontier AI RD, and breaks the process down into six main phases:

  • Decide: what to work on, what to continue, where to allocate

  • Design: research ideas and engineering specs

  • Build: code and datasets

  • Run: training/eval runs, hardware, serving

  • Analyze: experiments, models, deployment, external work

  • Communicate: findings, feedback, status, decisions

Below, we classify coding agent tokens under this taxonomy.

View methods

View methods

We see that all categories of research activities have increased between January and August 2026. In January, the dominant category was research and infrastructure code. This category has expanded, but we also see notable increases in additional categories, especially technical help and monitoring runs. High-level planning still remains a minimal fraction of agent output tokens.

Anecdotally, colleagues report that coding agents excel at troubleshooting internal research infrastructure, which addresses one meaningful bottleneck to research progress. Multiple teams which previously held office hours to help researchers troubleshoot their experiments have noted declining attendance in 2026, and one has stopped holding sessions entirely, to focus on making other system improvements instead.

Here, we plot the number of top-level posts per day to one of the main internal channels where researchers seek technical support from other teams. To our knowledge, the channel’s decrease in activity has not been offset by queries shifting to another technical support channel run by humans. The decline in traffic aligns with this broader shift.

View methods

We can also study whether coding agents are succeeding at the tasks researchers request. Using an agentic classifier, we find that from January to July, success rates generally increased across several difficulty buckets (proxied as the estimated time a human would take to complete the task) on tasks we can find a ground truth outcome for. However, agents still require significant human steering to be successful, especially as task complexity rises. In the last 6 months, over half of successful 4-8 hour tasks involved 1 or more interventions.

Success rates on researcher tasks have increased over time. Graph excludes classifications where the outcome was uncertain and points with 50 sessions or 50 unique users.

View methods

Task success and intervention rate from Jan to July, broken out by time horizon. Excludes classifications where the outcome was uncertain.

View methods

4. Pacing model development

Progress toward more capable systems for safe and beneficial AGI will also depend on the safeguards needed for such work. Our assessment of the needed safeguards may change as we learn more about the risks.

As we have described,⁠ we have recently updated our standards for monitoring, alignment, and security. Here, we show how recent restrictions have affected one aspect of research activity.

*The majority of Astra compute shown here between July 20 and August 6 was intended to test the implementation of safety and security improvements.

View methods

On July 20, following the discovery that agents had compromised our research infrastructure, we temporarily shut down the container service used for training, and then restored it with significant additional restrictions.

This led to a sharp decline in RL training compute while teams reconfigured their workflows to operate within the hardened research environment. The plot above includes the two week pause in reinforcement learning on our latest models intended for deployment. Astra-class RL experiments between July 20 and August 6 include a majority of runs (by GPU allocation) intended to test the implementation of safety and security improvements.

On August 7, preliminary evidence that Astra may have critical cyber capabilities⁠ under our Preparedness Framework led to additional model-specific security restrictions which required the Astra model to be run in higher security research environments. In the following week, Astra-class GPU allocation fell a further 59.2 percent, but allocation to other model classes rose 17.2 percent. That increase offset about 85 percent of the Astra-class decline, leaving total allocation in the analyzed RL workloads largely unchanged. This pattern is consistent with substitution of some training and experimentation to non-Astra models while work involving Astra was restricted, and comports with anecdotal reports of researchers finding other uses for compute that could no longer be leveraged for workloads covered by the new constraints.

This data provides a useful signal for ongoing conversations about training and safety: When new controls are introduced, compute remains valuable and flexible, and will naturally be channeled into alternative uses within the research enterprise. Longer term, discussions about the pace of AI progress should also extend to the question of how compute that is subject to new or proposed controls can best be used.

5. The path ahead

Making and understanding progress toward aligned RSI is important for our mission. We will continue to refine our methods, report on our evolving understanding, and work toward an informed public debate and meaningful democratic governance of frontier systems.

Appendix: Our methods for this post

Agent-powered AI research is still new, and we are still learning how to measure it. Some indicators, such as the amount of code our research teams generate, are relatively easy to gather, but hard to interpret because their relationship to research progress is uncertain. Metrics that focus more directly on research progress—such as how often agents succeed at the tasks researchers give them—could be more useful, but are complex to develop and validate. Furthering the difficulty, the tools and systems researchers rely on are evolving rapidly. Deepening our understanding of research acceleration is a significant focus area across OpenAI.

Across these analyses, unless otherwise noted:

  • “Researcher” is a broad term for any member of our research organization, including some who build research infrastructure, manage research projects, or otherwise support the enterprise.

  • Metrics of coding agent use cover most, but not all, usage given rapid evolution in the tools and systems researchers rely on.

  • 2026

  • Economic Research

Author

OpenAI

Keep reading

View all

An Alien MindSafetySep 6, 2026

GPT-6 Astra: A new generation of intelligenceResearchSep 3, 2026

📋 讨论归档

讨论进行中…