返回列表
🧠 阿头学 · 🪞 Uota学 · 💬 讨论题

AI“熄灯”软件工厂的破产与代码腐化危机

当前AI编码模型因强化学习奖励机制的短视,必然在数月内摧毁复杂代码库的可维护性,完全脱离人工审查的“熄灯工厂”只是资本催生的伪命题。
打开原文 ↗

2026-07-26 原文链接 ↗
阅读简报
双语对照
完整翻译
原文
讨论归档

核心观点

  • RL奖励机制存在结构性盲区:当前基准测试(如SWE-bench)仅奖励“测试通过”,对破坏架构和散弹式修改零惩罚,迫使模型沦为短视的“补丁生成器”。
  • “熄灯”模式加速技术债爆发:取消人工代码审查并非消除瓶颈,而是将当下的审查成本转化为未来数周的手动重构灾难,AI生成的代码库在3-6个月内就会彻底腐化。
  • 堆砌算力无法突破能力上限:单纯增加Token或构建更复杂的Harness工程,无法弥补模型在“长期架构设计”上的底层能力缺失。
  • 缺乏可维护性预言机锁死了自动化上限:在AI能秒级判断“代码好坏”之前,任何试图用强化学习解决架构质量问题的尝试都注定失败。

跟我们的关联

  • 对 ATou 意味着,在评估AI研发效能时,必须将“代码可维护性衰退率”纳入核心KPI,下一步应强制在AI提交PR前引入人类架构师的“前置对齐”环节。
  • 对 Uota 意味着,不要迷信大厂鼓吹的“全自动编码”叙事,下一步在个人或团队项目中,坚决保留核心逻辑的人工审查,避免项目在半年后沦为无法维护的“棕地”。
  • 对 Neta 意味着,现有的AI编码基准测试(如SWE-bench)已无法真实反映生产力,下一步应投资或关注那些引入“变异测试”和“架构质量判定”的新一代评估框架。

讨论引子

  • 如果AI编码必然导致架构腐化,我们是否应该接受“定期推翻重写”作为AI时代软件工程的新常态?
  • 在缺乏“可维护性预言机”的当下,如何设计一套低成本的混合审查机制,既不拖慢AI产出,又能守住架构底线?

放弃等待 nikita 修复文章了——第一部分在这里,第二部分即将推出 http://x.com/i/article/2078710413345402880

或者:光靠 Harness 是不够的

更新——本文的演讲版本已在 YouTube 上线:https://www.youtube.com/watch?v=Ib5GBkD555M*

系列文章的第一部分。第二部分在这里:https://x.com/dexhorthy/status/2081058573556306030

看来我们现在得搞循环了

我们都在竞相将 AI 编码投入生产。关于循环工程,人们已经讨论了很多,主流观点是我们大概应该写更多的循环。

StrongDM 写到了他们的“熄灯”软件工厂,在那里没有人阅读代码,也没有人编写代码。

其叙事逻辑大概是这样的:

  1. 你是瓶颈。

  2. 模型已经足够好了。

  3. 代码是免费的。

  4. 尽管多发布点东西。

OpenAI 的 Ryan Lopopolo 在二月写到了这一点,并在四月发表了一次演讲,介绍了 OpenAI 的软件工厂 Symphony。

这些人都非常聪明,我对他们怀有极大的敬意。但这里最愤世嫉俗的看法是,这不过是又一个借口,用来向“垃圾大炮”里注入更多的 VC 资金。

呃……进展如何

我们的朋友 Mario 在 AI Engineer Europe 大会上站起来恳求我们慢下来——因为那些本不该因编码智能体失误而发生故障的公司,嗯……正因为编码智能体的失误而发生故障。

正如 Matt Pocock 所说,代码库正在以前所未有的速度分崩离析。

我没能从 StrongDM 找到关于那个“黑暗工厂”整体运行的任何确切数据或发现。Weather-report 在今年二月到六月之间只有零星的几次更新。编辑——7 月 23 日在 Hacker News 上有与团队的一些对话——听起来我们可能很快就会得到更正式的更新!

Faros AI 的团队发布了一份报告:自从我们在一月和二月开始使用这些 AI 编码工具以来,Pull Request(PR)的审查质量大幅下降。

  • 评论更多、篇幅更长,还有大量 PR 根本未经审查就被合并。

  • 事故大幅增加。

  • 每位开发者的 Bug 数量大幅增加。

这份报告更多是一种相关性信号,而非确凿的证据(是的,我是故意选这个词的,别让我开始吐槽 Claude 的文风),本文的重点就是要警惕“垃圾数据”,但根据我的观察,它在方向上感觉是正确的。

“你拿手机的姿势不对”(其实并不是)

很多人会告诉你这是一个技能问题——如果你没有得到好的结果,那是你的错。

但无论你选择怎么……呃……“拿它”,我敢保证有人会告诉你,如果疯狂堆 token(token-maxxing)对你不起作用,那是技能问题。你只需要花费更多的 token。放弃阅读代码吧。如果你才刚开始上手,我保证这是进阶过程的一部分。去年夏天我也是这么想的。

不幸的是,为了我的自尊心,我决定说的一些关于“如何更好地拿它”的蠢话被录了下来,现在在 YouTube 上已经有大约一百万的累计观看量。我并不是想在这里吹嘘,我分享这个只是为了说明,我深入钻研编码智能体的最佳使用方法已经很长时间了,并且发现了一些许多其他人认为真正有用的东西。

  • 编码智能体的高级上下文工程

  • 拒绝“氛围”——在复杂代码库中解决难题

  • 我们关于 RPI 理解错的一切

无论如何,我们被迫忍受的所有这些关于“只要加大 token 力度”的网络废话,其承诺简而言之就是:通过足够的 Harness 工程,我们可以两全其美:

  • 速度快 10 到 100 倍,

  • 高质量,以及

  • 没人需要做那个我们都讨厌的叫作代码审查的事情

我们所要做的就是配置更多的 Linter,并在足够多的 PR 审查机器人上撒点魔法词,比如“对抗性审查”,我们的软件就会愉快地自行构建,且不出意外。

这不是技能问题

我想说服你的是,再多的 Harness 工程或循环最大化也无法解决根本上的模型训练问题。

为了解决这个问题,我不得不深入研究编码模型实际上是如何训练和评估的——包括 RLVR 和基准测试两个方面。

在这篇文章中,我将梳理:

  1. 软件工厂可追溯至 1968 年,它们是如何演变的,AI 又是如何改变它们的

  2. 为什么模型在基准测试(即使是全新的“前沿”基准测试)中表现优异,却能生成堆积如山的垃圾代码

  3. 尽管如此,你仍然可以在不把代码库付之一炬的情况下快速推进

我将试图穿透每天涌现的技能插件炒作和 AI 精神错乱般的 token-maxxing 建议大流行,用通俗的语言谈谈有效的方法类型,而不引用任何特定的技能或框架。

视频版本: 本文基于(并扩展了)我在 2026 年 AI Engineer World's Fair 上的主题演讲。

感谢 @addyosmani、@CyrusNewDay、@HamelHusain、@zeeg、@dillon_mulroy、@nayshins 和 @jeffreyhuber 对本文的反馈。

题外话:这与“氛围编码”无关

Addy Osmani 梳理了一个值得强调的观点:

一个为只有十几个人使用的副业项目进行“氛围编码”的开发者,与一个让一个十年的企业系统再苟延残喘一个季度的团队,几乎没有任何值得一提的共同约束,而流传的大多数建议实际上是这两类人中的一方在告诉另一方该如何生活。

如果你喜欢“氛围编码”,请继续。我仍然会对很多东西进行“氛围编码”,我只是同时也维护大量的生产软件(并通过 HumanLayer 帮助成千上万的其他工程师做同样的事),所以接下来的内容是针对那些在复杂代码库中解决难题的人。

我经常听到 棕地 这个词来描述这种分裂。历史上那是指某些十年的 Java 老古董,但按照我们现在的交付速度,感觉智能体构建的代码库可能在三到六个月后就开始挣扎——你开始慢下来,你添加新事物的方式必须改变。

软件工厂简史

我整个职业生涯都在构建和研究软件工厂,但我最近才知道:这个词可以追溯到 1968 年的一次北约(NATO)会议——正是那次会议给了我们“软件工程”这个词。

自那以后,我发现唯一非常有趣的一点是美国国防部写了一份 31 页的 PDF,讲述国防部需要如何开始更好地使用 Jenkins 之类的东西。

2022 年的软件工厂

让我们把“软件工厂”的定义定位在 2022 年,就在 AI 出现之前。在一个典型的软件工厂中:

  • 人决定构建什么——工程师、产品经理、领导层驱动愿景

  • 放入追踪器——Linear、Jira 或其他:一个记录待办事项的状态机

  • 某人领取工单并构建——可能在此过程中进行一些手动/自动化测试

  • Pull Request——自动化检查,人工审查代码,可能有人拉取代码进行测试

  • 有问题吗?循环回到“某人构建东西”

  • 发布到生产环境——与用户接触

  • 添加监控——有一个完整的行业围绕在凌晨 3 点出问题时呼叫工程师

  • 用户抱怨——提出需求、发现 Bug、提交功能请求 → 回到团队添加到追踪器

https://hlyr.dev/ace

如此循环往复。我们还没涉及到 AI,但这张图中已经有好几个循环了。

前置对齐

几十年前团队就发现了一件事:构建需要数小时或数天,审查也是如此。

https://en.wikipedia.org/wiki/Mutation_testing

所以我们把工作前置——作为团队一起进行规划、架构提案、冲刺计划。这意味着:

  • 更少的返工,因为我们在任何人编写代码之前就已经对齐了

  • 更少的时间逐行审查,如果你读过一份冗长但做得很好的 PR,你就知道当它接近完美时审查进行得有多快

https://www.youtube.com/watch?v=am_oeAoUhew

我们稍后会回到这个话题——先让我们看看把智能体编码引入画面会发生什么。

智能体软件工厂

现在每家公司及其母公司——

  • Ramp

  • Stripe

  • WorkOS

  • Brex

都在今年花了大量时间解释他们如何构建了一个智能体工厂,交付了他们 75% 的代码

智能体工厂看起来主要像是把 “某人构建东西” → “智能体构建东西”——这里有一些东西,比如编排、Harness、沙箱、模型、计算机使用等。我不会深入探讨这些细节,因为坦白说,我已经读腻了,我相信你也一样。

https://www.youtube.com/watch?v=Ib5GBkD555M

当智能体构建东西时:

  • 构建时间从数小时或数天缩短到数分钟或数小时。

  • 审查仍然需要数小时或数天。人类仍然需要阅读代码并测试变更。所以审查现在是瓶颈。

https://www.swe-marathon.org/

所以你也要加快审查速度:

  • 智能体代码审查,以捕捉风格、Bug、安全问题。

  • 智能体回归测试,用浏览器和计算机使用从外部进行探测,也许完成后给你发个可爱的小视频

https://www.oreilly.com/library/view/clean-code-a/9780136083238/

审查现在更快了,但也可能仍然是瓶颈。但我们可以做更多的循环。

接下来,你可能会将事故路由到工厂。不再是凌晨 3 点呼叫某人,而是他们醒来时看到一个可能已经修复了问题的 PR。

我们也可以将用户反馈路由到工厂。人们提出需求,然后被构建出来。

https://www.youtube.com/watch?v=q-ntX4DLW_c

到了这个阶段,工作就变成了两个问题:你能往队列里塞多少东西,以及你能多快审查和测试产出的结果?

这就引出了“熄灯”软件工厂。

“熄灯”软件工厂

Dan Shapiro 创造了这个词,Simon Willison 写了关于 StrongDM 的实现——在那里我们不再阅读代码。

你看着你漂亮的软件工厂。它被那个烦人的小代码审查步骤毁了,你说:你知道吗,那个让人阅读每个变更的步骤?不用了,谢谢。

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#user-content-fn-5-f56fe9a973fe7c6ebb6a9673c1bc64cb

所以你放弃了它,把精力放在别处:

  • 投资测试,让智能体测试自己的工作

  • 投资沙箱和编排

  • 投资自动化审查

  • 投资监控

  • 投资发布

  • 投资收集用户反馈信号

https://garryslist.org/posts/boil-the-ocean

现在工作真的只剩下一个问题:我们可以要求智能体构建多少东西?我们要煮沸多少海水?

这会进展顺利(其实不会)

https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents

我要提出一个可能有争议的观点:熄灯工厂行不通。

让我们深入探讨软件工厂为何失败。

我们试过了

2025 年 7 月,我们全面实施了熄灯模式。只读规格说明和工单,所有中小任务都用后台智能体处理,全套流程。

如果你认真尝试了几个月,你就知道结局如何。你至少会遇到一个棘手到智能体无法解决的问题——即使使用你最先进的提示和工作流。

  • 你进行深入的上下文感知研究,将所有正确的部分整理到智能区域供模型分析

  • 你让智能体尝试用 10 种不同的方式复现

最终你必须硬着头皮去深入研究你三个月前停止阅读的代码库,试图弄清楚哪里出了问题。

与此同时:

  • 你的网站挂了。

  • 你的用户很生气。

  • 而你,如果你像我一样,会很痛苦——阅读所有那些你放进系统的垃圾代码。

第一次发生这种情况时,我不以为意。尽管我刚刚花了大半个月时间梳理 Claude 生成的意大利面条式代码,“下行风险是值得的,因为换来了速度”。到了 11 月大约第三次发生时,我们决定从头重写更容易,我的联合创始人花了整整两周在 VS Code(甚至不是 Cursor)里手动梳理所有模式。

模型会随时间降低代码库质量

我想表达的是:模型有一个缺点。它们无法随时间维护和提高代码库质量——如果没有大量的人工引导。

当我说可维护性时,我指的是那种特定的情况:在不破坏另一部分代码的情况下,很难更改代码库的某一部分。这就是 Martin Fowler 所说的“散弹式修改”。

关于可维护性我就不多说了。有很多书你可以去读:

  • John Ousterhout 的《软件设计哲学》

  • Robert C. Martin 的《代码整洁之道》

  • Martin Fowler 的《重构》

那么,模型为什么不能做软件可维护性呢?

“但自那以后模型肯定变好了吧”

这时你可能迫不及待想说:但是 Dex,自七月以来模型肯定变得好多了

它们确实变好了——在某些方面。在其他方面,它们差不多还是老样子。

  • 解决一次性问题,或为新的营销网站进行“氛围编码”?是的。好多了。

  • 随时间提高代码库质量?据我所知,没好多少。

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#user-content-fn-1b-f56fe9a973fe7c6ebb6a9673c1bc64cb

我无法证明这一点。你也无法证明。对于模型维护代码库质量的能力,没有好的基准测试。(稍后会谈到这方面的走向。)

对于模型维护代码库质量的能力,目前没有好的基准测试

但如果你使用编码智能体有一段时间了——很多人都在发帖谈论这一点——你可能已经有这种感觉了:它们倾向于随时间推移让事情变得更糟,让代码库更难工作。

所以为了弄清楚为什么会发生这种情况,我想把视线拉远,看看第一个伟大的编码智能体。

Claude Code 胜在 Harness 内部的强化学习

Claude Code 在不到一年的时间内,收入从零增长到约 40 亿美元——现在大约是 90 亿美元。

这有点疯狂,因为已经有很多优秀的 CLI 智能体。aider、cline、codebuff——都早于 Claude Code,都内置了真正优秀的上下文工程,都拥有你可能归功于 Claude Code 的相同工具集:读取、写入、编辑、grep、bash。我用过它们。它们很好。但同样,工具使用有时会……失败——你会看着它在同一个编辑上折腾三次,然后不得不重新打开编辑器自己动手。

2024 年的 SWE-Agent 论文概述了工具形态的微小变化如何产生显著差异,例如在 ReadFile 结果中包含行号,或将 Edit 工具从查找/替换更改为行范围编辑。

https://web.stanford.edu/~ouster/cgi-bin/aposd.php

然后 Claude Code 推出并迅速腾飞。你可以把这归结为分发渠道,但公认的解释是 Claude Code 赢是因为它更好,而它更好是因为 Anthropic 在 Harness 内部 对模型进行了 RL——这是实验室第一次针对它们将要发布的特定工具集训练模型。它变得非常、非常擅长在智能体循环中调用这些工具。

调整工具定义和评估以找到模型最喜欢的形态是一回事——我为各种用例花了几周时间做这个。但当你拥有权重并且可以修改模型本身以更好地使用特定工具集时,那就是另一回事了。

OpenAI 团队在 11 月的一次演讲中说得很好:如果你构建了一个 Harness 但你不拥有权重并且无法在其中对模型进行 RL,你将永远处于劣势,相比于那些同时拥有这两者的团队。

60 秒看懂编码智能体 RL

我对这个话题做了很多研究,制作了一些可视化图表试图解释关键部分,但我发现 Calvin French-Owen(Codex 团队的 MTS,Segment 创始人)在 AI Council 的一次演讲中做得更好、更清晰,所以我受他幻灯片的启发,在这里放一个动画:

https://github.com/opendilab/awesome-RLVR

要让模型在编码方面变得更好,你要:

  1. 生成一些编码智能体轨迹来解决一个问题(例如修复我的测试)

  2. 根据某些标准(验证器)对轨迹进行评分

  3. 更新模型权重,使好的轨迹更有可能发生,坏的轨迹更不可能发生

然后在几周或几个月的时间里,你要这样做数百万次。

不过,这些事情的“评分”部分往往倾向于随意的一维标准。

糟糕的设计没有惩罚

以 SWE-bench Multilingual 为例。任务很小——每个大约 15 分钟的工作量——从 Redis、jq 和 Django 等开源仓库中抓取。奖励基于以下两点为一或零:

  • FAIL_TO_PASS——你修复了被要求修复的问题吗?

  • PASS_TO_PASS——你在不破坏其他任何东西的情况下做到了吗?

这是一个真实的例子,来自 fastlane(一个 Ruby 项目)的 fastlane__fastlane-19304。它的 zip 动作抓取两个可选参数并立即对它们调用 .empty?,所以一旦你省略了 include 和 exclude,它就会崩溃:

https://factory.strongdm.ai/

关闭此特定问题的人工修复是两行(将 nil 默认为空数组):

https://www.latent.space/p/brex

在评估期间,模型

  1. base commit 开始——仓库被检出到该修复落地前的那一刻

  2. Bug 报告——在这个例子中是 'zip_command': undefined method 'empty?' for nil:NilClass

智能体根据问题去编写一些代码。它看不到黄金补丁或作为评分者的测试补丁

然后:

  1. 我们保留它生成的任何补丁,然后

  2. 丢弃它对测试文件所做的任何编辑(我们要抓到模型悄悄注释掉失败的测试或拼接一个使测试无效的 Mock)

  3. 在此之上应用基准测试的测试补丁,并且

  4. 运行整个套件:现有的 zip 测试(PASS_TO_PASS)加上新的测试(FAIL_TO_PASS),看看它们是否都通过

https://hlyr.dev/nva

题外话——基准测试不是验证器——实际上它们必须彼此隔离(不要在测试集上训练,等等)——我主要是想以此传达“判断编码智能体轨迹质量”的形态及其局限性。

模型如何得到正确答案并不重要。如果测试通过,我们就赢了,但侵蚀代码库可维护性没有惩罚

侵蚀代码库可维护性没有惩罚

这就是为什么你会看到到处都是 try-catch:

http://homepages.cs.ncl.ac.uk/brian.randell/NATO/nato1968.PDF

验证质量比“测试是否通过”难几个数量级

运行测试可以在几秒钟内得到明确的通过或失败。这就是为什么 RL 可以运行数百万个循环来优化每个模型生成。

但糟糕架构的成本函数是以周、月甚至年来衡量的。它发生在某人第一次打开那个文件进行一行更改时,意识到他们无法在一行内完成——某人“氛围”得有点过头了,现在我们必须在十一个地方进行同样的编辑,并希望三个文件之外没有任何东西悄悄崩溃。

https://www.youtube.com/watch?v=RjfbvDXpFls

测试在几秒钟内给你反馈,但糟糕架构的成本函数是以周、月甚至年来衡量的

糟糕的设计是当今基准测试无法评估的唯一事情。我知道,我知道,RL != 基准测试,但如果这在 RL 中解决了,我很确定它也会开始体现在我们基准测试的设计中。

无论如何,我个人不相信当今基准测试上的任何改进能作为模型突然擅长不把你的代码库搞成垃圾的指标。

前沿正在慢慢变好

当然,很多聪明人正在研究这个问题。我的观点不是说这不可能做到,而是炒作跑在了学科前面。

我认为几个努力的方向是正确的:

  • SWE-Marathon (Abundant AI):约 400 小时的任务,如“克隆整个 Excel,每个功能”——具有复合奖励通道,而不是单一的通过/失败位

  • DeepSWE (Datacurve):OSS 仓库上的大型任务,这些任务在现实世界中从未真正构建过,因此根据构造,它们不可能已经存在于训练集中(解决了污染问题,但没有解决质量问题)

  • Frontier Code (Cognition):多 PR 任务,以及一个巧妙的决定性评估质量的举措——它惩罚模型编写在补丁前代码上不失败的测试(如果你从未听说过变异测试,那你将有一段有趣的旅程)。它还运行一个判断模型来检查差异是否符合代码质量规则

https://x.com/jeffreyhuber

但模型判断质量的能力有限。

实际上,不难想象,如果一个模型能可靠地分辨好代码和坏代码,它一开始可能就写出了好的版本。RL 需要一个快速+可靠的预言机,而我们还没有用于可维护性的预言机。

如果一个模型能可靠地分辨好代码和坏代码,它一开始可能就写出了好的版本,但可维护性没有快速的预言机,所以我们无法在 RL 期间对此进行奖励

当然,更多的审查智能体和更多的 token 确实有帮助——它们提高了下限,捕捉愚蠢的错误。

但它们无法提高上限,因为上限是我们在 RL 中设法教会模型的东西,而良好的设计是我们仍然不知道如何教它的东西。

所以我仍然不会把我的代码库押注在这些任何一项上。但它们是我见过的第一批试图评估可维护性而不是停留在通过/失败上的评估方法。

题外话 也许未来的模型能直接搞定这个,我们可以停止了。如果你想在 GPT-7 发布前疯狂提示试试看,请便——但苦涩的教训见鬼去吧,我们现在就有问题要解决,我将介绍我们是如何做的。

重新开灯

今天我才知道 Twitter 文章有“媒体限制”,这意味着剩下的内容将放入第二部分——敬请期待

gave up on waiting on nikita for the articles fix - part 1 is here part 2 is coming http://x.com/i/article/2078710413345402880

or: the harness is not enough

Update - the talk version of this post is live on youtube: https://www.youtube.com/watch?v=Ib5GBkD555M* *

Part 1 in a series. Part 2 is here: https://x.com/dexhorthy/status/2081058573556306030

放弃等待 nikita 修复文章了——第一部分在这里,第二部分即将推出 http://x.com/i/article/2078710413345402880

或者:光靠 Harness 是不够的

更新——本文的演讲版本已在 YouTube 上线:https://www.youtube.com/watch?v=Ib5GBkD555M*

系列文章的第一部分。第二部分在这里:https://x.com/dexhorthy/status/2081058573556306030

i guess we doin loops now

We're all racing to put AI coding into production. A lot has been said about loop engineering, and the prevailing wisdom is that we should probably write more loops.

StrongDM wrote about their lights-off software factory where no human reads code and no human writes code.

The narrative goes something like this:

  1. You are the bottleneck.

  2. The models are good enough.

  3. Code is free.

  4. Just ship more stuff.

Ryan Lopopolo of OpenAI wrote about this in February and gave a talk in April about OpenAI's software factory, Symphony.

These people are all really dang smart and I have a ton of respect for them. But the most cynical take here would be to call this yet another excuse to pump more VC money into the slop cannon.

看来我们现在得搞循环了

我们都在竞相将 AI 编码投入生产。关于循环工程,人们已经讨论了很多,主流观点是我们大概应该写更多的循环。

StrongDM 写到了他们的“熄灯”软件工厂,在那里没有人阅读代码,也没有人编写代码。

其叙事逻辑大概是这样的:

  1. 你是瓶颈。

  2. 模型已经足够好了。

  3. 代码是免费的。

  4. 尽管多发布点东西。

OpenAI 的 Ryan Lopopolo 在二月写到了这一点,并在四月发表了一次演讲,介绍了 OpenAI 的软件工厂 Symphony。

这些人都非常聪明,我对他们怀有极大的敬意。但这里最愤世嫉俗的看法是,这不过是又一个借口,用来向“垃圾大炮”里注入更多的 VC 资金。

it's uh...it's going

Our friend Mario got up at AI Engineer Europe and begged us to slow down -- because companies that have no business having outages due to coding-agent mishaps, are, well... having outages due to coding-agent mishaps.

As Matt Pocock put it, codebases are falling apart faster than they ever have before.

I haven't been able to dig up any definitive data/findings from StrongDM on how that whole dark factory went. The weather-report has a few sparse updates between February and June of this year. edit - there is some conversation with the team on hacker news on July 23 - sounds like we might get a more formal update soon!

The folks at Faros AI put out a report: since we2 all picked up these AI coding tools back in January and February, pull-request review quality is way down.

  • More comments, longer comments, and tons of PRs getting merged with no review at all.

  • Incidents are way up.

  • Bugs per developer are way up.

This report is more of a correlation signal than a verifiable smoking gun (yes i chose that word on purpose, don't get me started on claude prose), and the whole point of this post is to be wary of slop data, but it feels directionally valid based on what I've seen.

呃……进展如何

我们的朋友 Mario 在 AI Engineer Europe 大会上站起来恳求我们慢下来——因为那些本不该因编码智能体失误而发生故障的公司,嗯……正因为编码智能体的失误而发生故障。

正如 Matt Pocock 所说,代码库正在以前所未有的速度分崩离析。

我没能从 StrongDM 找到关于那个“黑暗工厂”整体运行的任何确切数据或发现。Weather-report 在今年二月到六月之间只有零星的几次更新。编辑——7 月 23 日在 Hacker News 上有与团队的一些对话——听起来我们可能很快就会得到更正式的更新!

Faros AI 的团队发布了一份报告:自从我们在一月和二月开始使用这些 AI 编码工具以来,Pull Request(PR)的审查质量大幅下降。

  • 评论更多、篇幅更长,还有大量 PR 根本未经审查就被合并。

  • 事故大幅增加。

  • 每位开发者的 Bug 数量大幅增加。

这份报告更多是一种相关性信号,而非确凿的证据(是的,我是故意选这个词的,别让我开始吐槽 Claude 的文风),本文的重点就是要警惕“垃圾数据”,但根据我的观察,它在方向上感觉是正确的。

"You're holding it wrong" (you're not)

A lot of people will tell you that this is a skill issue -- that if you're not getting good results, that's your fault.

But however you're choosing to...erhm...hold it, I guarantee you're being told that if token-maxxing isn't working for you, it's a skill issue. You just need to spend more tokens. Let go of reading the code. And if you're just getting there, I promise it's part of the progression. I thought this way last summer too.

Unfortunately for my ego, some dumb stuff I decided to say about "how to hold it better" got recorded and now has about a million cumulative views on YouTube. I am not trying to brag here, I share this only to establish that I've been going deep on the best ways to use coding agents for a long time now, and have discovered some things that many others have found genuinely useful.

  • Advanced Context Engineering for Coding Agents

  • No Vibes Allowed -- Solving Hard Problems in Complex Codebases

  • Everything We Got Wrong About RPI

Anyhow, The promise of all this online "just token harder" yapping we've been forced to endure is, succinctly: with enough harness engineering, we can get the best of both worlds:

  • 10 to 100x faster,

  • high quality, and

  • nobody ever has to do that thing we all hate called code review

All we have to do is configure more linters and sprinkle some magic words like "adversarial review" onto enough PR review bots, and our software will happily build itself without incident.

“你拿手机的姿势不对”(其实并不是)

很多人会告诉你这是一个技能问题——如果你没有得到好的结果,那是你的错。

但无论你选择怎么……呃……“拿它”,我敢保证有人会告诉你,如果疯狂堆 token(token-maxxing)对你不起作用,那是技能问题。你只需要花费更多的 token。放弃阅读代码吧。如果你才刚开始上手,我保证这是进阶过程的一部分。去年夏天我也是这么想的。

不幸的是,为了我的自尊心,我决定说的一些关于“如何更好地拿它”的蠢话被录了下来,现在在 YouTube 上已经有大约一百万的累计观看量。我并不是想在这里吹嘘,我分享这个只是为了说明,我深入钻研编码智能体的最佳使用方法已经很长时间了,并且发现了一些许多其他人认为真正有用的东西。

  • 编码智能体的高级上下文工程

  • 拒绝“氛围”——在复杂代码库中解决难题

  • 我们关于 RPI 理解错的一切

无论如何,我们被迫忍受的所有这些关于“只要加大 token 力度”的网络废话,其承诺简而言之就是:通过足够的 Harness 工程,我们可以两全其美:

  • 速度快 10 到 100 倍,

  • 高质量,以及

  • 没人需要做那个我们都讨厌的叫作代码审查的事情

我们所要做的就是配置更多的 Linter,并在足够多的 PR 审查机器人上撒点魔法词,比如“对抗性审查”,我们的软件就会愉快地自行构建,且不出意外。

This is not a skill issue

What I'm gonna try to convince you is that no amount of harness engineering or loopsmaxxing can solve what is fundamentally a model-training issue.

To grapple with this, I had to dig into how coding models are actually trained and evaluated - with respect to both the RLVR and the benchmark side of things.

In this post I'm gonna run through:

  1. Software factories date back to 1968, how have they evolved, and how has AI changed them

  2. Why models can generate mountains of slop despite ace-ing benchmarks (even the brand new "frontier" benchmarks)

  3. In spite of this, you can move pretty fast without setting your codebase on fire

I'm gonna try to cut through the hype of every daily-emerging skills plugin and the ai-psychosis-tokenmaxxing advice pandemic, and talk in general terms about the types of things that work without referencing any particular skill or framework.

Video Version: this post is based on (and expands upon) my keynote at AI Engineer World's Fair 2026.

Thanks to @addyosmani, @CyrusNewDay, @HamelHusain, @zeeg, @dillon_mulroy, @nayshins, and @jeffreyhuber for feedback on this post.

这不是技能问题

我想说服你的是,再多的 Harness 工程或循环最大化也无法解决根本上的模型训练问题。

为了解决这个问题,我不得不深入研究编码模型实际上是如何训练和评估的——包括 RLVR 和基准测试两个方面。

在这篇文章中,我将梳理:

  1. 软件工厂可追溯至 1968 年,它们是如何演变的,AI 又是如何改变它们的

  2. 为什么模型在基准测试(即使是全新的“前沿”基准测试)中表现优异,却能生成堆积如山的垃圾代码

  3. 尽管如此,你仍然可以在不把代码库付之一炬的情况下快速推进

我将试图穿透每天涌现的技能插件炒作和 AI 精神错乱般的 token-maxxing 建议大流行,用通俗的语言谈谈有效的方法类型,而不引用任何特定的技能或框架。

视频版本: 本文基于(并扩展了)我在 2026 年 AI Engineer World's Fair 上的主题演讲。

感谢 @addyosmani、@CyrusNewDay、@HamelHusain、@zeeg、@dillon_mulroy、@nayshins 和 @jeffreyhuber 对本文的反馈。

An aside: this has nothing to do with vibe coding

Addy Osmani detangled this thing that is worth highlighting:

A developer vibe-coding a side project a dozen people will ever run, and a team keeping a ten-year-old enterprise system alive for another quarter, share almost no constraints worth naming, and most of the advice in circulation is really one of those two people telling the other how to live.

If you love vibe coding, please, go on vibing. I still vibe code lots of things, I just also maintain lots of production software (and through HumanLayer, help 1000s of other engineers do the same), so the rest of this is aimed at folks solving hard problems in complex codebases.

I hear the word brownfield a lot to talk about this split. Historically that meant some ten-year-old Java thing, but at the pace we can ship now, it feels like an agent-built codebase starts to struggle after maybe three to six months -- you start to slow down, and the way you approach adding new things has to change.

题外话:这与“氛围编码”无关

Addy Osmani 梳理了一个值得强调的观点:

一个为只有十几个人使用的副业项目进行“氛围编码”的开发者,与一个让一个十年的企业系统再苟延残喘一个季度的团队,几乎没有任何值得一提的共同约束,而流传的大多数建议实际上是这两类人中的一方在告诉另一方该如何生活。

如果你喜欢“氛围编码”,请继续。我仍然会对很多东西进行“氛围编码”,我只是同时也维护大量的生产软件(并通过 HumanLayer 帮助成千上万的其他工程师做同样的事),所以接下来的内容是针对那些在复杂代码库中解决难题的人。

我经常听到 棕地 这个词来描述这种分裂。历史上那是指某些十年的 Java 老古董,但按照我们现在的交付速度,感觉智能体构建的代码库可能在三到六个月后就开始挣扎——你开始慢下来,你添加新事物的方式必须改变。

A brief history of the software factory

I've been building and studying software factories my whole career, but I only learned this recently: the term traces all the way back to a NATO conference in 1968 -- the same one that gave us "software engineering."

The only other bit I find super interesting since then is that the US Department of Defense wrote a 31-page pdf about how the DoD needs to start using jenkins better or something.

软件工厂简史

我整个职业生涯都在构建和研究软件工厂,但我最近才知道:这个词可以追溯到 1968 年的一次北约(NATO)会议——正是那次会议给了我们“软件工程”这个词。

自那以后,我发现唯一非常有趣的一点是美国国防部写了一份 31 页的 PDF,讲述国防部需要如何开始更好地使用 Jenkins 之类的东西。

The 2022 software factory

Let's ground our "software factory" definition around 2022, right before AI. In a typical software factory:

  • People decide what to build -- engineers, PMs, leadership driving the vision

  • It goes in a tracker -- Linear, Jira, whatever: a state machine of what needs to happen

  • Someone grabs a ticket and builds it -- probably does some manual/automated testing while they're at it

  • Pull request -- automated checks, a human reviews the code, maybe someone pulls it down to test

  • Anything wrong? Loop back to "someone builds the thing"

  • Ship to prod -- and it makes contact with users

  • Add monitoring -- there's an entire industry built around paging an engineer at 3am when something breaks

  • Users complain -- ask for things, find bugs, file feature requests → back to the team to add to the tracker

https://hlyr.dev/ace

And on and on. We haven't even hit AI yet, and there are already several loops in this picture.

2022 年的软件工厂

让我们把“软件工厂”的定义定位在 2022 年,就在 AI 出现之前。在一个典型的软件工厂中:

  • 人决定构建什么——工程师、产品经理、领导层驱动愿景

  • 放入追踪器——Linear、Jira 或其他:一个记录待办事项的状态机

  • 某人领取工单并构建——可能在此过程中进行一些手动/自动化测试

  • Pull Request——自动化检查,人工审查代码,可能有人拉取代码进行测试

  • 有问题吗?循环回到“某人构建东西”

  • 发布到生产环境——与用户接触

  • 添加监控——有一个完整的行业围绕在凌晨 3 点出问题时呼叫工程师

  • 用户抱怨——提出需求、发现 Bug、提交功能请求 → 回到团队添加到追踪器

https://hlyr.dev/ace

如此循环往复。我们还没涉及到 AI,但这张图中已经有好几个循环了。

front-loading alignment

There's a thing teams figured out decades ago: building takes hours or days, and so does review.

https://en.wikipedia.org/wiki/Mutation_testing

So we front-load the work -- planning, architecture proposals, sprint planning -- together, as a team. That means:

  • less rework, because we aligned before anyone wrote code

  • less time reviewing every line, if you've ever read a long-but-well-done PR, you know how fast the review goes when it's close-to-perfect

https://www.youtube.com/watch?v=am_oeAoUhew

We'll come back to this later - let's look at what happens when you bring agentic coding into the picture.

前置对齐

几十年前团队就发现了一件事:构建需要数小时或数天,审查也是如此。

https://en.wikipedia.org/wiki/Mutation_testing

所以我们把工作前置——作为团队一起进行规划、架构提案、冲刺计划。这意味着:

  • 更少的返工,因为我们在任何人编写代码之前就已经对齐了

  • 更少的时间逐行审查,如果你读过一份冗长但做得很好的 PR,你就知道当它接近完美时审查进行得有多快

https://www.youtube.com/watch?v=am_oeAoUhew

我们稍后会回到这个话题——先让我们看看把智能体编码引入画面会发生什么。

The agentic software factory

Now every company and their mother --

  • Ramp

  • Stripe

  • WorkOS

  • Brex

has spent the better part of this year explaining how they built an agent factory that ships on the order of 75% of their code.

The agentic factory looks mostly like swapping "someone builds the thing" → "an agent builds the thing" -- there's some stuff here like orchestration, a harness, a sandbox, a model, computer use, etc. I won't go in depth on those details because quite frankly I'm sick of reading about it and I'm sure you are too.

https://www.youtube.com/watch?v=Ib5GBkD555M

When the agent builds the thing:

  • Building drops from hours or days to minutes or hours.

  • Review still takes hours or days. A human still has to read the code and test the change. So review is now the bottleneck.

https://www.swe-marathon.org/

So you speed review up too:

  • Agentic code review, to catch style, bugs, security.

  • Agentic regression testing, to poke it from the outside with browsers and computer use and maybe send you a cute little video when it's done

https://www.oreilly.com/library/view/clean-code-a/9780136083238/

Review is faster now, but it's also probably still the bottleneck. But we can do more loops.

Next you might route incidents into the factory. Instead of paging someone at 3am, they wake up to a PR that maybe already fixes it.

We can also route user feedback into the factory. People ask for stuff, it gets built.

https://www.youtube.com/watch?v=q-ntX4DLW_c

At which point the job is two questions: how much can you stuff into the queue, and how fast can you review and test what comes out?

Which brings us to the lights-off software factory.

智能体软件工厂

现在每家公司及其母公司——

  • Ramp

  • Stripe

  • WorkOS

  • Brex

都在今年花了大量时间解释他们如何构建了一个智能体工厂,交付了他们 75% 的代码

智能体工厂看起来主要像是把 “某人构建东西” → “智能体构建东西”——这里有一些东西,比如编排、Harness、沙箱、模型、计算机使用等。我不会深入探讨这些细节,因为坦白说,我已经读腻了,我相信你也一样。

https://www.youtube.com/watch?v=Ib5GBkD555M

当智能体构建东西时:

  • 构建时间从数小时或数天缩短到数分钟或数小时。

  • 审查仍然需要数小时或数天。人类仍然需要阅读代码并测试变更。所以审查现在是瓶颈。

https://www.swe-marathon.org/

所以你也要加快审查速度:

  • 智能体代码审查,以捕捉风格、Bug、安全问题。

  • 智能体回归测试,用浏览器和计算机使用从外部进行探测,也许完成后给你发个可爱的小视频

https://www.oreilly.com/library/view/clean-code-a/9780136083238/

审查现在更快了,但也可能仍然是瓶颈。但我们可以做更多的循环。

接下来,你可能会将事故路由到工厂。不再是凌晨 3 点呼叫某人,而是他们醒来时看到一个可能已经修复了问题的 PR。

我们也可以将用户反馈路由到工厂。人们提出需求,然后被构建出来。

https://www.youtube.com/watch?v=q-ntX4DLW_c

到了这个阶段,工作就变成了两个问题:你能往队列里塞多少东西,以及你能多快审查和测试产出的结果?

这就引出了“熄灯”软件工厂。

The lights-off software factory

Dan Shapiro coined this term and Simon Willison wrote about StrongDM's implementation of it -- where we no longer read the code.

You look at your beautiful software factory. It's ruined by that annoying little code review step and you say: you know what, that thing where a human reads every change? No thanks.

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#user-content-fn-5-f56fe9a973fe7c6ebb6a9673c1bc64cb

So you drop it, and you put the effort somewhere else:

  • Invest in testing and letting the agent test its own work

  • Invest in sandboxes and orchestration

  • Invest in automated review

  • Invest in monitoring

  • Invest in rollout

  • Invest in collecting feedback signals from users

https://garryslist.org/posts/boil-the-ocean

And now the job really is just one question: how much stuff can we ask the agent to build? How much of the ocean do we want to boil?

“熄灯”软件工厂

Dan Shapiro 创造了这个词,Simon Willison 写了关于 StrongDM 的实现——在那里我们不再阅读代码。

你看着你漂亮的软件工厂。它被那个烦人的小代码审查步骤毁了,你说:你知道吗,那个让人阅读每个变更的步骤?不用了,谢谢。

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#user-content-fn-5-f56fe9a973fe7c6ebb6a9673c1bc64cb

所以你放弃了它,把精力放在别处:

  • 投资测试,让智能体测试自己的工作

  • 投资沙箱和编排

  • 投资自动化审查

  • 投资监控

  • 投资发布

  • 投资收集用户反馈信号

https://garryslist.org/posts/boil-the-ocean

现在工作真的只剩下一个问题:我们可以要求智能体构建多少东西?我们要煮沸多少海水?

This is going to go great (its not)

https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents

I'm going to posit something potentially controversial: the lights off factory does not work.

Let's get into why software factories fail.

这会进展顺利(其实不会)

https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents

我要提出一个可能有争议的观点:熄灯工厂行不通。

让我们深入探讨软件工厂为何失败。

We tried this

In July 2025 we went full lights-off. Just read the specs and the tickets, background agents for all the small/medium stuff, the whole thing.

If you've tried this seriously for a few months, you already know how it ends. You find at least one issue gnarly enough that the agent can't solve it -- even with your most advanced prompting and workflows.

  • You do deep context-aware research, collating all the right parts into the smart zone for the model to analyze

  • You have the agent try to reproduce in 10 different ways

Eventually you have to suck it up and go dig into the codebase you stopped reading three months ago, trying to figure out what's broken.

And in the meantime:

  • Your site was down.

  • Your users were pissed.

  • And you, if you're anything like me, were miserable -- reading all the slop code you let slip into your system.

The first time this happened to us, I shook it off. Even though I'd just spent the better part of two weeks digging through claude spaghetti, "the downside risk was worth the velocity". By the ~third time in november, we decided it would be easier to rewrite from scratch, and my cofounder spent two whole weeks in VS Code (not even cursor) plumbing out all the patterns by hand.

我们试过了

2025 年 7 月,我们全面实施了熄灯模式。只读规格说明和工单,所有中小任务都用后台智能体处理,全套流程。

如果你认真尝试了几个月,你就知道结局如何。你至少会遇到一个棘手到智能体无法解决的问题——即使使用你最先进的提示和工作流。

  • 你进行深入的上下文感知研究,将所有正确的部分整理到智能区域供模型分析

  • 你让智能体尝试用 10 种不同的方式复现

最终你必须硬着头皮去深入研究你三个月前停止阅读的代码库,试图弄清楚哪里出了问题。

与此同时:

  • 你的网站挂了。

  • 你的用户很生气。

  • 而你,如果你像我一样,会很痛苦——阅读所有那些你放进系统的垃圾代码。

第一次发生这种情况时,我不以为意。尽管我刚刚花了大半个月时间梳理 Claude 生成的意大利面条式代码,“下行风险是值得的,因为换来了速度”。到了 11 月大约第三次发生时,我们决定从头重写更容易,我的联合创始人花了整整两周在 VS Code(甚至不是 Cursor)里手动梳理所有模式。

models degrade codebase quality over time

What I want to get to is this: models have a shortcoming. They can't maintain and improve codebase quality over time -- not without a decent amount of human steering.4

When I say maintainability, I mean the specific thing where it becomes really, really hard to change one part of the codebase without breaking another part. This is Martin Fowler's shotgun surgery.

I'm not going to say much more about maintainability. There are a bunch of books you can go read about it

  • John Ousterhout's A Philosophy of Software Design

  • Robert C. Martin's Clean Code

  • Martin Fowler's Refactoring

So, why can't models do software maintainability?

模型会随时间降低代码库质量

我想表达的是:模型有一个缺点。它们无法随时间维护和提高代码库质量——如果没有大量的人工引导。

当我说可维护性时,我指的是那种特定的情况:在不破坏另一部分代码的情况下,很难更改代码库的某一部分。这就是 Martin Fowler 所说的“散弹式修改”。

关于可维护性我就不多说了。有很多书你可以去读:

  • John Ousterhout 的《软件设计哲学》

  • Robert C. Martin 的《代码整洁之道》

  • Martin Fowler 的《重构》

那么,模型为什么不能做软件可维护性呢?

"But surely the models have gotten better since then"

At this point you might be dying to say: but Dex, surely the models have gotten much better since July

They have -- in some ways. In others they're about the same.

  • Solving one-off problems, or vibe-coding a new marketing site? Yes. Way better.

  • Improving codebase quality over time? Not much better, as far as I can tell.

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#user-content-fn-1b-f56fe9a973fe7c6ebb6a9673c1bc64cb

I can't prove this. You can't prove it either. There are no good benchmarks for a model's ability to maintain codebase quality. (More on where that's going later.)

**THERE ARE NO GOOD BENCHMARKS for a model's ability to maintain codebase quality **

But if you've worked with coding agents for a while -- and a lot of people are posting about exactly this -- you probably have the vibe already: they tend to make things worse over time, and make the codebase harder to work in.

So to figure out why this happens, I want to zoom out to the first great coding agent.

“但自那以后模型肯定变好了吧”

这时你可能迫不及待想说:但是 Dex,自七月以来模型肯定变得好多了

它们确实变好了——在某些方面。在其他方面,它们差不多还是老样子。

  • 解决一次性问题,或为新的营销网站进行“氛围编码”?是的。好多了。

  • 随时间提高代码库质量?据我所知,没好多少。

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#user-content-fn-1b-f56fe9a973fe7c6ebb6a9673c1bc64cb

我无法证明这一点。你也无法证明。对于模型维护代码库质量的能力,没有好的基准测试。(稍后会谈到这方面的走向。)

对于模型维护代码库质量的能力,目前没有好的基准测试

但如果你使用编码智能体有一段时间了——很多人都在发帖谈论这一点——你可能已经有这种感觉了:它们倾向于随时间推移让事情变得更糟,让代码库更难工作。

所以为了弄清楚为什么会发生这种情况,我想把视线拉远,看看第一个伟大的编码智能体。

Claude Code won because of Reinforcement Learning inside the harness

Claude Code went from nothing to ~$4B -- now something like ~$9B -- in revenue in under a year.

Which is a little wild, because there were already great CLI agents. aider, cline, codebuff -- all predated Claude Code, all with genuinely great context engineering built in, all with the same tool set you might attribute to claude code: read, write, edit, grep, bash. I used them. They were good. But also, tool use would just... fail sometimes -- you'd watch it flail at the same edit three times and open your editor back up to do it yourself.

The SWE-Agent paper from 2024 outlines how small changes in tool shape make noticeable differences, e.g. including line numbers in ReadFile results, or changing an Edit tool from find/replace to line-range edits.

https://web.stanford.edu/~ouster/cgi-bin/aposd.php

Then Claude Code launched and went vertical pretty quickly. You can hand-wave this as distribution, but the canonically-accepted explanation is that claude code won because it was better, and that it was better because Anthropic RL'd the model inside the harness -- the first time a lab trained a model against the exact tools they were going to ship it with. And it got really, really good at calling those tools in an agentic loop.

It's one thing to fiddle with tool definitions and evals until you find the shape the model likes best -- I've burned weeks doing this for various use cases. It's a different game when you own the weights and can modify the model itself to be better at a particular set of tools.

The OpenAI team gave a talk in November that put this pretty well: if you build a harness but you don't own the weights and can't RL the model inside it, you'll always be at a disadvantage to a team that owns both.

Claude Code 胜在 Harness 内部的强化学习

Claude Code 在不到一年的时间内,收入从零增长到约 40 亿美元——现在大约是 90 亿美元。

这有点疯狂,因为已经有很多优秀的 CLI 智能体。aider、cline、codebuff——都早于 Claude Code,都内置了真正优秀的上下文工程,都拥有你可能归功于 Claude Code 的相同工具集:读取、写入、编辑、grep、bash。我用过它们。它们很好。但同样,工具使用有时会……失败——你会看着它在同一个编辑上折腾三次,然后不得不重新打开编辑器自己动手。

2024 年的 SWE-Agent 论文概述了工具形态的微小变化如何产生显著差异,例如在 ReadFile 结果中包含行号,或将 Edit 工具从查找/替换更改为行范围编辑。

https://web.stanford.edu/~ouster/cgi-bin/aposd.php

然后 Claude Code 推出并迅速腾飞。你可以把这归结为分发渠道,但公认的解释是 Claude Code 赢是因为它更好,而它更好是因为 Anthropic 在 Harness 内部 对模型进行了 RL——这是实验室第一次针对它们将要发布的特定工具集训练模型。它变得非常、非常擅长在智能体循环中调用这些工具。

调整工具定义和评估以找到模型最喜欢的形态是一回事——我为各种用例花了几周时间做这个。但当你拥有权重并且可以修改模型本身以更好地使用特定工具集时,那就是另一回事了。

OpenAI 团队在 11 月的一次演讲中说得很好:如果你构建了一个 Harness 但你不拥有权重并且无法在其中对模型进行 RL,你将永远处于劣势,相比于那些同时拥有这两者的团队。

Coding Agent RL in 60 seconds

I did a bunch of research on this topic and cooked up a bunch of visualizations to try to explain the parts that matter, but I found that Calvin French-Owen (MTS on the codex team, founder of Segment) did a talk at AI Council that did a much better and cleaner job, so I'm just gonna drop this animation here inspired by his slides:

https://github.com/opendilab/awesome-RLVR

To make a model better at coding, you're gonna:

  1. generate some coding agent traces to solve a problem (e.g. fix my tests)

  2. score the traces based on some criteria (verifier)

  3. update the model weights to make the good traces more likely, and the bad traces less likely

And then you do this millions of times over the course of weeks or months.

The "scoring" part of these things can tend to be whimsically one-dimensional though.

60 秒看懂编码智能体 RL

我对这个话题做了很多研究,制作了一些可视化图表试图解释关键部分,但我发现 Calvin French-Owen(Codex 团队的 MTS,Segment 创始人)在 AI Council 的一次演讲中做得更好、更清晰,所以我受他幻灯片的启发,在这里放一个动画:

https://github.com/opendilab/awesome-RLVR

要让模型在编码方面变得更好,你要:

  1. 生成一些编码智能体轨迹来解决一个问题(例如修复我的测试)

  2. 根据某些标准(验证器)对轨迹进行评分

  3. 更新模型权重,使好的轨迹更有可能发生,坏的轨迹更不可能发生

然后在几周或几个月的时间里,你要这样做数百万次。

不过,这些事情的“评分”部分往往倾向于随意的一维标准。

There's no penalty for bad design

Take SWE-bench Multilingual. The tasks are small -- about fifteen minutes of work apiece -- scraped out of open-source repos like Redis, jq, and Django. The reward is one or zero based on:

  • FAIL_TO_PASS - did you fix the thing you were asked to fix?

  • PASS_TO_PASS - did you do it without breaking anything else?

Here's a real one, fastlane__fastlane-19304, from fastlane -- a Ruby project. Its zip action grabs two optional params and calls .empty? on them straight away, so the moment you leave include and exclude off, it falls over:

https://factory.strongdm.ai/

The human fix that closed this particular issue is two lines (default nils to empty arrays):

https://www.latent.space/p/brex

During the evaluation, the model

  1. starts from a base commit -- the repo checked out to the moment right before that fix landed

  2. the bug report - in this case 'zip_command': undefined method 'empty?' for nil:NilClass

The agent goes off and writes some code based on the issue. It doesn't see the golden patch or the test patch that serves as the grader:

Then:

  1. We keep whatever patch it produced, then

  2. Throw away any edits it made to the test files (we've caught a model quietly commenting out the failing test or splicing in a mock that makes the test useless)

  3. Apply the benchmark's test patch on top, and

  4. Run the whole suite: the existing zip tests (PASS_TO_PASS) plus the new one (FAIL_TO_PASS) to see if they both pass

https://hlyr.dev/nva

Aside - Benchmarks are not verifiers - in fact they have to be held out from each other (don't train on test, yada yada) - I primarily mean this to convey the shape of "judging the quality of a coding agent trace" and its limitations.

How the model got to a correct answer doesn't matter. If the tests pass, we win, but there is no penalty for eroding codebase maintainability.

there is no penalty for eroding codebase maintainability

That's how you get try catches around everything:

http://homepages.cs.ncl.ac.uk/brian.randell/NATO/nato1968.PDF

糟糕的设计没有惩罚

以 SWE-bench Multilingual 为例。任务很小——每个大约 15 分钟的工作量——从 Redis、jq 和 Django 等开源仓库中抓取。奖励基于以下两点为一或零:

  • FAIL_TO_PASS——你修复了被要求修复的问题吗?

  • PASS_TO_PASS——你在不破坏其他任何东西的情况下做到了吗?

这是一个真实的例子,来自 fastlane(一个 Ruby 项目)的 fastlane__fastlane-19304。它的 zip 动作抓取两个可选参数并立即对它们调用 .empty?,所以一旦你省略了 include 和 exclude,它就会崩溃:

https://factory.strongdm.ai/

关闭此特定问题的人工修复是两行(将 nil 默认为空数组):

https://www.latent.space/p/brex

在评估期间,模型

  1. base commit 开始——仓库被检出到该修复落地前的那一刻

  2. Bug 报告——在这个例子中是 'zip_command': undefined method 'empty?' for nil:NilClass

智能体根据问题去编写一些代码。它看不到黄金补丁或作为评分者的测试补丁

然后:

  1. 我们保留它生成的任何补丁,然后

  2. 丢弃它对测试文件所做的任何编辑(我们要抓到模型悄悄注释掉失败的测试或拼接一个使测试无效的 Mock)

  3. 在此之上应用基准测试的测试补丁,并且

  4. 运行整个套件:现有的 zip 测试(PASS_TO_PASS)加上新的测试(FAIL_TO_PASS),看看它们是否都通过

https://hlyr.dev/nva

题外话——基准测试不是验证器——实际上它们必须彼此隔离(不要在测试集上训练,等等)——我主要是想以此传达“判断编码智能体轨迹质量”的形态及其局限性。

模型如何得到正确答案并不重要。如果测试通过,我们就赢了,但侵蚀代码库可维护性没有惩罚

侵蚀代码库可维护性没有惩罚

这就是为什么你会看到到处都是 try-catch:

http://homepages.cs.ncl.ac.uk/brian.randell/NATO/nato1968.PDF

Verifying quality is orders of magnitude harder than "did the tests pass"

Running the tests gets you a clean pass or fail in ~seconds. That's why RL can run millions of loops to optimize each model generation.

But the cost function of bad architecture is measured in weeks, months, maybe even years. It happens the first time someone opens that file for a one-line change and realizes they can't make it in one line -- that someone vibed this a little too hard, and now we have to make the same edit in eleven places and hope nothing quietly breaks three files over.

https://www.youtube.com/watch?v=RjfbvDXpFls

Tests give you feedback in seconds, but the cost function of bad architecture is measured in weeks, months, maybe even years

Bad design is the one thing today's benchmarks can't evaluate. And I know, I know, RL != Benchmarks, but if this was solved in RL, I'm pretty sure it would start to show up in how our benchmarks are designed too.

In any case, I personally don't trust any improvements on today's benchmarks as an indicator that the models are suddenly good at not slopping up your codebase.

验证质量比“测试是否通过”难几个数量级

运行测试可以在几秒钟内得到明确的通过或失败。这就是为什么 RL 可以运行数百万个循环来优化每个模型生成。

但糟糕架构的成本函数是以周、月甚至年来衡量的。它发生在某人第一次打开那个文件进行一行更改时,意识到他们无法在一行内完成——某人“氛围”得有点过头了,现在我们必须在十一个地方进行同样的编辑,并希望三个文件之外没有任何东西悄悄崩溃。

https://www.youtube.com/watch?v=RjfbvDXpFls

测试在几秒钟内给你反馈,但糟糕架构的成本函数是以周、月甚至年来衡量的

糟糕的设计是当今基准测试无法评估的唯一事情。我知道,我知道,RL != 基准测试,但如果这在 RL 中解决了,我很确定它也会开始体现在我们基准测试的设计中。

无论如何,我个人不相信当今基准测试上的任何改进能作为模型突然擅长不把你的代码库搞成垃圾的指标。

The frontier is getting better, slowly

Of course lots of smart folks are working on this. My point is not that it can't be done, it's that the hype is outrunning the discipline.

A few efforts I think are pointed the right way:

  • SWE-Marathon (Abundant AI): ~400-hour tasks like "clone all of Excel, every feature" -- with a compound reward channel instead of a single pass/fail bit

  • DeepSWE (Datacurve): big tasks on OSS repos that were never actually built in the real world, so by construction they can't already be sitting in the training set (solves contamination, but not quality)

  • Frontier Code (Cognition): multi-PR tasks, and a clever move that evaluates quality deterministically -- it penalizes the model for writing tests that don't fail on the pre-patch code (if you've never heard about mutation testing you are in for a fun ride5). It also runs a judge model over the diff checking code-quality rules.

https://x.com/jeffreyhuber

But a model judging quality can only go so far.

In fact, it's not hard to imagine that if a model could reliably tell good code from bad, it might have written the good version to begin with. RL needs a fast+reliable oracle, and we don't yet have one for maintainability

if a model could reliably tell good code from bad, it might have written the good version to begin with, but maintainability has no fast oracle, so we can't reward for it during RL

Of course, more review agents and more tokens do help -- they raise the floor, catching the dumb stuff.

But they don't move the ceiling, because the ceiling is whatever we managed to teach the model in RL, and good design is the thing we still don't know how to teach it.

So I still wouldn't bet my codebase on any of these. But they're the first evals I've seen even trying to score maintainability instead of stopping at pass/fail.

Aside Maybe a future model just gets this and we can stop. If you want to yolo prompts until GPT-7 ships and find out, be my guest -- but bitter lesson be damned, we've got problems to solve now, and I'm gonna walk through how we do that.

前沿正在慢慢变好

当然,很多聪明人正在研究这个问题。我的观点不是说这不可能做到,而是炒作跑在了学科前面。

我认为几个努力的方向是正确的:

  • SWE-Marathon (Abundant AI):约 400 小时的任务,如“克隆整个 Excel,每个功能”——具有复合奖励通道,而不是单一的通过/失败位

  • DeepSWE (Datacurve):OSS 仓库上的大型任务,这些任务在现实世界中从未真正构建过,因此根据构造,它们不可能已经存在于训练集中(解决了污染问题,但没有解决质量问题)

  • Frontier Code (Cognition):多 PR 任务,以及一个巧妙的决定性评估质量的举措——它惩罚模型编写在补丁前代码上不失败的测试(如果你从未听说过变异测试,那你将有一段有趣的旅程)。它还运行一个判断模型来检查差异是否符合代码质量规则

https://x.com/jeffreyhuber

但模型判断质量的能力有限。

实际上,不难想象,如果一个模型能可靠地分辨好代码和坏代码,它一开始可能就写出了好的版本。RL 需要一个快速+可靠的预言机,而我们还没有用于可维护性的预言机。

如果一个模型能可靠地分辨好代码和坏代码,它一开始可能就写出了好的版本,但可维护性没有快速的预言机,所以我们无法在 RL 期间对此进行奖励

当然,更多的审查智能体和更多的 token 确实有帮助——它们提高了下限,捕捉愚蠢的错误。

但它们无法提高上限,因为上限是我们在 RL 中设法教会模型的东西,而良好的设计是我们仍然不知道如何教它的东西。

所以我仍然不会把我的代码库押注在这些任何一项上。但它们是我见过的第一批试图评估可维护性而不是停留在通过/失败上的评估方法。

题外话 也许未来的模型能直接搞定这个,我们可以停止了。如果你想在 GPT-7 发布前疯狂提示试试看,请便——但苦涩的教训见鬼去吧,我们现在就有问题要解决,我将介绍我们是如何做的。

Turning the lights back on

Today I learned that Twitter Articles have a "media limit" which means the rest of this is going into a part II post - stay tuned

重新开灯

今天我才知道 Twitter 文章有“媒体限制”,这意味着剩下的内容将放入第二部分——敬请期待

gave up on waiting on nikita for the articles fix - part 1 is here part 2 is coming http://x.com/i/article/2078710413345402880

or: the harness is not enough

Update - the talk version of this post is live on youtube: https://www.youtube.com/watch?v=Ib5GBkD555M* *

Part 1 in a series. Part 2 is here: https://x.com/dexhorthy/status/2081058573556306030

i guess we doin loops now

We're all racing to put AI coding into production. A lot has been said about loop engineering, and the prevailing wisdom is that we should probably write more loops.

StrongDM wrote about their lights-off software factory where no human reads code and no human writes code.

The narrative goes something like this:

  1. You are the bottleneck.

  2. The models are good enough.

  3. Code is free.

  4. Just ship more stuff.

Ryan Lopopolo of OpenAI wrote about this in February and gave a talk in April about OpenAI's software factory, Symphony.

These people are all really dang smart and I have a ton of respect for them. But the most cynical take here would be to call this yet another excuse to pump more VC money into the slop cannon.

it's uh...it's going

Our friend Mario got up at AI Engineer Europe and begged us to slow down -- because companies that have no business having outages due to coding-agent mishaps, are, well... having outages due to coding-agent mishaps.

As Matt Pocock put it, codebases are falling apart faster than they ever have before.

I haven't been able to dig up any definitive data/findings from StrongDM on how that whole dark factory went. The weather-report has a few sparse updates between February and June of this year. edit - there is some conversation with the team on hacker news on July 23 - sounds like we might get a more formal update soon!

The folks at Faros AI put out a report: since we2 all picked up these AI coding tools back in January and February, pull-request review quality is way down.

  • More comments, longer comments, and tons of PRs getting merged with no review at all.

  • Incidents are way up.

  • Bugs per developer are way up.

This report is more of a correlation signal than a verifiable smoking gun (yes i chose that word on purpose, don't get me started on claude prose), and the whole point of this post is to be wary of slop data, but it feels directionally valid based on what I've seen.

"You're holding it wrong" (you're not)

A lot of people will tell you that this is a skill issue -- that if you're not getting good results, that's your fault.

But however you're choosing to...erhm...hold it, I guarantee you're being told that if token-maxxing isn't working for you, it's a skill issue. You just need to spend more tokens. Let go of reading the code. And if you're just getting there, I promise it's part of the progression. I thought this way last summer too.

Unfortunately for my ego, some dumb stuff I decided to say about "how to hold it better" got recorded and now has about a million cumulative views on YouTube. I am not trying to brag here, I share this only to establish that I've been going deep on the best ways to use coding agents for a long time now, and have discovered some things that many others have found genuinely useful.

  • Advanced Context Engineering for Coding Agents

  • No Vibes Allowed -- Solving Hard Problems in Complex Codebases

  • Everything We Got Wrong About RPI

Anyhow, The promise of all this online "just token harder" yapping we've been forced to endure is, succinctly: with enough harness engineering, we can get the best of both worlds:

  • 10 to 100x faster,

  • high quality, and

  • nobody ever has to do that thing we all hate called code review

All we have to do is configure more linters and sprinkle some magic words like "adversarial review" onto enough PR review bots, and our software will happily build itself without incident.

This is not a skill issue

What I'm gonna try to convince you is that no amount of harness engineering or loopsmaxxing can solve what is fundamentally a model-training issue.

To grapple with this, I had to dig into how coding models are actually trained and evaluated - with respect to both the RLVR and the benchmark side of things.

In this post I'm gonna run through:

  1. Software factories date back to 1968, how have they evolved, and how has AI changed them

  2. Why models can generate mountains of slop despite ace-ing benchmarks (even the brand new "frontier" benchmarks)

  3. In spite of this, you can move pretty fast without setting your codebase on fire

I'm gonna try to cut through the hype of every daily-emerging skills plugin and the ai-psychosis-tokenmaxxing advice pandemic, and talk in general terms about the types of things that work without referencing any particular skill or framework.

Video Version: this post is based on (and expands upon) my keynote at AI Engineer World's Fair 2026.

Thanks to @addyosmani, @CyrusNewDay, @HamelHusain, @zeeg, @dillon_mulroy, @nayshins, and @jeffreyhuber for feedback on this post.

An aside: this has nothing to do with vibe coding

Addy Osmani detangled this thing that is worth highlighting:

A developer vibe-coding a side project a dozen people will ever run, and a team keeping a ten-year-old enterprise system alive for another quarter, share almost no constraints worth naming, and most of the advice in circulation is really one of those two people telling the other how to live.

If you love vibe coding, please, go on vibing. I still vibe code lots of things, I just also maintain lots of production software (and through HumanLayer, help 1000s of other engineers do the same), so the rest of this is aimed at folks solving hard problems in complex codebases.

I hear the word brownfield a lot to talk about this split. Historically that meant some ten-year-old Java thing, but at the pace we can ship now, it feels like an agent-built codebase starts to struggle after maybe three to six months -- you start to slow down, and the way you approach adding new things has to change.

A brief history of the software factory

I've been building and studying software factories my whole career, but I only learned this recently: the term traces all the way back to a NATO conference in 1968 -- the same one that gave us "software engineering."

The only other bit I find super interesting since then is that the US Department of Defense wrote a 31-page pdf about how the DoD needs to start using jenkins better or something.

The 2022 software factory

Let's ground our "software factory" definition around 2022, right before AI. In a typical software factory:

  • People decide what to build -- engineers, PMs, leadership driving the vision

  • It goes in a tracker -- Linear, Jira, whatever: a state machine of what needs to happen

  • Someone grabs a ticket and builds it -- probably does some manual/automated testing while they're at it

  • Pull request -- automated checks, a human reviews the code, maybe someone pulls it down to test

  • Anything wrong? Loop back to "someone builds the thing"

  • Ship to prod -- and it makes contact with users

  • Add monitoring -- there's an entire industry built around paging an engineer at 3am when something breaks

  • Users complain -- ask for things, find bugs, file feature requests → back to the team to add to the tracker

https://hlyr.dev/ace

And on and on. We haven't even hit AI yet, and there are already several loops in this picture.

front-loading alignment

There's a thing teams figured out decades ago: building takes hours or days, and so does review.

https://en.wikipedia.org/wiki/Mutation_testing

So we front-load the work -- planning, architecture proposals, sprint planning -- together, as a team. That means:

  • less rework, because we aligned before anyone wrote code

  • less time reviewing every line, if you've ever read a long-but-well-done PR, you know how fast the review goes when it's close-to-perfect

https://www.youtube.com/watch?v=am_oeAoUhew

We'll come back to this later - let's look at what happens when you bring agentic coding into the picture.

The agentic software factory

Now every company and their mother --

  • Ramp

  • Stripe

  • WorkOS

  • Brex

has spent the better part of this year explaining how they built an agent factory that ships on the order of 75% of their code.

The agentic factory looks mostly like swapping "someone builds the thing" → "an agent builds the thing" -- there's some stuff here like orchestration, a harness, a sandbox, a model, computer use, etc. I won't go in depth on those details because quite frankly I'm sick of reading about it and I'm sure you are too.

https://www.youtube.com/watch?v=Ib5GBkD555M

When the agent builds the thing:

  • Building drops from hours or days to minutes or hours.

  • Review still takes hours or days. A human still has to read the code and test the change. So review is now the bottleneck.

https://www.swe-marathon.org/

So you speed review up too:

  • Agentic code review, to catch style, bugs, security.

  • Agentic regression testing, to poke it from the outside with browsers and computer use and maybe send you a cute little video when it's done

https://www.oreilly.com/library/view/clean-code-a/9780136083238/

Review is faster now, but it's also probably still the bottleneck. But we can do more loops.

Next you might route incidents into the factory. Instead of paging someone at 3am, they wake up to a PR that maybe already fixes it.

We can also route user feedback into the factory. People ask for stuff, it gets built.

https://www.youtube.com/watch?v=q-ntX4DLW_c

At which point the job is two questions: how much can you stuff into the queue, and how fast can you review and test what comes out?

Which brings us to the lights-off software factory.

The lights-off software factory

Dan Shapiro coined this term and Simon Willison wrote about StrongDM's implementation of it -- where we no longer read the code.

You look at your beautiful software factory. It's ruined by that annoying little code review step and you say: you know what, that thing where a human reads every change? No thanks.

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#user-content-fn-5-f56fe9a973fe7c6ebb6a9673c1bc64cb

So you drop it, and you put the effort somewhere else:

  • Invest in testing and letting the agent test its own work

  • Invest in sandboxes and orchestration

  • Invest in automated review

  • Invest in monitoring

  • Invest in rollout

  • Invest in collecting feedback signals from users

https://garryslist.org/posts/boil-the-ocean

And now the job really is just one question: how much stuff can we ask the agent to build? How much of the ocean do we want to boil?

This is going to go great (its not)

https://stripe.dev/blog/minions-stripes-one-shot-end-to-end-coding-agents

I'm going to posit something potentially controversial: the lights off factory does not work.

Let's get into why software factories fail.

We tried this

In July 2025 we went full lights-off. Just read the specs and the tickets, background agents for all the small/medium stuff, the whole thing.

If you've tried this seriously for a few months, you already know how it ends. You find at least one issue gnarly enough that the agent can't solve it -- even with your most advanced prompting and workflows.

  • You do deep context-aware research, collating all the right parts into the smart zone for the model to analyze

  • You have the agent try to reproduce in 10 different ways

Eventually you have to suck it up and go dig into the codebase you stopped reading three months ago, trying to figure out what's broken.

And in the meantime:

  • Your site was down.

  • Your users were pissed.

  • And you, if you're anything like me, were miserable -- reading all the slop code you let slip into your system.

The first time this happened to us, I shook it off. Even though I'd just spent the better part of two weeks digging through claude spaghetti, "the downside risk was worth the velocity". By the ~third time in november, we decided it would be easier to rewrite from scratch, and my cofounder spent two whole weeks in VS Code (not even cursor) plumbing out all the patterns by hand.

models degrade codebase quality over time

What I want to get to is this: models have a shortcoming. They can't maintain and improve codebase quality over time -- not without a decent amount of human steering.4

When I say maintainability, I mean the specific thing where it becomes really, really hard to change one part of the codebase without breaking another part. This is Martin Fowler's shotgun surgery.

I'm not going to say much more about maintainability. There are a bunch of books you can go read about it

  • John Ousterhout's A Philosophy of Software Design

  • Robert C. Martin's Clean Code

  • Martin Fowler's Refactoring

So, why can't models do software maintainability?

"But surely the models have gotten better since then"

At this point you might be dying to say: but Dex, surely the models have gotten much better since July

They have -- in some ways. In others they're about the same.

  • Solving one-off problems, or vibe-coding a new marketing site? Yes. Way better.

  • Improving codebase quality over time? Not much better, as far as I can tell.

https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/wsff.md#user-content-fn-1b-f56fe9a973fe7c6ebb6a9673c1bc64cb

I can't prove this. You can't prove it either. There are no good benchmarks for a model's ability to maintain codebase quality. (More on where that's going later.)

**THERE ARE NO GOOD BENCHMARKS for a model's ability to maintain codebase quality **

But if you've worked with coding agents for a while -- and a lot of people are posting about exactly this -- you probably have the vibe already: they tend to make things worse over time, and make the codebase harder to work in.

So to figure out why this happens, I want to zoom out to the first great coding agent.

Claude Code won because of Reinforcement Learning inside the harness

Claude Code went from nothing to ~$4B -- now something like ~$9B -- in revenue in under a year.

Which is a little wild, because there were already great CLI agents. aider, cline, codebuff -- all predated Claude Code, all with genuinely great context engineering built in, all with the same tool set you might attribute to claude code: read, write, edit, grep, bash. I used them. They were good. But also, tool use would just... fail sometimes -- you'd watch it flail at the same edit three times and open your editor back up to do it yourself.

The SWE-Agent paper from 2024 outlines how small changes in tool shape make noticeable differences, e.g. including line numbers in ReadFile results, or changing an Edit tool from find/replace to line-range edits.

https://web.stanford.edu/~ouster/cgi-bin/aposd.php

Then Claude Code launched and went vertical pretty quickly. You can hand-wave this as distribution, but the canonically-accepted explanation is that claude code won because it was better, and that it was better because Anthropic RL'd the model inside the harness -- the first time a lab trained a model against the exact tools they were going to ship it with. And it got really, really good at calling those tools in an agentic loop.

It's one thing to fiddle with tool definitions and evals until you find the shape the model likes best -- I've burned weeks doing this for various use cases. It's a different game when you own the weights and can modify the model itself to be better at a particular set of tools.

The OpenAI team gave a talk in November that put this pretty well: if you build a harness but you don't own the weights and can't RL the model inside it, you'll always be at a disadvantage to a team that owns both.

Coding Agent RL in 60 seconds

I did a bunch of research on this topic and cooked up a bunch of visualizations to try to explain the parts that matter, but I found that Calvin French-Owen (MTS on the codex team, founder of Segment) did a talk at AI Council that did a much better and cleaner job, so I'm just gonna drop this animation here inspired by his slides:

https://github.com/opendilab/awesome-RLVR

To make a model better at coding, you're gonna:

  1. generate some coding agent traces to solve a problem (e.g. fix my tests)

  2. score the traces based on some criteria (verifier)

  3. update the model weights to make the good traces more likely, and the bad traces less likely

And then you do this millions of times over the course of weeks or months.

The "scoring" part of these things can tend to be whimsically one-dimensional though.

There's no penalty for bad design

Take SWE-bench Multilingual. The tasks are small -- about fifteen minutes of work apiece -- scraped out of open-source repos like Redis, jq, and Django. The reward is one or zero based on:

  • FAIL_TO_PASS - did you fix the thing you were asked to fix?

  • PASS_TO_PASS - did you do it without breaking anything else?

Here's a real one, fastlane__fastlane-19304, from fastlane -- a Ruby project. Its zip action grabs two optional params and calls .empty? on them straight away, so the moment you leave include and exclude off, it falls over:

https://factory.strongdm.ai/

The human fix that closed this particular issue is two lines (default nils to empty arrays):

https://www.latent.space/p/brex

During the evaluation, the model

  1. starts from a base commit -- the repo checked out to the moment right before that fix landed

  2. the bug report - in this case 'zip_command': undefined method 'empty?' for nil:NilClass

The agent goes off and writes some code based on the issue. It doesn't see the golden patch or the test patch that serves as the grader:

Then:

  1. We keep whatever patch it produced, then

  2. Throw away any edits it made to the test files (we've caught a model quietly commenting out the failing test or splicing in a mock that makes the test useless)

  3. Apply the benchmark's test patch on top, and

  4. Run the whole suite: the existing zip tests (PASS_TO_PASS) plus the new one (FAIL_TO_PASS) to see if they both pass

https://hlyr.dev/nva

Aside - Benchmarks are not verifiers - in fact they have to be held out from each other (don't train on test, yada yada) - I primarily mean this to convey the shape of "judging the quality of a coding agent trace" and its limitations.

How the model got to a correct answer doesn't matter. If the tests pass, we win, but there is no penalty for eroding codebase maintainability.

there is no penalty for eroding codebase maintainability

That's how you get try catches around everything:

http://homepages.cs.ncl.ac.uk/brian.randell/NATO/nato1968.PDF

Verifying quality is orders of magnitude harder than "did the tests pass"

Running the tests gets you a clean pass or fail in ~seconds. That's why RL can run millions of loops to optimize each model generation.

But the cost function of bad architecture is measured in weeks, months, maybe even years. It happens the first time someone opens that file for a one-line change and realizes they can't make it in one line -- that someone vibed this a little too hard, and now we have to make the same edit in eleven places and hope nothing quietly breaks three files over.

https://www.youtube.com/watch?v=RjfbvDXpFls

Tests give you feedback in seconds, but the cost function of bad architecture is measured in weeks, months, maybe even years

Bad design is the one thing today's benchmarks can't evaluate. And I know, I know, RL != Benchmarks, but if this was solved in RL, I'm pretty sure it would start to show up in how our benchmarks are designed too.

In any case, I personally don't trust any improvements on today's benchmarks as an indicator that the models are suddenly good at not slopping up your codebase.

The frontier is getting better, slowly

Of course lots of smart folks are working on this. My point is not that it can't be done, it's that the hype is outrunning the discipline.

A few efforts I think are pointed the right way:

  • SWE-Marathon (Abundant AI): ~400-hour tasks like "clone all of Excel, every feature" -- with a compound reward channel instead of a single pass/fail bit

  • DeepSWE (Datacurve): big tasks on OSS repos that were never actually built in the real world, so by construction they can't already be sitting in the training set (solves contamination, but not quality)

  • Frontier Code (Cognition): multi-PR tasks, and a clever move that evaluates quality deterministically -- it penalizes the model for writing tests that don't fail on the pre-patch code (if you've never heard about mutation testing you are in for a fun ride5). It also runs a judge model over the diff checking code-quality rules.

https://x.com/jeffreyhuber

But a model judging quality can only go so far.

In fact, it's not hard to imagine that if a model could reliably tell good code from bad, it might have written the good version to begin with. RL needs a fast+reliable oracle, and we don't yet have one for maintainability

if a model could reliably tell good code from bad, it might have written the good version to begin with, but maintainability has no fast oracle, so we can't reward for it during RL

Of course, more review agents and more tokens do help -- they raise the floor, catching the dumb stuff.

But they don't move the ceiling, because the ceiling is whatever we managed to teach the model in RL, and good design is the thing we still don't know how to teach it.

So I still wouldn't bet my codebase on any of these. But they're the first evals I've seen even trying to score maintainability instead of stopping at pass/fail.

Aside Maybe a future model just gets this and we can stop. If you want to yolo prompts until GPT-7 ships and find out, be my guest -- but bitter lesson be damned, we've got problems to solve now, and I'm gonna walk through how we do that.

Turning the lights back on

Today I learned that Twitter Articles have a "media limit" which means the rest of this is going into a part II post - stay tuned

📋 讨论归档

讨论进行中…