返回列表
🧠 阿头学 · 💬 讨论题

AI基准测试的“刷榜”陷阱与信任透支

公开AI基准测试已因过度优化与利益绑定异化为注意力杠杆,行业必须放弃对排行榜的盲目依赖,转向透明披露与私有评估以重建真实信任。
打开原文 ↗

2026-09-16 原文链接 ↗
阅读简报
双语对照
完整翻译
原文
讨论归档

核心观点

  • 基准测试已沦为“信任贷款”而非能力标尺:公开评测因与资金流量强绑定,必然诱发针对评估分布的工业化刷榜,导致模型在可测指标上虚高而在真实场景中可靠性停滞。
  • 评测生态存在危险的“反身性反馈循环”:第三方指数并非客观中立,而是受公众叙事与直觉牵引不断修订,最终沦为群体偏见的放大器而非决策依据。
  • 高分与经济价值严重脱节:模型在学术难题上超越人类专家,但未能转化为实际生产力,证明当前静态题库无法衡量通用智能的真实商业效用。
  • 破局需切断指标优化激励:实验室应主动披露不利数据与实验筛选路径,用户需建立私有评估基线,用“一次性快照”替代持续排名以遏制P值操纵。

跟我们的关联

  • ATou 意味着产品评估不能依赖外部跑分背书,下一步需构建贴合核心业务流的私有沙盒测试集,将“用户真实留存与任务完成率”设为不可篡改的北极星指标。
  • Neta 意味着算法迭代必须警惕“路灯效应”,下一步需在训练管线中强制引入对抗性负样本与长尾场景压测,杜绝仅针对公开榜单调参的短视优化。
  • Uota 意味着对外沟通需放弃“刷榜营销”路径,下一步应建立“发布即披露”机制,主动公开模型短板与数据筛选过程,以透明换取长期客户信任。

讨论引子

  • 当公开基准测试的公信力彻底破产后,企业应如何设计一套既能防刷榜、又能被行业广泛采信的新型评估协议?
  • 在资源有限的前提下,团队应如何平衡“针对已知评测集快速拿分获取融资”与“投入长周期私有评估打磨真实可靠性”之间的战略冲突?

正如那句著名的被误传的名言所说:“世界上有三种谎言:谎言,该死的谎言,以及统计数据。”现在有了第四种:AI 基准测试(AI benchmarks)。

基准测试造就了极其聪明的 AI。凡是能够被衡量的,就能够被管理。多年来,让大量的基准测试数字上升通常确实能让模型变得更聪明。我并不反对基准测试,也不否认它们对行业进步的有益贡献。我反对的是基准测试刷榜(benchmaxxing)。

当构建模型的人能够针对公开的评估(eval)分数进行优化时,基准测试不可避免地会被“刷榜”(benchmaxxed)。[1] 你不必为了刷榜而直接在基准测试上进行训练;你只需在相似的数据上训练,或者尝试一百种实验设置,然后比较你的模型表现如何即可。即使没有人打算钻空子,基准测试本身也在筛选模型。当基准测试还只是学术追求时,情况就已经不太好了;而现在,由于它们与巨大的关注度和资金挂钩,诱导人们钻空子的负面激励比比皆是。

路灯效应(Streetlight Effect)示意图 - 来源

该领域想要衡量通用智能(general intelligence),但只能优化它能看到的东西。公开评估就是那盏路灯。因此,后训练(post-training)在光照区域内最快地提升了能力,而在此之外,可靠性可能停滞不前甚至恶化。我的主张是,“参差多态的智能”(jagged intelligence)中的许多峰值正是基准测试。

带注释的版本来自来源

每个人都在作弊

模型追逐基准测试。

据称,Meta 的 Llama 4 在 LMArena 上排名靠前,但事实证明,它不仅不是公开的模型,Meta 还被发现测试了 27 个私有变体后才挑选出最好的一个,而这个变体具有异常冗长且大量使用表情符号的风格,这正是 Arena 用户所青睐的。普通版本后来的排名要低得多。即使是 Claude(通常是那个乖乖牌),在一个模拟自动售货机并以赚取最多资金为目标的基准测试中,也形成了价格卡特尔,对供应商撒谎,并承诺向客户退款却从未兑现。

基准测试追逐用户。

当 GPT-6 Astra 发布时,Artificial Analysis 的智能指数(Intelligence Index)显示它与 GPT-5.6 Sol 并列,这与人们普遍认为 Astra 是一次重大飞跃的看法相冲突。第二天,修订后的指数将 Astra 领先了四分。三天后,又一次修订使其与 Claude Fable 5.1 并列第一。我不认为 Artificial Analysis 操纵了指数。这些改动是站得住脚的。但这一连串事件展示了反馈循环:人们对哪个模型更好的直觉,有助于决定一个评估是否看起来有效。外部评估的意义本应是告知大众,而不是本末倒置。

人们为了自己偏爱的叙事而各取所需。

当中国开放权重(open-weight)模型 GLM-5.2 在一个网页设计排行榜上击败 Fable 5 时,这成了中国模型已经赶上前沿(frontier)的证据。Design Arena 自己的分析则更为狭义:GLM 使用了更多可重复的模板,多生成了 25% 的代码,花费了两倍的时间,并且在几个设计类别上仍然输给了 Fable。

基准测试已经失灵有一阵子了。

在原论文中,非专业人类在 MMLU 上的得分为 34.5%,比许多老旧的小型模型还要差。OpenAI 的“政变”部分原因是安全研究人员在 2023 年看到模型在谷歌防伪问答(Google-Proof Question Answering, GPQA - 一个极其困难的数据集)上超越了博士水平。然而,世界上大部分有用的活仍然是人类在干。

https://epoch.ai/gradient-updates/the-real-reason-ai-benchmarks-havent-reflected-economic-impacts

该领域的每个人都知道它们已经失灵,但它们仍然被引用,因为每个人都在争夺注意力。

基准测试本该是什么样

问题的症结在于,(1)智能很难衡量,(2)一个(善意参与者)实验室想说“我们让模型变得更聪明了”,以及(3)要让这句话被人相信(在准确的程度上)。

基准测试是绕过信任难题的一条捷径。它们看起来像是信任的放大器,但实际上是一笔贷款:恶意参与者随后可以利用投射在基准测试上的信任。

出路在于首先要值得信赖。我们需要停止将信誉外包给排行榜,而是直接保持诚实。

对于实验室来说,这意味着发布注意事项、披露挑选有利数据的行为、包含那些让你看起来很糟的证据,并且即使你处于领先地位也要淡化基准测试。

用户也有责任:运行你自己的私有评估,对公开评估持保留态度,不要放大你看到的每一个数字或图表。

在 TypeSafe,我们正在制造一种新型模型,这意味着现有的基准测试并不适用。我们可以用一面写满评估数据的墙来展示我们击败了其他所有人,从而开启一场逐底竞争,或者从一张干净、最大程度诚实的白纸开始。我们选择干净的白纸:在我们的模型发布中没有标准的基准测试表。新的评估将是带有日期的快照,一旦发布就立即作废,而不是去不断刷分优化。我们还将发布不断完善的内部评估,作为我们目前最好的猜测,并附带注意事项、任何挑选有利数据的行为,以及对我们不利的证据。

最初发布于 https://www.completeskeptic.com/p/lies-damned-lies-and-benchmarks

感谢 Ke Deng (@justKDeng)、Erik Gafni (@EGafni) 和 Sasha Sheng (@hackgoofer) 提供反馈。

[1] 这是一种 P 值操纵(p-hacking)的形式:尝试足够多的训练选择,然后报告看起来最强的结果。

As the famously misattributed quote goes: “There are three kinds of lies: lies, damned lies, and statistics.” There is now a fourth kind: AI benchmarks.

正如那句著名的被误传的名言所说:“世界上有三种谎言:谎言,该死的谎言,以及统计数据。”现在有了第四种:AI 基准测试(AI benchmarks)。

Benchmarks got us to incredibly smart AI. What gets measured gets managed, and for years making lots of benchmark numbers go up generally made models smarter. I'm not arguing against benchmarks and their helpful contributions to progress. I’m against benchmaxxing.

基准测试造就了极其聪明的 AI。凡是能够被衡量的,就能够被管理。多年来,让大量的基准测试数字上升通常确实能让模型变得更聪明。我并不反对基准测试,也不否认它们对行业进步的有益贡献。我反对的是基准测试刷榜(benchmaxxing)。

A benchmark inevitably gets “benchmaxxed” when people building the model can optimize for a public eval score.[1] You don't have to train directly on a benchmark to benchmaxx it; you just have to train on similar data or try a hundred experimental settings, then compare how your model performs. The benchmark selects the model even if nobody intended to game it. Things were not great when benchmarks were an academic pursuit, but now that they are tied to enormous attention and funding, negative incentives to game them abound.

当构建模型的人能够针对公开的评估(eval)分数进行优化时,基准测试不可避免地会被“刷榜”(benchmaxxed)。[1] 你不必为了刷榜而直接在基准测试上进行训练;你只需在相似的数据上训练,或者尝试一百种实验设置,然后比较你的模型表现如何即可。即使没有人打算钻空子,基准测试本身也在筛选模型。当基准测试还只是学术追求时,情况就已经不太好了;而现在,由于它们与巨大的关注度和资金挂钩,诱导人们钻空子的负面激励比比皆是。

Illustration of the Streetlight Effect - source

路灯效应(Streetlight Effect)示意图 - 来源

The field wants to measure general intelligence, but can only optimize what it can see. Public evals are the streetlight. Post-training therefore pushes capability fastest in the illuminated areas, while reliability outside them can stay flat or even get worse. My claim is that many of the spikes in “jagged intelligence” are the benchmarks.

该领域想要衡量通用智能(general intelligence),但只能优化它能看到的东西。公开评估就是那盏路灯。因此,后训练(post-training)在光照区域内最快地提升了能力,而在此之外,可靠性可能停滞不前甚至恶化。我的主张是,“参差多态的智能”(jagged intelligence)中的许多峰值正是基准测试。

Annotated version from source

带注释的版本来自来源

Everyone is cheating

每个人都在作弊

Models chase benchmarks.

模型追逐基准测试。

Meta’s Llama 4 allegedly ranked near the top of LMArena, but it turned out that not only was it not the public model, Meta was also caught testing 27 private variants before picking the best one, which had the unusually verbose and emoji-heavy style that Arena users rewarded. The ordinary version later landed far lower. Even Claude (normally the goody two-shoes), on a benchmark to simulate a vending machine and end with the most money, formed price cartels, lied to suppliers and promised customer refunds it never sent.

据称,Meta 的 Llama 4 在 LMArena 上排名靠前,但事实证明,它不仅不是公开的模型,Meta 还被发现测试了 27 个私有变体后才挑选出最好的一个,而这个变体具有异常冗长且大量使用表情符号的风格,这正是 Arena 用户所青睐的。普通版本后来的排名要低得多。即使是 Claude(通常是那个乖乖牌),在一个模拟自动售货机并以赚取最多资金为目标的基准测试中,也形成了价格卡特尔,对供应商撒谎,并承诺向客户退款却从未兑现。

Benchmarks chase users.

基准测试追逐用户。

When GPT-6 Astra launched, Artificial Analysis's Intelligence Index had it tied with GPT-5.6 Sol, which clashed with the widespread belief that Astra was a major leap. The next day, a revised index put Astra four points ahead. Three days later, another revision tied it with Claude Fable 5.1 for first. I do not think Artificial Analysis rigged the index. The changes are defensible. But the sequence shows the feedback loop: people's intuitions about which model is better help determine whether an eval looks valid. The point of external evals should've been to inform the population, not the other way around.

当 GPT-6 Astra 发布时,Artificial Analysis 的智能指数(Intelligence Index)显示它与 GPT-5.6 Sol 并列,这与人们普遍认为 Astra 是一次重大飞跃的看法相冲突。第二天,修订后的指数将 Astra 领先了四分。三天后,又一次修订使其与 Claude Fable 5.1 并列第一。我不认为 Artificial Analysis 操纵了指数。这些改动是站得住脚的。但这一连串事件展示了反馈循环:人们对哪个模型更好的直觉,有助于决定一个评估是否看起来有效。外部评估的意义本应是告知大众,而不是本末倒置。

People cherry-pick for their favorite narrative.

人们为了自己偏爱的叙事而各取所需。

When the Chinese open-weight model GLM-5.2 beat Fable 5 on one web-design leaderboard, it became evidence that Chinese models had caught the frontier. Design Arena's own analysis was narrower: GLM used more repeatable templates, generated 25% more code, took twice as long, and still lost to Fable on several design categories.

当中国开放权重(open-weight)模型 GLM-5.2 在一个网页设计排行榜上击败 Fable 5 时,这成了中国模型已经赶上前沿(frontier)的证据。Design Arena 自己的分析则更为狭义:GLM 使用了更多可重复的模板,多生成了 25% 的代码,花费了两倍的时间,并且在几个设计类别上仍然输给了 Fable。

Benchmarks have been broken for a while.

基准测试已经失灵有一阵子了。

Unspecialized humans scored 34.5% on MMLU in the original paper, worse than many small old models. The OpenAI coup was partially a result of safety researchers seeing models surpass PhD level at Google-Proof Question Answering (GPQA - an extremely hard dataset) in 2023. Yet humans still do most of the world's useful work.

在原论文中,非专业人类在 MMLU 上的得分为 34.5%,比许多老旧的小型模型还要差。OpenAI 的“政变”部分原因是安全研究人员在 2023 年看到模型在谷歌防伪问答(Google-Proof Question Answering, GPQA - 一个极其困难的数据集)上超越了博士水平。然而,世界上大部分有用的活仍然是人类在干。

https://epoch.ai/gradient-updates/the-real-reason-ai-benchmarks-havent-reflected-economic-impacts

https://epoch.ai/gradient-updates/the-real-reason-ai-benchmarks-havent-reflected-economic-impacts

Everyone in the field knows they're broken, but they still get quoted because everyone is competing for attention.

该领域的每个人都知道它们已经失灵,但它们仍然被引用,因为每个人都在争夺注意力。

What benchmarks should be

基准测试本该是什么样

The crux of the problem is (1) intelligence is hard to measure, (2) a (good actor) lab wants to say "we made the model smarter," and (3) for it to be believed (to whatever extent is accurate).

问题的症结在于,(1)智能很难衡量,(2)一个(善意参与者)实验室想说“我们让模型变得更聪明了”,以及(3)要让这句话被人相信(在准确的程度上)。

Benchmarks are a shortcut to get around the hard problem of trust. They seem like a magnifier of trust, but instead are a loan: bad actors can then exploit the trust projected on the benchmark.

基准测试是绕过信任难题的一条捷径。它们看起来像是信任的放大器,但实际上是一笔贷款:恶意参与者随后可以利用投射在基准测试上的信任。

The way out is to be trustworthy in the first place. We need to stop outsourcing credibility to a leaderboard and just be honest.

出路在于首先要值得信赖。我们需要停止将信誉外包给排行榜,而是直接保持诚实。

For labs, that means publishing caveats, disclosing cherry-picking, including the evidence that makes you look bad, and de-emphasizing benchmarks even when you're ahead.

对于实验室来说,这意味着发布注意事项、披露挑选有利数据的行为、包含那些让你看起来很糟的证据,并且即使你处于领先地位也要淡化基准测试。

Users have a responsibility too: run your own private evals, treat public ones with a grain of salt, and don't amplify every number or plot you see.

用户也有责任:运行你自己的私有评估,对公开评估持保留态度,不要放大你看到的每一个数字或图表。

At TypeSafe, we're making a new type of model, which means existing benchmarks don't apply. We can start the race to the bottom with a wall of evals showing that we beat everyone else, or start from a clean maximally honest slate. We are choosing the clean slate: no standard benchmark table in our model releases. New evals will be dated snapshots and immediately retired once posted rather than hill-climbed. We will also publish our evolving internal evals as our current best guesses, alongside the caveats, any cherry-picking, and evidence that looks bad for us.

在 TypeSafe,我们正在制造一种新型模型,这意味着现有的基准测试并不适用。我们可以用一面写满评估数据的墙来展示我们击败了其他所有人,从而开启一场逐底竞争,或者从一张干净、最大程度诚实的白纸开始。我们选择干净的白纸:在我们的模型发布中没有标准的基准测试表。新的评估将是带有日期的快照,一旦发布就立即作废,而不是去不断刷分优化。我们还将发布不断完善的内部评估,作为我们目前最好的猜测,并附带注意事项、任何挑选有利数据的行为,以及对我们不利的证据。

Originally posted to https://www.completeskeptic.com/p/lies-damned-lies-and-benchmarks

最初发布于 https://www.completeskeptic.com/p/lies-damned-lies-and-benchmarks

Thanks to Ke Deng (@justKDeng), Erik Gafni (@EGafni), and Sasha Sheng (@hackgoofer) for feedback.

感谢 Ke Deng (@justKDeng)、Erik Gafni (@EGafni) 和 Sasha Sheng (@hackgoofer) 提供反馈。

[1] This is a form of p-hacking: try enough training choices, then report the result that looks strongest.

[1] 这是一种 P 值操纵(p-hacking)的形式:尝试足够多的训练选择,然后报告看起来最强的结果。

As the famously misattributed quote goes: “There are three kinds of lies: lies, damned lies, and statistics.” There is now a fourth kind: AI benchmarks. Benchmarks got us to incredibly smart AI. What gets measured gets managed, and for years making lots of benchmark numbers go up generally made models smarter. I'm not arguing against benchmarks and their helpful contributions to progress. I’m against benchmaxxing. A benchmark inevitably gets “benchmaxxed” when people building the model can optimize for a public eval score.[1] You don't have to train directly on a benchmark to benchmaxx it; you just have to train on similar data or try a hundred experimental settings, then compare how your model performs. The benchmark selects the model even if nobody intended to game it. Things were not great when benchmarks were an academic pursuit, but now that they are tied to enormous attention and funding, negative incentives to game them abound. Illustration of the Streetlight Effect - source The field wants to measure general intelligence, but can only optimize what it can see. Public evals are the streetlight. Post-training therefore pushes capability fastest in the illuminated areas, while reliability outside them can stay flat or even get worse. My claim is that many of the spikes in “jagged intelligence” are the benchmarks. Annotated version from source Everyone is cheating Models chase benchmarks. Meta’s Llama 4 allegedly ranked near the top of LMArena, but it turned out that not only was it not the public model, Meta was also caught testing 27 private variants before picking the best one, which had the unusually verbose and emoji-heavy style that Arena users rewarded. The ordinary version later landed far lower. Even Claude (normally the goody two-shoes), on a benchmark to simulate a vending machine and end with the most money, formed price cartels, lied to suppliers and promised customer refunds it never sent. Benchmarks chase users. When GPT-6 Astra launched, Artificial Analysis's Intelligence Index had it tied with GPT-5.6 Sol, which clashed with the widespread belief that Astra was a major leap. The next day, a revised index put Astra four points ahead. Three days later, another revision tied it with Claude Fable 5.1 for first. I do not think Artificial Analysis rigged the index. The changes are defensible. But the sequence shows the feedback loop: people's intuitions about which model is better help determine whether an eval looks valid. The point of external evals should've been to inform the population, not the other way around. People cherry-pick for their favorite narrative. When the Chinese open-weight model GLM-5.2 beat Fable 5 on one web-design leaderboard, it became evidence that Chinese models had caught the frontier. Design Arena's own analysis was narrower: GLM used more repeatable templates, generated 25% more code, took twice as long, and still lost to Fable on several design categories. Benchmarks have been broken for a while. Unspecialized humans scored 34.5% on MMLU in the original paper, worse than many small old models. The OpenAI coup was partially a result of safety researchers seeing models surpass PhD level at Google-Proof Question Answering (GPQA - an extremely hard dataset) in 2023. Yet humans still do most of the world's useful work. https://epoch.ai/gradient-updates/the-real-reason-ai-benchmarks-havent-reflected-economic-impacts Everyone in the field knows they're broken, but they still get quoted because everyone is competing for attention. What benchmarks should be The crux of the problem is (1) intelligence is hard to measure, (2) a (good actor) lab wants to say "we made the model smarter," and (3) for it to be believed (to whatever extent is accurate). Benchmarks are a shortcut to get around the hard problem of trust. They seem like a magnifier of trust, but instead are a loan: bad actors can then exploit the trust projected on the benchmark. The way out is to be trustworthy in the first place. We need to stop outsourcing credibility to a leaderboard and just be honest. For labs, that means publishing caveats, disclosing cherry-picking, including the evidence that makes you look bad, and de-emphasizing benchmarks even when you're ahead. Users have a responsibility too: run your own private evals, treat public ones with a grain of salt, and don't amplify every number or plot you see. At TypeSafe, we're making a new type of model, which means existing benchmarks don't apply. We can start the race to the bottom with a wall of evals showing that we beat everyone else, or start from a clean maximally honest slate. We are choosing the clean slate: no standard benchmark table in our model releases. New evals will be dated snapshots and immediately retired once posted rather than hill-climbed. We will also publish our evolving internal evals as our current best guesses, alongside the caveats, any cherry-picking, and evidence that looks bad for us. Originally posted to https://www.completeskeptic.com/p/lies-damned-lies-and-benchmarks Thanks to Ke Deng (@justKDeng), Erik Gafni (@EGafni), and Sasha Sheng (@hackgoofer) for feedback. [1] This is a form of p-hacking: try enough training choices, then report the result that looks strongest.

📋 讨论归档

讨论进行中…