opus 5.5 用 medium 还是 high,我拿自己的 repo 测了一下
opus 5.5 出来当天,我拿自己 11 个 repo 里的 22 个真实提交做了一轮 medium 对 high 的配对测试: 两档完成度一样,high 慢三分之一、多花三成多 token,但盲评里多抓出几个测试查不出来的真 bug。 再对一下 Artificial Analysis 的 Intelligence Index 数据,最后把默认档定成了 high。 文末按不同用法给了推荐。
opus 5.5 出来当天,我拿自己 11 个 repo 里的 22 个真实提交做了一轮 medium 对 high 的配对测试: 两档完成度一样,high 慢三分之一、多花三成多 token,但盲评里多抓出几个测试查不出来的真 bug。 再对一下 Artificial Analysis 的 Intelligence Index 数据,最后把默认档定成了 high。 文末按不同用法给了推荐。
On launch day I replayed 22 real commits from 11 of my own repos, running Opus 5.5 at medium and high effort side by side. Both finished every task; high took about a third longer and 35% more tokens, but blind judges credited it with catching several bugs the tests couldn't see. Cross-checked against the Artificial Analysis Intelligence Index, I made high my default. Recommendations by usage profile at the end.
© Xingfan Xia 2024 - 2026 · CC BY-NC 4.0