任务域决策包
成本收益阈值约在 300 行 / 单次调用范围
c
我的 Python 函数第二次调用时返回错误结果
r
locate@code_engineering/single_file_debug
C'
Sonnet 4.6 足以处理约 300 行以内的孤立 bug。涉及共享状态或多步因果链时升级到 Opus 4.8。
@code_engineering/single_file_debug
成本收益阈值约在 300 行 / 单次调用范围
我的 Python 函数第二次调用时返回错误结果
locate@code_engineering/single_file_debug
Sonnet 4.6 足以处理约 300 行以内的孤立 bug。涉及共享状态或多步因果链时升级到 Opus 4.8。
bug is isolated, context fits one prompt, cost matters
claude/sonnet-4.6bug involves subtle state, multiple suspects, correctness is critical
claude/opus-4.8you need precise root-cause location, not symptom description
openai/gpt-5.5terminal-heavy debugging workflow
openai/gpt-5.5| ID | 关系 | 执行体 | 结果 | 输出 |
|---|---|---|---|---|
| CT-D1-001 | identify@single_file_debug | claude/opus-4.8 | 已接受 | 能准确分类 bug 类型,包括隐蔽状态问题 |
| CT-D1-002 | identify@single_file_debug | claude/sonnet-4.6 | 已接受 | 常见模式接近 Opus,且受开发者偏好 |
| CT-D1-003 | identify@single_file_debug | openai/gpt-5.5 | 有边界 | 分类准确,但有时停留在症状层面 |
| CT-D1-004 | validate@single_file_debug | claude/opus-4.8 | 已接受 | 比前代 Opus 更容易标记自身代码缺陷,且更能抵抗不诚实摘要 |
| CT-D1-005 | validate@single_file_debug | claude/sonnet-4.6 | 有边界 | 孤立 bug 可靠,但相比 Opus 4.8 更可能静默漏掉细微缺陷 |
Opus 4.8 is 4x less likely to miss code flaws than Opus 4.7
E-GLOBAL-004Sonnet 4.6 preferred by developers over prior flagship Opus 4.5 in controlled preference testing
Anthropic Sonnet 4.6 release, 2026GPT-5.5 weaker than Opus 4.8 on precise root-cause location vs symptom description in debugging
根据相邻证据推断: SWE-Bench Pro gap