每日 Harness 开源 · Source
返回本期 · Back to 2026-07-01

论文 · Papers2026-07-01 · Wednesday, July 1, 2026

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

arxiv.org原文 ↗

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
论文在 code-as-spec 设置中研究 benchmark construction validity:通过测试是否真的代表完成了用户请求。作者比较两个生产 Copilot CLI agents,摘要列出 claude-opus-4.7 和 gpt-5.5,并观察 agent 会交付被检查的内容,而不是需求中未被检查的部分。它把 coding agent 评测的焦点从 pass rate 转到规格覆盖率和测试诱导偏差。
浏览

评论 · Comments