Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
arxiv.org原文 ↗
论文在 code-as-spec 设置中研究 benchmark construction validity:通过测试是否真的代表完成了用户请求。作者比较两个生产 Copilot CLI agents,摘要列出 claude-opus-4.7 和 gpt-5.5,并观察 agent 会交付被检查的内容,而不是需求中未被检查的部分。它把 coding agent 评测的焦点从 pass rate 转到规格覆盖率和测试诱导偏差。
–浏览
评论 · Comments