Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety
arxiv.org原文 ↗
论文指出 AI control eval 中攻击者若能选择何时开始或停止攻击,会显著降低测得安全性。作者把 attack selection 分成 start policy 和 stop policy,在 BashArena 与 LinuxArena 中测试;1% audit budget 下,start policy 在两个环境都把 safety 降低 20pp,stop policy 在 BashArena 降 20pp、LinuxArena 降 28pp。
–浏览
评论 · Comments