SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents
arxiv.org原文 ↗
SafeClawBench 用 600 个受控对抗任务覆盖直接/间接 prompt injection、tool-return injection、memory poisoning、memory extraction 和歧义推断等 6 类攻击。论文最有价值的数字是:12,000 行匹配分析中,347 个沙箱伤害有 291 个发生在语义检查通过的行里。它说明 tool-using agent 的安全评测不能只看模型文本是否“拒绝”,还要看审计证据和外部状态变化。
–浏览
评论 · Comments