每日 Harness 开源 · Source
返回本期 · Back to 2026-06-19

论文 · Papers2026-06-19 · Friday, June 19, 2026

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents

arxiv.org原文 ↗

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents
SafeClawBench 用 600 个受控对抗任务覆盖直接/间接 prompt injection、tool-return injection、memory poisoning、memory extraction 和歧义推断等 6 类攻击。论文最有价值的数字是:12,000 行匹配分析中,347 个沙箱伤害有 291 个发生在语义检查通过的行里。它说明 tool-using agent 的安全评测不能只看模型文本是否“拒绝”,还要看审计证据和外部状态变化。
浏览

评论 · Comments