Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
arxiv.org原文 ↗
研究证明跨多次无人值守迭代分散的攻击证据无法被任何单轨迹监控器可靠区分,几何衰减风险也只让攻击者等待固定时间。LoopHarness 保留非衰减循环状态,并在 mediated commit 与仲裁检测下把未授权不可逆动作期望数界定为 B+m-1+m/delta_M,与循环长度 N 无关;评测覆盖 Agent-SafetyBench、外状态攻击、消融和白盒红队。
–浏览
评论 · Comments