The ways we contain Claude across products
anthropic.com原文 ↗
Anthropic 这篇工程文章讨论 claude.ai、Claude Code、Cowork 三种 agent 产品的 containment 架构。文中给出几个具体数字:Claude Code 用户批准约 93% 的 permission prompts,auto mode 可在执行前捕获约 83% overeager behaviors;Claude Opus 4.7 在 Gray Swan benchmark 单次攻击成功率约 0.1%,100 次 adaptive attempts 后约 5-6%。文章的判断是模型层防御必然有漏网率,环境隔离、外部内容权限和模型防御必须叠加。
–浏览
评论 · Comments