PrivacyAlign: Contextual Privacy Alignment for LLM Agents
arxiv.org原文 ↗
PrivacyAlign 把 agent 隐私问题具体化为“在什么对象、什么条件下可以分享什么”的情境判断,而不是简单的敏感词屏蔽。数据集有 1,350 个样本和 3,516 条详细标注,来自 599 名独立标注者,作者还把人类解释喂给 LLM judge,并用 annotation-conditioned reward modeling 训练小型开源 agent。看点在于它把隐私对齐的基准从代理标签拉回到人类规范本身,尤其适合评估会发消息、发帖、调工具的 agent。
–浏览
评论 · Comments