每日 Harness 开源 · Source
返回本期 · Back to 2026-06-25

论文 · Papers2026-06-25 · Thursday, June 25, 2026

AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability

arxiv.org原文 ↗

AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability
AdversaBench 从 seed prompt 出发应用 5 种结构化 mutation,查询目标模型后再用三 judge 加 meta-judge 确认 failure。实验覆盖 reasoning、instruction following 和 tool use 三类共 45 个 seed,每个 seed 都能产生确认失败;作者还报告 judge pairwise agreement 为 80-87%,但 Cohen kappa 接近 0。这个细节很有用:高表面一致率不等于可靠裁判,red-teaming 自动化必须处理 label skew 和 judge 偏差。
浏览

评论 · Comments