AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability
arxiv.org原文 ↗
AdversaBench 从 seed prompt 出发应用 5 种结构化 mutation,查询目标模型后再用三 judge 加 meta-judge 确认 failure。实验覆盖 reasoning、instruction following 和 tool use 三类共 45 个 seed,每个 seed 都能产生确认失败;作者还报告 judge pairwise agreement 为 80-87%,但 Cohen kappa 接近 0。这个细节很有用:高表面一致率不等于可靠裁判,red-teaming 自动化必须处理 label skew 和 judge 偏差。
–浏览
评论 · Comments