每日 Harness 开源 · Source
返回本期 · Back to 2026-06-07

论文 · Papers2026-06-07 · Sunday, June 7, 2026

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges

arxiv.org原文 ↗

评测方法其他垂直
Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges
作者发现 LLM judge 在重复或中性复评时很稳定,但初判之后遭遇有目标的挑战会显著可逆。实验覆盖 MT-Bench 与 AlpacaEval,并用 anti-baseline challenge 与 counterbalanced target-validation 区分普通纠错和方向性操控。文章提出 ERS,把 reversal susceptibility 和 directional effects 合并度量,补上了静态 judge agreement 之外的一块评测盲区。
浏览

评论 · Comments