RoPoLL 把 LLM-as-judge 多评审团放进 Huber contamination 模型,证明普通 PoLL 共识在任何正污染下都可能因为单个偏置 judge 产生无界 bias。方法上它不换 jury,只把聚合函数换成鲁棒均值估计,实例化为 tuning-free 的 geometric median,并给出有限样本误差界和 minimax lower bound。实验覆盖 13 个 4B 到 675B 的开源 judge、3 个 reward-model benchmark 和最高 50% 的 corruption rate;3-judge、38B 的 RoPoLL committee 在 HelpSteer-2 30% bimodal-random corruption 下还超过 675B Mistral-Large-3,说明评审聚合的统计鲁棒性不是边角问题。
–浏览
评论 · Comments