Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems
arxiv.org原文 ↗
中间件在生成前做 NLI 事实核验、五信号投毒检测和带阻尼的 Trust Index,在 TruthfulQA/Llama 3.3 70B 得到 91% 准确率、100% precision、100% 注入 recall。三模型 ROC-AUC 0.73 - 0.81,但实体替换等语义弱化仍漏检,FEVER 结果较弱也迫使每个数据域重新标定阈值。
–浏览
评论 · Comments