每日 Harness 开源 · Source
返回本期 · Back to 2026-07-07

博客文章 · Blog Posts2026-07-07 · Tuesday, July 7, 2026

Pruning RAG context down to what the answer actually needs

kapa.ai原文 ↗

Pruning RAG context down to what the answer actually needs
Kapa.ai 在 RAG pipeline 的 reranker 与 generator 之间加入一个小模型 pruner,让它同时看问题和所有候选 chunk,再按 1-5 等级判断哪些 chunk 真会参与答案。生产回放显示它丢掉约 68% context、保留约 96% recall,并在扣除自身调用成本后降低约 34% 查询费用。文章还指出 rerank score 不是跨 query 校准的绝对分数,因此简单阈值剪枝会误删组合相关的 chunk。
浏览

评论 · Comments