每日 Harness 开源 · Source
返回本期 · Back to 2026-07-02

论文 · Papers2026-07-02 · Thursday, July 2, 2026

ECHO: Prune to act, trace to learn with selective turn memory in agentic RL

arxiv.org原文 ↗

ECHO: Prune to act, trace to learn with selective turn memory in agentic RL
ECHO 把上下文裁剪和 RL credit tracing 放在同一个设计里处理。每个完成的 environment turn 被压缩成带 source index 的 memory record,策略上下文按需从这些记录重建;成功终局的正向 credit 再沿 source index 回流到支持答案的证据和选择动作。BrowseComp-Plus 上 ECHO held-out accuracy 达 43.4%,高于 GRPO 的 28.9% 和 rolling-summary baseline SUPO 的 36.1%,同时用更少 turns 和更低 trajectory volume,说明“可寻址压缩”比普通 summary 更适合训练闭环。
浏览

评论 · Comments