每日 Harness 开源 · Source
返回本期 · Back to 2026-06-12

论文 · Papers2026-06-12 · Friday, June 12, 2026

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

arxiv.org原文 ↗

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
TRACE 面向 RLVR 中 rollout 成本高但 reward contrast 不足的问题,特别是多轮 ReAct rollout 里同一个 terminal reward 会被所有中间决策共享。方法把每个 thought-action-observation turn 视为树中的语义节点,在 prompt roots 和 intermediate prefixes 上自适应分配 continuation 预算,优先探索最可能产生 mixed terminal rewards 的位置。等采样成本下,它让 Qwen3-14B 在 Multi-Hop QA 平均准确率比竞争 baseline 高 2.8 点。论文的价值在于把 rollout allocation 从 prompt 粒度推进到 prefix 粒度,使 outcome-only reward 能提供更密的 policy update 信号。
浏览

评论 · Comments