TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
arxiv.org原文 ↗
TRACE 面向 RLVR 中 rollout 成本高但 reward contrast 不足的问题,特别是多轮 ReAct rollout 里同一个 terminal reward 会被所有中间决策共享。方法把每个 thought-action-observation turn 视为树中的语义节点,在 prompt roots 和 intermediate prefixes 上自适应分配 continuation 预算,优先探索最可能产生 mixed terminal rewards 的位置。等采样成本下,它让 Qwen3-14B 在 Multi-Hop QA 平均准确率比竞争 baseline 高 2.8 点。论文的价值在于把 rollout allocation 从 prompt 粒度推进到 prefix 粒度,使 outcome-only reward 能提供更密的 policy update 信号。
–浏览
评论 · Comments