每日 Harness 开源 · Source
返回本期 · Back to 2026-06-19

论文 · Papers2026-06-19 · Friday, June 19, 2026

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

arxiv.org原文 ↗

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
论文指出 RL rollout 的瓶颈不是普通 serving 场景的固定模型推理,而是策略持续变化、高温长生成和 batch size 逐渐缩小。EfficientRollout 从目标模型诱导量化 drafter,并用系统感知开关只在 memory-bound 阶段启用 self-speculative decoding。它最高减少 19.6% rollout latency 和 12.7% end-to-end latency,同时保持最终模型质量,适合看 RL 训练系统优化。
浏览

评论 · Comments